Tenta descrever o sabor de um sanduíche para a pessoa do caixa de um restaurante. Não apontar para a foto no cardápio, não escolher de uma lista, apenas usar suas palavras para fazer seu pedido: "Quero um pão macio, tipo brioche, com uma carne que tenha aquele gosto defumado mas não seco, um molho meio ácido mas não exatamente mostarda, e umas folhas que deem crocância mas que não sejam alface."
Por mais hábil que seja a pessoa que faz o sanduíche, o resultado ainda assim provavelmente vai ter erros e você vai ter que fazer vários pequenos ajustes, provavelmente enfurecer quem estiver na fila com fome atrás de você. Mas e se você só tivesse mostrado a foto do sanduíche que você queria?
Parece trivial, mas existe um longo debate no campo do aprendizado humano que compara a linguagem textual com a visual. Com o relativo retrocesso trazido pelas IAs, esse debate voltou com tudo, assim como os problemas antigos no desenvolvimento de software.
Agora imagine que em vez de sanduíche você precisa descrever verbalmente um sistema de atendimento ao cidadão de um serviço público brasileiro: "tem que ser fácil de acessar, mas não pode deixar vazar informação. Tem que seguir as leis, mas não pode ser muito rígido". Imagine você e seus colegas tendo que falar para os desenvolvedores todas as exceções, fluxos paralelos e contextos regionais que um sistema desses precisa. Só com palavras (ou pelo menos com a maioria sendo palavras). Pois foi assim que a maioria dos sistemas computacionais para usuários externos é criada até hoje na Engenharia de Requisitos.
Com a evolução do campo de conhecimento de Interação Humano-Computador, técnicas de ergonomia, psicologia e design passaram a ser integradas ao desenvolvimento de softwares. Assim, quando o design trouxe suas ferramentas para os projetos de sistemas, portou consigo técnicas de prototipagem e validação visual de conceitos muito úteis, mas que não foram completamente incorporadas pela Engenharia de Requisitos, ainda que tenham trazido melhorias.
Em The prompt is not an interface, Joshua Leigh reconstrói quarenta anos de evolução em design de interação para argumentar que a interface de prompt é um retrocesso para tarefas visuais e espaciais. O argumento se apoia na teoria de dupla codificação de Paivio (1971), segundo a qual o cérebro humano processa informação visual e verbal por sistemas cognitivos separados. O verbal opera em sequência, limitado pela linearidade da linguagem. O visual opera em paralelo, especializado em relações espaciais, processando mais informação ao mesmo tempo. Imagens também são lembradas mais facilmente do que palavras.
Pesquisas recentes confirmam e refinam essa distinção. Potter et al. (2014, MIT) mostraram que o cérebro consegue processar uma imagem em 13 milissegundos. Zeng et al. (2025) usaram fMRI com pictogramas chineses, estímulos visualmente idênticos interpretados ora como palavras ora como objetos, e confirmaram que os dois tipos de reconhecimento ativam mecanismos neurais distintos.
O retrocesso trazido pelos prompts de IA, portanto, está em voltarmos a depender predominantemente de texto para especificar o que queremos de um sistema, como se fazia na época do DOS. Na prática, isso significou relegar a comunicação visual a um papel secundário no contexto de sistemas computacionais. Como resultado, problemas que existiam desde a Engenharia de Requisitos agora estão sendo potencializados em escala sem precedentes.
O efeito do uso quase exclusivo ou preferencial de linguagem verbal como ferramenta de elicitação de requisitos e especificação já é um clássico na área de software. O NaPiRE, um levantamento global conduzido com 228 empresas em 10 países, identificou que requisitos incompletos são o problema número um da engenharia de requisitos, reportados por 48% dos respondentes, seguidos por falhas de comunicação (41%) e requisitos abstratos demais (33%) (Méndez Fernández et al., 2017). Além disso, requisitos mal escritos e ambíguos estão consistentemente entre as cinco principais causas de falha em projetos de software.
Com a IA, o prompt assumiu o papel do requisito. Chen et al. (2025) propõem uma taxonomia de defeitos em prompts e argumentam que eles funcionam como código-fonte em sistemas baseados em LLM, só que escritos em linguagem natural ambígua e executados num motor probabilístico. Os defeitos resultantes vão de problemas triviais de formatação a falhas de segurança e desinformação. Pesquisa recente também mostra que programadores iniciantes sistematicamente omitem detalhes cruciais nos prompts, e que a especificidade do prompt afeta diretamente a qualidade do código gerado (Lucchetti et al., 2025).
Ou seja, uma das principais dificuldades no desenvolvimento de sistemas é comunicação. Profissionais de tecnologia operam com linguagens e jargões diferentes entre si, e mais diferentes ainda das linguagens de quem requisitou o sistema e de quem vai usá-lo. Essa dificuldade está no centro de como as IAs populares de hoje foram desenvolvidas. Designers raramente participaram das decisões estruturais desses sistemas. O padrão se repete: são chamados para arrumar o que já foi construído, o mesmo papel que já tinham quando o produto era só uma interface. Via de regra, não estavam na sala quando as decisões estruturais foram tomadas.
As maiores empresas de IA do mundo, incluindo OpenAI, Google DeepMind e Anthropic, foram fundadas e são lideradas por pesquisadores de aprendizado de máquina e engenheiros de software. Nenhuma delas tem um designer entre seus fundadores. Vagas de design existem nessas empresas, mas como função de suporte ao produto, não de sistema. As decisões estruturais, da arquitetura dos modelos à escolha da interface de prompt como modo primário de interação, foram tomadas por pessoas para quem texto é a linguagem natural de trabalho. Designers entram depois, para deixar o que já foi construído com cara de produto coerente para o usuário final.
Portanto, não surpreende que a principal forma de interação com IAs hoje seja um prompt. LLMs são, literalmente, Language Models, construídos por e para pessoas para quem o texto é a ferramenta natural de pensamento. O terminal, para um engenheiro de software, não é um retrocesso. É o ambiente preferido. A interface de prompt não é um acidente, mas sim a consequência lógica de quem estava na sala quando as decisões foram tomadas.
Nesse sentido, Leigh apresenta o problema como uma dicotomia clara, em que engenheiros pensam de forma verbal discursiva e designers pensam em imagens. Como profissional que atua nos dois campos, discordo da polaridade. O que observo é que os recursos visuais e verbais de cada campo nem sempre se traduzem bem entre si.
O que observo na prática é que engenheiros também desenham o tempo todo para se fazer entender. Criam diagramas, representações, abstrações. Desenham no quadro branco, rabiscam em cadernos, mandam capturas de tela anotadas no chat e nos tickets do Jira. A linguagem predominante da engenharia não é só texto. Explicar uma arquitetura complexa sem desenhar um diagrama é impossível. Escrever requisitos só com imagens também é.
Da mesma forma, todo projeto de interface ou de experiência que já vi precisou de texto escrito em algum momento, seja para documentar iterações, instruir o uso do design system ou registrar comentários de quem avalia o trabalho. Quando atuei diretamente com design, nunca consegui me comunicar só com imagens.
Em suma, engenheiros também se comunicam visualmente, e designers também escrevem. A diferença entre um engenheiro e um designer não é que um pensa em palavras e o outro em imagens. É que eles foram ensinados a valorizar linguagens diferentes. Como consequência, tradução entre visual e textual sempre foi problemática. A IA não criou esse problema. Ela o amplificou ao obrigar todo mundo a ser capaz de descrever perfeitamente o que precisa a um sistema que pode falhar 60% das vezes em contextos complexos.
Leigh termina o artigo dele dizendo que o futuro da interação com IA não é digitar o que você quer, é mostrar o que você quer dizer. Me parece que esse sempre foi o caso e segue sendo. Eu sempre coloco uma imagem ou uma captura de tela para ajudar a IA a contextualizar o que pedi. Com técnicas como OCR e integração com Python, mostrar à IA o que você quer dizer já é uma realidade, não apenas uma promessa. Afinal, o prompt da IA é funcionalmente muito próximo de um terminal de computador, e o terminal sempre soube lidar com arquivos, imagens e dados além de texto.
Diante dessa possibilidade, a tentação é concluir que o problema é de quem usa a ferramenta. Não é, ou pelo menos não é só isso. Um sistema que depende de descrição verbal perfeita para funcionar e que erra até 60% das vezes em contextos complexos é um sistema com falhas de projeto. A interface de prompt como modo primário de interação foi uma escolha, não uma inevitabilidade. Interfaces visuais, manipulação direta, canvas com IA embutida, tudo isso já existe em estágio inicial.
Mas enquanto essas interfaces não se consolidam, e mesmo depois que se consolidarem, a capacidade de articular problemas complexos continua sendo uma competência humana insubstituível. Isso nos deixa com duas conclusões importantes.
A primeira é que as IAs podem e devem ser melhores. Sistemas construídos por times mais diversos, com designers envolvidos nas decisões estruturais desde o início, muito provavelmente produziriam interfaces e interações que ajudariam a melhorar as taxas de sucesso da IA em negócios.
A segunda é que, mesmo com interfaces e interações melhores, a pessoa precisa saber o que quer dizer antes de mostrar ou digitar qualquer coisa. Isso não é uma questão de ferramenta. É uma questão de formação, de repertório, de capacidade de operar em mais de uma linguagem ao mesmo tempo.
Enquanto as decisões estruturais de sistemas de IA continuarem sendo tomadas por times sem diversidade disciplinar, o resultado vai continuar refletindo a linguagem de quem estava na sala. Interfaces melhores vão chegar. IAs que praticamente não alucinam também. Mas ferramentas melhores não eliminam a necessidade de entender o problema, ter uma boa ideia do que deve ser a solução e aplicar tudo isso de forma bem-sucedida. A responsabilidade é dos dois lados.
Sem essa combinação, a IA não traduz entre linguagens. Ela só amplifica problemas preexistentes, quase sempre com um temperinho a mais de alucinação de IA.
Paivio, A. (1971). Imagery and Verbal Processes. Oxford University Press. taylorfrancis.com
Potter, M.C., Wyble, B., Hagmann, C.E. & McCourt, E.S. (2014). "Detecting meaning in RSVP at 13 ms per picture." Attention, Perception, and Psychophysics, 76(2), 270–279. springer.com
Zeng, J. et al. (2025). "Neural Distinction between Visual Word and Object Recognition." Journal of Neuroscience, 45(28). jneurosci.org
Méndez Fernández, D., Wagner, S. et al. (2017). "Naming the Pain in Requirements Engineering." Empirical Software Engineering, 22(5), 2298–2338. springer.com
Chen, Z. et al. (2025). "A Taxonomy of Prompt Defects in LLM Systems." arxiv.org
Lucchetti, A. et al. (2025). "More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation." arxiv.org
Leigh, J. (2026). "The prompt is not an interface." UX Collective / The Conditions. medium.com
AuthenHallu (2025). "Detecting Hallucinations in Authentic LLM–Human Interactions." arxiv.org
Try describing the taste of a sandwich to the person at the counter. Not pointing to the photo on the menu, not choosing from a list, just using your own words to place your order: "I want a soft bun, like brioche, with meat that has that smoky flavour but isn't dry, a sauce that's sort of tangy but not exactly mustard, and some leaves that add crunch but aren't lettuce."
However skilled the person making the sandwich, the result will probably still have errors and you'll have to make several small adjustments, probably infuriating whoever is standing hungry behind you in line. But what if you had just shown the photo of the sandwich you wanted?
It sounds trivial, but there is a long-running debate in the field of human learning comparing textual and visual language. With the relative regression brought by AI, that debate is back in full force, along with long-standing problems in software development.
Now imagine that instead of a sandwich you need to verbally describe a citizen service system for a Brazilian public agency: "it has to be easy to access, but can't leak information. It has to follow the law, but can't be too rigid." Imagine you and your colleagues having to tell the developers about every exception, parallel flow and regional context such a system requires. Only with words (or at least mostly words). That is how most software systems for external users are still built today in Requirements Engineering.
As the field of Human-Computer Interaction evolved, techniques from ergonomics, psychology and design were integrated into software development. When design brought its tools into system projects, it carried useful prototyping and visual concept validation techniques, but these were never fully incorporated into Requirements Engineering, even though they brought improvements.
In The prompt is not an interface, Joshua Leigh reconstructs forty years of interaction design evolution to argue that the prompt interface is a regression for visual and spatial tasks. The argument draws on Paivio's (1971) dual coding theory, according to which the human brain processes visual and verbal information through separate cognitive systems. The verbal system operates sequentially, constrained by the linearity of language. The visual system operates in parallel, specialised in spatial relationships, processing more information simultaneously. Images are also remembered more easily than words.
Recent research confirms and refines this distinction. Potter et al. (2014, MIT) showed that the brain can process an image in 13 milliseconds. Zeng et al. (2025) used fMRI with Chinese pictographs — visually identical stimuli interpreted as either words or objects — and confirmed that the two types of recognition activate distinct neural mechanisms.
The regression brought by AI prompts, therefore, lies in going back to depending predominantly on text to specify what we want from a system, as was done in the DOS era. In practice, this meant relegating visual communication to a secondary role in the context of computational systems. As a result, problems that have existed since the early days of Requirements Engineering are now being amplified at an unprecedented scale.
The effect of near-exclusive reliance on verbal language as a tool for requirements elicitation and specification is already a classic in the software field. NaPiRE, a global survey conducted with 228 companies across 10 countries, found that incomplete requirements are the number one problem in requirements engineering, reported by 48% of respondents, followed by communication failures (41%) and requirements that are too abstract (33%) (Méndez Fernández et al., 2017). Poorly written and ambiguous requirements are consistently among the top five causes of failure in software projects.
With AI, the prompt has assumed the role of the requirement. Chen et al. (2025) propose a taxonomy of prompt defects and argue that prompts function as source code in LLM-based systems, except they are written in ambiguous natural language and executed on a probabilistic engine. The resulting defects range from trivial formatting problems to security failures and misinformation. Recent research also shows that novice programmers systematically omit crucial details in their prompts, and that prompt specificity directly affects the quality of generated code (Lucchetti et al., 2025).
In other words, one of the main difficulties in systems development is communication. Technology professionals operate with different languages and jargons among themselves, and even more different from the languages of those who commissioned the system and those who will use it. This difficulty is at the heart of how today's popular AIs were developed. Designers rarely participated in the structural decisions behind these systems. The pattern repeats: they are called in to fix what has already been built, the same role they had when the product was just an interface. As a rule, they were not in the room when the structural decisions were made.
The world's largest AI companies, including OpenAI, Google DeepMind and Anthropic, were founded and are led by machine learning researchers and software engineers. None of them has a designer among its founders. Design roles exist at these companies, but as a product support function, not a system function. The structural decisions — from model architecture to the choice of the prompt interface as the primary mode of interaction — were made by people for whom text is the natural language of work. Designers come in afterwards, to make what has already been built look like a coherent product for the end user.
It is therefore no surprise that the primary form of interaction with AI today is a prompt. LLMs are, literally, Language Models, built by and for people for whom text is the natural tool of thought. The terminal, for a software engineer, is not a regression. It is the preferred environment. The prompt interface is not an accident, but the logical consequence of who was in the room when the decisions were made.
In this sense, Leigh frames the problem as a clear dichotomy in which engineers think verbally and discursively while designers think in images. As a professional who works in both fields, I disagree with the polarity. What I observe is that the visual and verbal resources of each field do not always translate well between each other.
What I see in practice is that engineers also draw constantly to make themselves understood. They create diagrams, representations, abstractions. They draw on whiteboards, sketch in notebooks, send annotated screenshots in chat and in Jira tickets. The predominant language of engineering is not just text. Explaining a complex architecture without drawing a diagram is impossible. Writing requirements using only images is too.
Similarly, every interface or experience project I have seen needed written text at some point, whether to document iterations, instruct the use of the design system or record comments from those evaluating the work. When I worked directly in design, I was never able to communicate with images alone.
In short, engineers also communicate visually, and designers also write. The difference between an engineer and a designer is not that one thinks in words and the other in images. It is that they were taught to value different languages. As a consequence, translation between visual and textual has always been problematic. AI did not create this problem. It amplified it by requiring everyone to describe perfectly what they need to a system that can fail 60% of the time in complex contexts.
Leigh ends his article by saying the future of AI interaction is not typing what you want, but showing what you mean. It seems to me this has always been the case and continues to be. I always include an image or a screenshot to help the AI contextualise what I have asked. With techniques such as OCR and Python integration, showing the AI what you mean is already a reality, not just a promise. After all, the AI prompt is functionally very close to a computer terminal, and the terminal has always known how to handle files, images and data beyond text.
Given this possibility, the temptation is to conclude that the problem lies with the user. It does not, or at least not exclusively. A system that depends on perfect verbal description to function and that errs up to 60% of the time in complex contexts is a system with design flaws. The prompt interface as the primary mode of interaction was a choice, not an inevitability. Visual interfaces, direct manipulation, canvas with embedded AI — all of this already exists in early stages.
But while these interfaces are not yet consolidated, and even after they are, the capacity to articulate complex problems remains an irreplaceable human competence. This leaves us with two important conclusions.
The first is that AIs can and should be better. Systems built by more diverse teams, with designers involved in structural decisions from the start, would very likely produce interfaces and interactions that would help improve AI success rates in business.
The second is that, even with better interfaces and interactions, the person needs to know what they want to say before showing or typing anything. This is not a question of tooling. It is a question of education, of repertoire, of the capacity to operate in more than one language at once.
As long as the structural decisions behind AI systems continue to be made by teams without disciplinary diversity, the result will continue to reflect the language of whoever was in the room. Better interfaces will come. AIs that barely hallucinate will too. But better tools do not eliminate the need to understand the problem, have a good idea of what the solution should be, and apply all of this successfully. The responsibility belongs to both sides.
Without this combination, AI does not translate between languages. It only amplifies pre-existing problems, almost always with a little extra seasoning of AI hallucination.
Paivio, A. (1971). Imagery and Verbal Processes. Oxford University Press. taylorfrancis.com
Potter, M.C., Wyble, B., Hagmann, C.E. & McCourt, E.S. (2014). "Detecting meaning in RSVP at 13 ms per picture." Attention, Perception, and Psychophysics, 76(2), 270–279. springer.com
Zeng, J. et al. (2025). "Neural Distinction between Visual Word and Object Recognition." Journal of Neuroscience, 45(28). jneurosci.org
Méndez Fernández, D., Wagner, S. et al. (2017). "Naming the Pain in Requirements Engineering." Empirical Software Engineering, 22(5), 2298–2338. springer.com
Chen, Z. et al. (2025). "A Taxonomy of Prompt Defects in LLM Systems." arxiv.org
Lucchetti, A. et al. (2025). "More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation." arxiv.org
Leigh, J. (2026). "The prompt is not an interface." UX Collective / The Conditions. medium.com
AuthenHallu (2025). "Detecting Hallucinations in Authentic LLM–Human Interactions." arxiv.org