Try describing the taste of a sandwich to the person at the counter. Not pointing to the photo on the menu, not choosing from a list, just using your own words to place your order: “I want a soft bun, like brioche, with meat that has that smoky flavour but isn’t dry, a sauce that’s sort of tangy but not exactly mustard, and some leaves that add crunch but aren’t lettuce.”
However skilled the person making the sandwich, the result will probably still have errors and you’ll have to make several small adjustments, probably infuriating whoever is standing hungry behind you in line. But what if you had just shown the photo of the sandwich you wanted?
It sounds trivial, but there is a long-running debate in the field of human learning comparing textual and visual language. With the relative regression brought by AI, that debate is back in full force, along with long-standing problems in software development.
Now imagine that instead of a sandwich you need to verbally describe a citizen service system for a Brazilian public agency: “it has to be easy to access, but can’t leak information. It has to follow the law, but can’t be too rigid.” Imagine you and your colleagues having to tell the developers about every exception, parallel flow and regional context such a system requires. Only with words (or at least mostly words). That is how most software systems for external users are still built today in Requirements Engineering.
As the field of Human-Computer Interaction evolved, techniques from ergonomics, psychology and design were integrated into software development. When design brought its tools into system projects, it carried useful prototyping and visual concept validation techniques, but these were never fully incorporated into Requirements Engineering, even though they brought improvements.
In The prompt is not an interface, Joshua Leigh reconstructs forty years of interaction design evolution to argue that the prompt interface is a regression for visual and spatial tasks. The argument draws on Paivio’s (1971) dual coding theory, according to which the human brain processes visual and verbal information through separate cognitive systems. The verbal system operates sequentially, constrained by the linearity of language. The visual system operates in parallel, specialised in spatial relationships, processing more information simultaneously. Images are also remembered more easily than words.
Recent research confirms and refines this distinction. Potter et al. (2014, MIT) showed that the brain can process an image in 13 milliseconds. Zeng et al. (2025) used fMRI with Chinese pictographs — visually identical stimuli interpreted as either words or objects — and confirmed that the two types of recognition activate distinct neural mechanisms.
The regression brought by AI prompts, therefore, lies in going back to depending predominantly on text to specify what we want from a system, as was done in the DOS era. In practice, this meant relegating visual communication to a secondary role in the context of computational systems. As a result, problems that have existed since the early days of Requirements Engineering are now being amplified at an unprecedented scale.
The effect of near-exclusive reliance on verbal language as a tool for requirements elicitation and specification is already a classic in the software field. NaPiRE, a global survey conducted with 228 companies across 10 countries, found that incomplete requirements are the number one problem in requirements engineering, reported by 48% of respondents, followed by communication failures (41%) and requirements that are too abstract (33%) (Méndez Fernández et al., 2017). Poorly written and ambiguous requirements are consistently among the top five causes of failure in software projects.
With AI, the prompt has assumed the role of the requirement. Chen et al. (2025) propose a taxonomy of prompt defects and argue that prompts function as source code in LLM-based systems, except they are written in ambiguous natural language and executed on a probabilistic engine. The resulting defects range from trivial formatting problems to security failures and misinformation. Recent research also shows that novice programmers systematically omit crucial details in their prompts, and that prompt specificity directly affects the quality of generated code (Lucchetti et al., 2025).
In other words, one of the main difficulties in systems development is communication. Technology professionals operate with different languages and jargons among themselves, and even more different from the languages of those who commissioned the system and those who will use it. This difficulty is at the heart of how today’s popular AIs were developed. Designers rarely participated in the structural decisions behind these systems. The pattern repeats: they are called in to fix what has already been built, the same role they had when the product was just an interface. As a rule, they were not in the room when the structural decisions were made.
The world’s largest AI companies, including OpenAI, Google DeepMind and Anthropic, were founded and are led by machine learning researchers and software engineers. None of them has a designer among its founders. Design roles exist at these companies, but as a product support function, not a system function. The structural decisions — from model architecture to the choice of the prompt interface as the primary mode of interaction — were made by people for whom text is the natural language of work. Designers come in afterwards, to make what has already been built look like a coherent product for the end user.
It is therefore no surprise that the primary form of interaction with AI today is a prompt. LLMs are, literally, Language Models, built by and for people for whom text is the natural tool of thought. The terminal, for a software engineer, is not a regression. It is the preferred environment. The prompt interface is not an accident, but the logical consequence of who was in the room when the decisions were made.
In this sense, Leigh frames the problem as a clear dichotomy in which engineers think verbally and discursively while designers think in images. As a professional who works in both fields, I disagree with the polarity. What I observe is that the visual and verbal resources of each field do not always translate well between each other.
What I see in practice is that engineers also draw constantly to make themselves understood. They create diagrams, representations, abstractions. They draw on whiteboards, sketch in notebooks, send annotated screenshots in chat and in Jira tickets. The predominant language of engineering is not just text. Explaining a complex architecture without drawing a diagram is impossible. Writing requirements using only images is too.
Similarly, every interface or experience project I have seen needed written text at some point, whether to document iterations, instruct the use of the design system or record comments from those evaluating the work. When I worked directly in design, I was never able to communicate with images alone.
In short, engineers also communicate visually, and designers also write. The difference between an engineer and a designer is not that one thinks in words and the other in images. It is that they were taught to value different languages. As a consequence, translation between visual and textual has always been problematic. AI did not create this problem. It amplified it by requiring everyone to describe perfectly what they need to a system that can fail 60% of the time in complex contexts.
Leigh ends his article by saying the future of AI interaction is not typing what you want, but showing what you mean. It seems to me this has always been the case and continues to be. I always include an image or a screenshot to help the AI contextualise what I have asked. With techniques such as OCR and Python integration, showing the AI what you mean is already a reality, not just a promise. After all, the AI prompt is functionally very close to a computer terminal, and the terminal has always known how to handle files, images and data beyond text.
Given this possibility, the temptation is to conclude that the problem lies with the user. It does not, or at least not exclusively. A system that depends on perfect verbal description to function and that errs up to 60% of the time in complex contexts is a system with design flaws. The prompt interface as the primary mode of interaction was a choice, not an inevitability. Visual interfaces, direct manipulation, canvas with embedded AI — all of this already exists in early stages.
But while these interfaces are not yet consolidated, and even after they are, the capacity to articulate complex problems remains an irreplaceable human competence. This leaves us with two important conclusions.
The first is that AIs can and should be better. Systems built by more diverse teams, with designers involved in structural decisions from the start, would very likely produce interfaces and interactions that would help improve AI success rates in business.
The second is that, even with better interfaces and interactions, the person needs to know what they want to say before showing or typing anything. This is not a question of tooling. It is a question of education, of repertoire, of the capacity to operate in more than one language at once.
As long as the structural decisions behind AI systems continue to be made by teams without disciplinary diversity, the result will continue to reflect the language of whoever was in the room. Better interfaces will come. AIs that barely hallucinate will too. But better tools do not eliminate the need to understand the problem, have a good idea of what the solution should be, and apply all of this successfully. The responsibility belongs to both sides.
Without this combination, AI does not translate between languages. It only amplifies pre-existing problems, almost always with a little extra seasoning of AI hallucination.
References
Paivio, A. (1971). Imagery and Verbal Processes. Oxford University Press. taylorfrancis.com
Potter, M.C., Wyble, B., Hagmann, C.E. & McCourt, E.S. (2014). “Detecting meaning in RSVP at 13 ms per picture.” Attention, Perception, and Psychophysics, 76(2), 270–279. springer.com
Zeng, J. et al. (2025). “Neural Distinction between Visual Word and Object Recognition.” Journal of Neuroscience, 45(28). jneurosci.org
Méndez Fernández, D., Wagner, S. et al. (2017). “Naming the Pain in Requirements Engineering.” Empirical Software Engineering, 22(5), 2298–2338. springer.com
Chen, Z. et al. (2025). “A Taxonomy of Prompt Defects in LLM Systems.” arxiv.org
Lucchetti, A. et al. (2025). “More Than a Score: Probing the Impact of Prompt Specificity on LLM Code Generation.” arxiv.org
Leigh, J. (2026). “The prompt is not an interface.” UX Collective / The Conditions. medium.com
AuthenHallu (2025). “Detecting Hallucinations in Authentic LLM–Human Interactions.” arxiv.org