AI-driven UX drift quietly flattens product design, and traditional human review often fails to catch it. Learn how design and platform engineering teams can document AI intent, redesign review surfaces, and maintain product quality.
Jarmo Parkkinen
Jarmo on Lead UX Designer Eficodessa. Hän on työskennellyt käytettävyyden, käyttäjäkokemuksen, vuorovaikutussuunnittelun ja palvelumuotoilun parissa yli 25 vuoden ajan. Hän rakastaa monimutkaisia järjestelmiä ja tehokkaita käyttöliittymiä. Vapaa-aikanaan hän jakaa kuvia kissastaan Facebookissa.
Are your AI tools quietly flattening your product's UX or service quality? When adopting AI in your product workflows, the biggest risk isn't that AI breaks your design. But it may slowly average it out.
Learn about AI adoption in software development
Read moreI gave four frontier models an example of Google Labs' DESIGN.md, and asked them to create a "next" arrow at 44×44 and 88×88px. All four complied with every token the spec defined: color, stroke weight, caps, joins, canvas, no fill. They also agreed with each other on two things the spec never mentioned — the angle of the arrowhead, and the treatment of the line ends.
See DESIGN.md, the design-rule format open-sourced by Google.
This is not a benchmark, but it demonstrates a problem: A specification can be complete enough to pass compliance tests and still leave the model to fill the gaps. What models fill them with is the average of everything they were trained on.
Why AI-driven UX drift is hard to catch
The icons experiment was part of my own project, where I tested creating and dealing with UX material of my own. I set myself with strict UX and service design goals that I was not allowed to design away — the way you could not design away a client's core requirements. I gave it deliberately unusual interaction patterns on purpose: three overlapping navigation models, controls that appear only after specific events, back-and-forward behavior that changes depending on context - things that I assume are not well presented in the training data.
I assumed the odd patterns would be hard to build with AI. They were not, but keeping them intact was almost impossible.
Over a few weeks, the unusual patterns quietly became conventional ones. The wording I had chosen for a reason got "improved" during an unrelated bug fix. A colleague hit the same thing from the other direction: re-applying a design system that had worked before started producing pixel drift, odd positioning, and subtle shifts in tone of voice.
None of the above was visible in tests, as nothing was broken. Everything moved a few degrees toward the average.
The mechanism is not mysterious. AI is trained on normal things that exist in abundance in the training material. In technical work, that is mostly a gift - but in product design, it is close to the opposite. Most of what makes a product worth funding is the part that is not normal yet, definitely not available at a large scale in training data.
Why human review fails to catch AI drift
Here is the part I would rather not write. This did not happen behind my back. The tool showed me a diff for every change and asked for permission. I read them. I approved them all.
That is worth sitting with, because the standard answer to everything above is "no problem, a human reviews the output and takes responsibility". I was the human, I knew my goals, I was reviewing - I should maybe fire myself from my test project? Or are there alternative views?
A diff is an excellent instrument, and it deserves better than to be blamed here. Diff works best when the scope of a change is narrow or well defined: the effects are mostly local, and the non-local ones are made explicit by a coding convention or a failing test. Also, the work loop around of code and software development work is very tightly knit. These factors explain why developers absorbed AI assistance as fast as they did.
A change to a product assumption is not local in that way. Change one sentence at the top of a positioning document, and the meaning of everything below it may change while the text stays the same. Change a rule in a service workflow, and the consequence surfaces three steps later - in somebody else's journey. We are lacking the tools and conventions to catch this type of change when done by AI.
So I was maybe not reading carelessly. I was looking for the right object but in the wrong form. I read the words and missed the meaning. Volume finishes the job.
We have run this experiment several times outside AI, and we know how it ends. Early virus scanners announced everything they did, because users "needed to know security was happening". What the users learned instead was the reflex to dismiss the dialogs, including the odd one that would have been important. Web advertising taught the same lesson in reverse: put something in the shape of a banner and people stop seeing it. The human-factors literature has said this about automation for years — Parasuraman and Manzey found that automation bias and complacency turn on the person, the situation, and the system, and that instructions and training alone do not remove it. NIST's AI risk guidance, therefore, asks teams to examine how outputs are *presented* and to bring in human-factors expertise.
Sources:
Parasuraman & Manzey: Complacency and Bias in Human Use of Automation: An Attentional Integration
None of that means people are careless or that the review failed and they need to be re-educated. Instead, the task was designed for different content and profession.
How design agencies work with AI
Mine is not a lone finding. The big design firms have largely stopped arguing if generation is useful in product design and information worker tasks. They have moved on to exactly this new territory. Designit built Repsol, a shared AI-experience framework of interaction principles into reusable human-AI patterns so that "teams no longer start from scratch". frog's agentic-AI playbook prescribes guardrails, verification points, escalation rules, agents holding "the minimum authority required to do the job", and evaluation that is continuous rather than one-off. And IDEO has published The case against AI-generated users, warning about the dangers of shallow, generic and emotionless "data" entering to customer understanding.
So, encode your context, govern the loop, stop counting prompts. I agree with them, but I see a missing perspective: the daily work, and changes needed in workflow to keep the wanted direction.
References: Designit built Repsol and frog's agentic-AI playbook
Sources: IDEO: The case against AI-generated users
For designers: Document for AI what must stay true
The useful starting set is small: the design and product rules, including where each one does *not* apply, should always be available for the agent working on the task. Assumptions behind a journey should be linked with the screens and behaviors that depend on them, examples of acceptable and unacceptable outcomes, and all that good stuff that used to be left between swim lanes.
This is not a "single-prompt" task. An agent follows stale intent with exactly as much confidence as fresh intent. One big file is the wrong shape as brand, product vision, service journeys, content, and interaction patterns have different owners and change at different rates. Shipping all of it on every LLM-call costs tokens and buries the instruction that mattered.
I am experimenting with an index instead — something that points the agent at "widgets → buttons → switch" when that is what is needed, and not before. The structure matters less than the principle: the smallest relevant intent, at the moment it applies. Caveats? This must be maintained, so you need to get involved with Confluence and Jira work on planning, and get into GitHub to make the relevant changes during the development phase.
For managers: Design the human's work as carefully as the prompt
Writing product intent down reduces drift, and monitoring the intent effect allows goals to stay intact. But drifts need a review surface built for the profession doing the reviewing.
The designer needs to be part of the product team. They need to see the visual and service model changes in their own context. Show a service designer the assumption that moved and the journeys hanging off it. Show a product owner the goal, the evidence, and the trade-off.
Let the environment enforce the routine technical boundaries on its own — asking a designer to approve an unfamiliar shell command does not make that designer a security reviewer. It is noise that exhausts and diverts. Asking a product manager to approve a text diff does not establish that they approved the change to the product.
This is a joint job for design, development, and platform teams, and it needs one thing from whoever owns the pipeline: people in roles other than developer need access to the real AI-enabled workflow rather than its final output, and the authority to tune, stop, or reject it.
Pick one workflow where people are repeatedly approving, reconstructing, or repairing AI output a lot. Ask each reviewer what they actually need to understand in order to say yes. Build the review around that answer, and measure whether it catches anything.
AI will keep getting better at producing the work. Keeping each role's tooling up to pace isn't a designer problem; it's a pipeline problem. To fix AI drift, [Platform Engineering teams](https://www.eficode.com/software-tooling/platform-engineering) need to build review surfaces that allow product owners to see AI changes in context, not just in text diffs.
Read the ultimate guide to platform engineering
Read the guide- Platform Engineering
- AI
- Design and UX
Subscribe to our newsletter
Related blogs