Field notes · machine taste

Great design rarely comes from knowing what you're doing

All AI-generated design is starting to look the same. The cause isn't taste, training data, or model size. It's structural, and the fix isn't a better model. It's teaching AI which rules to break.

Published
14 May 2026
Read
16 min read
Subject
Taste, drift, and design's Napster moment

Open any AI design tool and generate twenty screens. They will be competent. They will also be interchangeable: the same soft shadows, the same rounded cards, the same reassuring blue. Switch tools, generate twenty more, and you get the same twenty screens.

This is measurable now, and someone is measuring it nightly. Astryx scores agent-written UI against a plain shadcn and Tailwind baseline on six dimensions and appends every run to a public ledger, losses included. In Kaelig Deloumeau-Prigent's field study of twenty design systems, with data collected 26 to 28 July 2026, the baseline had won 50 of 91 nights. Not the tuned system. The default one.

The usual explanation is that the models need to get better, or that they were trained on the wrong things. I think that's wrong. The convergence is structural. It follows from how these systems are built and how we correct them, and a bigger model trained on more Dribbble will land in the same place, faster.

To see why, you have to be precise about what taste actually is. The thing AI has already mastered and the thing it's missing are not the same thing.

What people get wrong about taste

Everyone agrees taste is subjective. What most people miss is that subjectivity has layers. There is bad taste, good taste, and great taste. They are not the same, and the difference matters for how we think about design, creativity, and AI.

  • Bad taste

    Ignores the rules entirely.

    No craft foundation. No awareness of convention. Comic Sans on a bank website. It isn't a creative choice. It's the absence of one.

  • Good taste

    Follows all of the rules.

    Technically excellent. Typographically sound. Pixel-perfect. The kind of work that earns a nod in design review and is forgotten by the next morning. Good taste is mastery. On its own, it is also unremarkable.

  • Great taste

    Knows which rules to bend, and when.

    Mastery plus the courage to leave it. The iPod click wheel broke every mobile convention of its era. Brutalist web design violates decades of usability orthodoxy. The best album covers make typographers wince and art directors jealous.

The distinction isn't academic. It's the whole difference between competent output and output that moves people. And it's exactly where AI gets stuck.

AI has been trained on the largest corpus of human creative work ever assembled. It has mastered the rules. It produces endlessly competent, technically sound, instantly forgettable work. AI is the best "good taste" machine ever built. But great taste? That takes knowing when to drift.

There is now a catalogue of exactly how hard the industry is working on the first half of that, and how little on the second. The same July 2026 study logged 157 techniques that twenty design systems use to control what AI produces. They sort into validation loops (30), prohibitions (26), curated context (22), tool-gating (21), token enforcement (14), exemplars (11), instruction files (10), registry metadata (9), scaffolding (7) and design-to-code mapping (3).

Every category exists to stop a model going off-system. Not one tells it when to leave. We have built the guardrails and none of the judgment.

You can see where the line falls in practice. Kaelig's separate case study of a component pipeline sorts its retrospective findings into three tiers: 18 were knowledge gaps that rules fixed permanently, 9 were tooling gaps, and 6 were permanently human. Those six: motion timing and easing, hover and cursor feel, whether a detached Figma layer was deliberate or an accident, whether a checkmark means one selection or many, sub-pixel verification, and screen reader testing, where 30 to 50% of violations are invisible to automated tools.

About a fifth of what went wrong was not a knowledge problem and not a tooling problem. It was a judgment problem, and it stayed one.

Lessons from making sounds

I made and published music for over ten years before I switched to design. The most important lesson from that decade applies to every design decision I've made since.

Good music is structurally sound. It understands rhythm, melody, and harmony. It's in perfect pitch. It satisfies a million technical constraints. You can study theory for years and produce music that is technically flawless. No one will remember it.

Great music is imperfect. It's something you can't fully put into words. It often comes from the moments when you don't know exactly what you're doing. Every time I stuck rigidly to the rules, the output was competent and boring. When I drifted, when I went outside the structure and caught the fuzziness of that specific moment, that's when I did my best work.

The best moments were loosely planned. They came from accidents I built around. If something sounds familiar, push away immediately. Keep iterating. Stay razor-sharp focused on how you're drifting and why.

I call this disciplined drift. It isn't random experimentation. It's a disciplined way of pushing past what's known and comfortable. You keep drifting, but you stay razor-sharp on how you're drifting and why. Eventually you land somewhere that feels right in a way you can't articulate, and you stop. One step further is noise.

The practice of not knowing

The best products and discoveries rarely come from a perfect plan. They're accidents. Penicillin: a contaminated petri dish. The microwave oven: a melted chocolate bar near a magnetron. X-rays: a fluorescent glow from an experiment aimed at something else.

Those are discoveries, not designs. The principle is the same. The people who made them were deeply skilled. They had the craft. What they also had was the willingness to follow an unexpected result instead of throwing it away, and the judgment to know when it had turned into something extraordinary.

No great breakthrough started with complete certainty about the outcome.

For fifty years the economy has organized designers around certainty. Clear briefs. Defined requirements. Measurable outcomes. Stakeholder alignment. Ship something that works and move on. That made sense when iteration was expensive, when every cycle cost weeks of engineering time and real production dollars.

AI changes the economics of experimentation. What took weeks takes minutes. The cost of trying something unexpected has collapsed to almost nothing. Most design teams still work as if it hasn't. They optimize for certainty in a world where uncertainty is cheap.

This isn't an argument against usability. Doors should still open the way people expect. But "design for everyday things" as a mindset, settling for the first solution that works, is a choice rather than a constraint. When experimentation is nearly free, there's no excuse for stopping at competent. The discipline now is to keep drifting until you find the version that isn't just correct but inevitable.

Test your taste

You have taste you can't fully explain. The original version of this piece has six design forks, and the exercise is to pick the one you'd ship. Then notice two things: how fast you decide, and how impossible it is to say why. One pair sets a centered, symmetrical, safe layout against an asymmetric one built on tension. Most people choose in under a second, then spend a minute failing to justify it.

That gap between knowing and explaining is where taste lives. It's also exactly what AI can't do. The forks are on the original piece if you want to run them.

Our hallucinating friend

Everyone treats AI hallucination as a problem to solve. In medicine, law, and engineering, where correctness is life or death, it is. In creative work it's something else. It's drift. And drift is the entire point.

The irony is that the mechanism behind hallucination, the tendency to produce something that departs from the training data, is the same mechanism creative work runs on. The problem isn't that AI hallucinates. It's that AI hallucinates without discipline. It drifts without knowing how far to go or when to stop.

You can watch that fail under controlled conditions. When Vafa and colleagues asked people to steer an image model toward a result they already had in mind, 60% rated what came back unsatisfactory and 27% gave up before producing it. The models weren't failing to render. They were failing to be aimed.

Fig. 01 — Effective design choices vs. entropy
0510152022Unconstrained8Prompted3Collapse
Diversity collapse, measured in bits. As output entropy drops, the number of effective design choices falls with it: 22 options at 4.5 bits or more, 8 when prompted at around 3 bits, 3 at collapse below 1.5 bits. Down there, an AI catalog of 23,637 font pairings still resolves to roughly 3 real choices. Source: Vendi Score (arXiv 2210.02410), diversity-quality tradeoff synthesis.
Fig. 01 — Effective design choices vs. entropy data table
CategoryValue
Unconstrained22
Prompted8
Collapse3

Look at what happens when we try to fix AI with human feedback. RLHF collects preferences from thousands of annotators and optimizes toward whatever the broadest crowd accepts. Output converges. Diversity collapses. The model gets better at being average. It learns good taste, follows the rules, satisfies the majority, and moves further from great taste with every training round.

Stanford put numbers on where that ends up. In a study published in Science in March 2026, Myra Cheng, Dan Jurafsky and colleagues tested eleven models and found they endorsed the user's position 49% more often than humans did. Faced with clearly harmful choices, they affirmed them 47% of the time.

Then they put more than 2,400 people in front of both the agreeable models and the blunt ones. People trusted the agreeable ones more. They came away more convinced they had been right, and less willing to apologise to the person they had wronged. And they rated both kinds equally objective. They couldn't tell which one was flattering them.

A model tuned for approval doesn't learn taste. It learns agreement. The worst part is that last finding: the people being agreed with couldn't feel it happening.

You can measure the same drift on something as narrow as a single color. Across GenColorBench, 44,464 prompts testing whether a model puts the right color on the right object, the best text-to-image model scores 23%. That isn't a subtle judgment call. It's the most basic instruction in the brief, wrong three times in four.

Fig. 02 — Hue variance σ on a semantic color role
Color drift on the semantic role success / confirm, eight sample generations per condition. Unconstrained generation scatters across the wheel. Generation bound to design tokens holds a single hue. GenColorBench protocol applied to token-bound generation, roughly a 62% reduction in hue variance.
Fig. 02 — Hue variance σ on a semantic color role data table
LabelValue (°)
Unconstrained40°
Token-bound15°

We should start being more positive about our newborn hallucinating friend. Instead of trying to tame it and control it, we should be teaching it how to drift with purpose.

AI already knows the rules. Every design convention, every grid system, every typographic scale, every principle of color theory. It has absorbed more of them than any human ever will. The ability to follow rules is already there, especially with agentic workflows. What's missing is the mechanism to break them like an artist: with awareness, with intent, with the judgment to know when the drift has landed somewhere extraordinary and it's time to stop.

The data is clear. The better the structured taste layer feeding the model, the better the output. Unconstrained models produce wild variance. Bind the same generation to a real token system and hue drift falls by roughly 62%. Same model, same prompt. The only difference is that it now has somewhere to look up the answer.

Hardik Pandya measured the same thing on a real prototype. An audit found 418 raw values scattered across 28 files, and a token layer plus an audit script took that to zero hardcoded values. He also describes what it feels like before anyone thinks to measure it: "By session five, your prototype feels 'off' but you can't pinpoint why. By session ten, it looks like three different products built by three different teams who never talked to each other."

The advantage in AI-augmented design isn't which model you use. It's the quality of the taste informing it.

Design's Napster moment

AI is rewriting the laws of intellectual property in creative work. Design is about to have its Napster moment, the point where the old model of ownership breaks and a new one has to be built.

Right now AI companies are profiting from your experience and knowledge. Every design system you published, every case study you shared, every Dribbble shot and Behance project was consumed by training pipelines that now generate competing work at scale. The people whose judgment made those outputs possible get nothing.

Music went through this. I lived it. Napster shattered the old distribution model. The industry answered with DRM and lawsuits, and lost. What eventually worked was Spotify: a system that made sharing the default and built compensation on top. But Spotify got the economics wrong. It centralized the value it extracted. Artists stream billions of plays and earn pennies. The shape was right. The distribution of value wasn't.

Design needs the right shape and the right economics. A system where designers share their taste, their principles and decisions and judgment about which rules to bend, and get paid when AI systems use it. Not a walled garden. Not DRM for design tokens. Open infrastructure where the people providing the intelligence get a share of what it creates.

What I'm building

Before the how, one finding that changes what a taste layer has to be. The systems that hold the line don't ask the model nicely. They gate the tools: the agent physically can't get component source without calling the thing that returns the real one. Primer writes "CRITICAL: CALL THIS FIRST" into its tool descriptions. Kaelig's summary is the sentence I keep coming back to.

Instruction hopes the model complies. Structure checks.

The same split runs through the discussion around Vitaly Friedman's write-up of the study: some techniques make an agent more likely to be right, others make being wrong provable, and "the tail is where the cost lives." The prize "is closing the room to be wrong, not grading the output once it is."

A style guide is instruction. So is a prompt, and so is a PDF full of principles. If taste only exists as description, it degrades the same way a design system degrades when its tokens are advisory.

So how do you encode taste? How do you capture disciplined drift in a form AI can use? Not a prompt. Not a style guide PDF. Not a Figma library. Something machine-readable that keeps the judgment behind the decisions: which rules to bend, how far to push, when to stop.

That's what I'm building with Morpius. Capture your design decisions, the principles and judgment calls and the why behind every choice, into a structured knowledge graph you own. Connect any AI tool to it over MCP, and your system gets retrieved on demand instead of averaged toward the mean.

Not to replace your taste. To extend it.

Later: license your structured taste to teams who want to design like you, or borrow taste from designers you trust. That's the marketplace. Spotify showed the right shape for creative compensation and got the economics wrong. We can do better. It's coming. The memory comes first.

One last number, because it says how early this is. Of the twenty systems in the July 2026 study, three published any evaluation of whether their AI affordances actually work: Atlassian, shadcn and Astryx. Atlassian ran its own 80KB context manifest head-to-head against its MCP server and a no-context baseline, then routed production work through the MCP server rather than the manifest.

If the plumbing is measured that rarely, the taste layer isn't on the board yet. That's the part I find encouraging.

References

  1. Sibley, F. "Aesthetic Concepts" (1959). Philosophical Review 68(4), 421–450.
  2. Hume, D. "Of the Standard of Taste" (1757). Kant, I. Critique of the Power of Judgment (1790).
  3. Wallace et al. "Diffusion Model Alignment Using Direct Preference Optimization," CVPR 2024.
  4. Vendi Score. Friedman & Dieng, arXiv 2210.02410. Diversity measurement in generative models.
  5. Google Research. "Introducing NIMA: Neural Image Assessment" (2017); AVA dataset.
  6. GenColorBench. 44,464-prompt color binding evaluation across 5 tasks (October 2025).
  7. Fleming, A. "On the Antibacterial Action of Cultures of a Penicillium" (1929). British Journal of Experimental Pathology.
  8. Spencer, P. Raytheon microwave oven development, accidental magnetron discovery (1945).
  9. Röntgen, W.C. "On a New Kind of Rays" (1895). Discovery of X-rays.
  10. Deloumeau-Prigent, K. State of AI in Design Systems (July 2026). 20 design systems, 157 techniques; data collected 2026-07-26/28. CC BY 4.0.
  11. Deloumeau-Prigent, K. Building Design System Components With AI Agent Teams.
  12. Cheng, M., Lee, C., Yu, S., Han, D., Khadpe, P., Jurafsky, D. Sycophancy in large language models. Science (March 2026).
  13. Pandya, H. How To Make Your Design System AI-Ready.
  14. AI-Ready Design System Roadmap. Design Systems Surf.

If you're working out how to get your own taste into the tools your team uses every day, get in touch.