A Morpius knowledge graph filling the frame — clusters of green, teal and orange nodes joined by directed edges, dense at the right and thinning to scattered points at the left

Morpius

Client
Morpius (self-initiated)
Role
Founder & Design Engineer
Timeline
2026 — present
Deliverables
Product strategy, Practitioner research, Workbench interface, Agentic build pipeline
Research conversations
25
Agents on the engineering team
17
Architecture decisions recorded
86

Morpius is the product I am building now — morpius.com, self-initiated and in active development. It is here for the reason the client studies cannot cover: nobody set the constraints but me.

Why agents produce generic work

Ask a capable model to design something and it returns the median of everything it has seen. For a design team that is the wrong answer, because a team's value is precisely the part that is not the median — the conventions it settled on, the patterns it threw out, and the reasons behind both.

Designers compensate by pasting the same context into every session. It works, and it decays. The gap is not model capability; it is that a team's judgment has no structured home.

What twenty-five conversations produced

Nineteen practitioners in the role I am building for, six advisors — a split I audited partway through rather than at the end, because it changes what the research is allowed to prove. Four findings changed the product.

Who I actually spoke to

25 people

  • 19Practitioners in the role
  • 6Advisors
A small sample, treated as one: enough to establish the problem is real and badly served, not enough to size a market.
Who I actually spoke to data table
GroupPeople
Practitioners in the role19
Advisors6
  • The competition

    Everyone serious had already built their own version

    Five practitioners, five hand-rolled proto-versions. The thing to beat is not another product but the user's own duct tape — inside a single session.

  • The ceiling

    Nobody asked for a model that designs at a senior level

    One put the plateau at 60–70% of the quality he would accept. The ask was not a better model, but that the 70% land on their standard rather than everyone's average.

  • The demand signal

    Authorship, not automation

    I expected enthusiasm for generation. What landed was the editable view — seeing what the system inferred and correcting it. Every practitioner got there by a different route.

  • The gap

    The journey dead-ends one step before the paid deliverable

    What a client pays for is the workflow model. Tooling covers import, extract and shape — then drops the designer into Visio to draw the thing they are paid for.

So the human gate is not a compliance cost to minimise. It is the feature — the reading that turned this from a generator with review attached into an editor with generation attached.

What it does, end to end

Pasted context fails three ways: it has no home, it decays, and it has to be re-explained. Fix only the first and you have built a filing cabinet — so Morpius is a loop, running on an ontology: a written map of what a team cares about and how it relates. A prompt is a paragraph you retype; a structure can be corrected in one place and handed to any agent.

  1. Get it in

    The machine reads the transcripts, docs and files and proposes a structure. Nothing gets retyped, and nothing enters because the machine said so.

  2. Put it to work

    An agent asks per task and gets an answer scoped to that task — never the substrate itself. This replaces the paragraph you paste every session.

  3. Keep it true

    Teams change their minds, and a substrate that does not know a decision was superseded starts quietly lying to every agent that reads it.

What a designer actually does
  1. 01Sourcestranscripts, docs, files
  2. 02Extractthe model reads and structures
  3. 03Proposelands as suggestions
  4. 04Reviewyou accept, edit or reject
  5. 05Graphthe substrate you shaped
  6. 06Compileasked for, per task
The substrate only ever gains what a human ratified — the one rule the whole product is arranged around.

You assert → it commits. AI infers → it proposes.

A verdict carries its reason, because a bare rejection teaches the system nothing.

Put it to work

The agent asks, per task. Handing it the whole substrate and letting it work out what matters is just a bigger paste — so Morpius never does. The selection is the product.

What a compile does
  1. 01Situatewhat kind of design problem is this, really
  2. 02Shaperetrieve across lanes, fuse, grade, recover
  3. 03Packettyped, with receipts and known gaps
Situating first stops the compile retrieving against the words in a request instead of the problem underneath it.

What comes back is typed: binding decisions with receipts, advisory guidance marked advisory, and — the field I would point at — what the ontology does not know. A model with no grounds to answer will answer anyway. A substrate that reports its own gaps turns that into a decision somebody can make.

Keep it true

The difference between a substrate and a folder. Everything carries where it came from, so a weak inference has a paper trail. A new decision supersedes an older one on the record; aged knowledge gets surfaced for re-ratification; decisions have owners. None of that demos well, and all of it is why the thing is still true in a year — the failure mode of every context tool in the research was not that it never got populated, but that it got populated once.

It is also where this stops being a designer tool: a designer asks what we decided, a PM asks why, an engineer asks what to build. One governed source beats three maintained documents, because the documents disagree by Thursday.

The Workbench — what I designed

If the value is authorship, the product lives or dies on one screen. A real team's judgment is not twelve nodes but two hundred and up, so the hard part is manipulation, not visualisation — bulk selection, retagging at scale, rewiring relationships, and position that means what the designer chose.

A real workbench, not a sample: 689 nodes extracted from my own buyer-interview research.
The same substrate as a table — Morpius's own design system in 35 concepts. Every row is a judgment with its type and reason attached.

Four surfaces. The canvas is spatial work; the table is the same substrate row by row, for the bulk operations a canvas is bad at; the review queue is where inferred proposals wait for a verdict; the graph view is the connective layer alone — relationships rather than things.

450 concepts, 337 relationships. Selecting a node shows how it is joined — recommended_for, constrained_by, implements. That is what pasted context cannot carry: not the things, but the argument between them.

How it's built — the agent team

I run the discovery, set the strategy and design the interface; seventeen agents research, write, test and ship. It holds because the work is cut into phases with a written contract between each — /deep-research returns one cited report, /write-prd and /to-issues turn a decision into vertical slices, /implement builds one slice against a test gate. Each phase hands the next an artifact rather than a conversation, so an agent joining at any point reads the output of the one before it, not a chat log.

17 agents6 haiku6 sonnet5 inherit

Agents have no memory between sessions, so anything I decide and do not record I re-litigate the following week. The decision record is not documentation written after the fact — it is the shared memory every agent reads before it touches anything. Hence eighty-six of them.

Where this goes

The industry is racing to make agents better at design. I think that is the wrong race. A frontier model is already a stronger generalist designer than the median practitioner and will keep improving without my help. What no amount of model scale fixes is knowing what your team decided, and why.

So the thing worth building is the layer underneath — the one that gets more valuable every time someone corrects it. Generators are a commodity race with well-funded runners; the substrate they draw on is not.

  1. Design theory

    The universal layer. Shared, stable, the same for everyone.

  2. System rules

    One design system's decisions: its components, tokens and patterns.

  3. Your judgment

    The layer nobody else has: what you choose when several options are defensible, and why.

Each layer overrides the one beneath it, and the top one is the only part a bigger model cannot reproduce.

One call is worth the record: I had ranked distribution above the owned interface — Morpius as the thing other tools quietly consume — and reversed it within the week. Accumulation happens through correction, correction is an interaction, and a product nobody opens for its own sake has nothing to compound.

Everything after that follows from one property: judgment, once structured, can move. To your own agents first, then across a team, then between teams — where structured expertise can be licensed, and a designer's accumulated judgment stops evaporating at the end of a project. That last step is a direction, not a shipped feature, and it carries real questions about what such a market owes the people whose work is in it. The sharpest answer I have came out of a research call: it is his idea, not mine.

Design is shifting from screens and hours to systems, structure and judgment. The teams that give their judgment a home first will compound it while everyone else retypes theirs every session.