
Morpius
- Client
- Morpius (self-initiated)
- Role
- Founder & Design Engineer
- Timeline
- 2026 — present
- Deliverables
- Product strategy, Practitioner research, Workbench interface, Agentic build pipeline
- Research conversations
- 25
- Agents on the engineering team
- 17
- Architecture decisions recorded
- 86
Morpius is the product I am building now — morpius.com, self-initiated and in active development. It is here for the reason the client studies cannot cover: nobody set the constraints but me.
Why agents produce generic work
Ask a capable model to design something and it returns the median of everything it has seen. For a design team that is the wrong answer, because a team's value is precisely the part that is not the median — the conventions it settled on, the patterns it threw out, and the reasons behind both.
Designers compensate by pasting the same context into every session. It works, and it decays. The gap is not model capability; it is that a team's judgment has no structured home.
What twenty-five conversations produced
Nineteen practitioners in the role I am building for, six advisors — a split I audited partway through rather than at the end, because it changes what the research is allowed to prove. Four findings changed the product.
25 people
- 19Practitioners in the role
- 6Advisors
Who I actually spoke to data table
| Group | People |
|---|---|
| Practitioners in the role | 19 |
| Advisors | 6 |
The competition
Everyone serious had already built their own version
Five practitioners, five hand-rolled proto-versions. The thing to beat is not another product but the user's own duct tape — inside a single session.
The ceiling
Nobody asked for a model that designs at a senior level
One put the plateau at 60–70% of the quality he would accept. The ask was not a better model, but that the 70% land on their standard rather than everyone's average.
The demand signal
Authorship, not automation
I expected enthusiasm for generation. What landed was the editable view — seeing what the system inferred and correcting it. Every practitioner got there by a different route.
The gap
The journey dead-ends one step before the paid deliverable
What a client pays for is the workflow model. Tooling covers import, extract and shape — then drops the designer into Visio to draw the thing they are paid for.
So the human gate is not a compliance cost to minimise. It is the feature — the reading that turned this from a generator with review attached into an editor with generation attached.
What it does, end to end
Pasted context fails three ways: it has no home, it decays, and it has to be re-explained. Fix only the first and you have built a filing cabinet — so Morpius is a loop, running on an ontology: a written map of what a team cares about and how it relates. A prompt is a paragraph you retype; a structure can be corrected in one place and handed to any agent.
Get it in
The machine reads the transcripts, docs and files and proposes a structure. Nothing gets retyped, and nothing enters because the machine said so.
Put it to work
An agent asks per task and gets an answer scoped to that task — never the substrate itself. This replaces the paragraph you paste every session.
Keep it true
Teams change their minds, and a substrate that does not know a decision was superseded starts quietly lying to every agent that reads it.
- 01Sourcestranscripts, docs, files
- 02Extractthe model reads and structures
- 03Proposelands as suggestions
- 04Reviewyou accept, edit or reject
- 05Graphthe substrate you shaped
- 06Compileasked for, per task
You assert → it commits. AI infers → it proposes.
A verdict carries its reason, because a bare rejection teaches the system nothing.
Put it to work
The agent asks, per task. Handing it the whole substrate and letting it work out what matters is just a bigger paste — so Morpius never does. The selection is the product.
- 01Situatewhat kind of design problem is this, really
- 02Shaperetrieve across lanes, fuse, grade, recover
- 03Packettyped, with receipts and known gaps
What comes back is typed: binding decisions with receipts, advisory guidance marked advisory, and — the field I would point at — what the ontology does not know. A model with no grounds to answer will answer anyway. A substrate that reports its own gaps turns that into a decision somebody can make.
Keep it true
The difference between a substrate and a folder. Everything carries where it came from, so a weak inference has a paper trail. A new decision supersedes an older one on the record; aged knowledge gets surfaced for re-ratification; decisions have owners. None of that demos well, and all of it is why the thing is still true in a year — the failure mode of every context tool in the research was not that it never got populated, but that it got populated once.
It is also where this stops being a designer tool: a designer asks what we decided, a PM asks why, an engineer asks what to build. One governed source beats three maintained documents, because the documents disagree by Thursday.
The Workbench — what I designed
If the value is authorship, the product lives or dies on one screen. A real team's judgment is not twelve nodes but two hundred and up, so the hard part is manipulation, not visualisation — bulk selection, retagging at scale, rewiring relationships, and position that means what the designer chose.
Four surfaces. The canvas is spatial work; the table is the same substrate row by row, for the bulk operations a canvas is bad at; the review queue is where inferred proposals wait for a verdict; the graph view is the connective layer alone — relationships rather than things.
How it's built — the agent team
I run the discovery, set the strategy and design the interface; seventeen agents
research, write, test and ship. It holds because the work is cut into phases with
a written contract between each — /deep-research returns one cited report,
/write-prd and /to-issues turn a decision into vertical slices, /implement
builds one slice against a test gate. Each phase hands the next an artifact
rather than a conversation, so an agent joining at any point reads the output of
the one before it, not a chat log.
17 agents6 haiku6 sonnet5 inherit
Agents have no memory between sessions, so anything I decide and do not record I re-litigate the following week. The decision record is not documentation written after the fact — it is the shared memory every agent reads before it touches anything. Hence eighty-six of them.
Where this goes
The industry is racing to make agents better at design. I think that is the wrong race. A frontier model is already a stronger generalist designer than the median practitioner and will keep improving without my help. What no amount of model scale fixes is knowing what your team decided, and why.
So the thing worth building is the layer underneath — the one that gets more valuable every time someone corrects it. Generators are a commodity race with well-funded runners; the substrate they draw on is not.
Design theory
The universal layer. Shared, stable, the same for everyone.
System rules
One design system's decisions: its components, tokens and patterns.
Your judgment
The layer nobody else has: what you choose when several options are defensible, and why.
Each layer overrides the one beneath it, and the top one is the only part a bigger model cannot reproduce.
One call is worth the record: I had ranked distribution above the owned interface — Morpius as the thing other tools quietly consume — and reversed it within the week. Accumulation happens through correction, correction is an interaction, and a product nobody opens for its own sake has nothing to compound.
Everything after that follows from one property: judgment, once structured, can move. To your own agents first, then across a team, then between teams — where structured expertise can be licensed, and a designer's accumulated judgment stops evaporating at the end of a project. That last step is a direction, not a shipped feature, and it carries real questions about what such a market owes the people whose work is in it. The sharpest answer I have came out of a research call: it is his idea, not mine.
Design is shifting from screens and hours to systems, structure and judgment. The teams that give their judgment a home first will compound it while everyone else retypes theirs every session.