The default story about AI capability is a leaderboard. The best model is the one at the top, it runs in someone else's data center, and everything else is a compromise. If you want good results, you rent the frontier.
We think that story leaves out the most interesting variable, which is time. A smaller model running on hardware you own, inside a system that remembers, routes and checks its work, gets better at your problems every week, and it doesn't have to beat the on a to do that. It only has to stop re-solving what it has already solved.
That's the case for decentralized AI: not one model everywhere, but many modest ones close to their owners, made capable by everything around them.
Capability isn't only in the weights
When a frontier model outperforms a small one on a task, part of the difference is raw reasoning, and part of it is context: the large model can reconstruct from scratch what the small one doesn't know. The second part is compressible. Give a smaller model the decision record, the documented pattern and the relevant history, and it doesn't need to reconstruct them. It needs to follow them, which is a much easier task.
So the right question isn't “how good is the model?” but “how good is the system the model sits in?” Six pieces of that system do most of the work.
1. Memory that compounds
The most wasteful thing in AI-assisted work is the session that ends and takes its reasoning with it. Capture that output as documents: what was decided, why, and what was tried and failed. Make it retrievable, and the next session starts from there. Over months, a growing share of the work stops being novel reasoning and becomes retrieval plus application, the part smaller models handle well. We wrote about our version of this in Tokens Are Output.
2. Harnessing as a force multiplier
A model in a good is a different tool from the same model in a chat box. The harness decides what the model sees, which tools it can reach, what it must check before it reports success, and what happens when it fails. A well-scoped task with the right files mounted and a clear definition of done is within reach of a model that would flounder on the open-ended version.
3. Route by difficulty
Most work isn't hard. Renaming, formatting, applying a documented pattern, summarizing a log: none of these need the most expensive model available. Send easy tasks to cheap, local models and escalate only what earns it. We already run the same work across several runtimes and a mix of local and hosted models; the next step is making the choice of model a routing decision based on evidence (which model has done well on this kind of task) rather than a habit.
4. Skills and guardrails in the harness
Procedures that took a frontier model to work out (how to review a backend, how to structure a hand-off, how to investigate a failing test) can be written down as skills and handed to any model. A skill is reusable intelligence: the expensive model pays for it once, and cheaper models inherit it. Guardrails belong in the same place, in the harness rather than in the weights, where they can be inspected, versioned and changed.
5. Know when to escalate
The piece that makes routing safe is ambiguity detection. A small model that confidently guesses on an ambiguous task is worse than useless. A small model that notices ambiguity and hands the task up (to a larger model, or to a person) is most of what you need. The goal isn't a small model that can do everything. It's a small model that knows which things it can't do.
6. Inference keeps getting cheaper
Underneath all of this, running models keeps getting cheaper. lets a model that once needed a data-center GPU run on a workstation. Newer methods go after memory directly: Google Research's TurboQuant, for example, compresses the key-value cache (the memory a model uses to hold a long context) to around three and a half bits per value with little measurable loss. Techniques like that turn “needs a server” into “fits on the machine you have.”
Why it matters beyond cost
Access. Most people don't own a large GPU, and memory is one of the scarcest and most expensive parts of any machine. An approach that only works at the frontier leaves most of the world renting intelligence. One that works on modest hardware doesn't.
Ownership. A model and a knowledge base on your own machine keep your documents, your decisions and your half-formed ideas under your own roof.
Discipline. Constraints are good for architecture. When you can't paper over a weak design with a bigger model, you have to fix the design: tighter context, clearer tasks, better documentation, checks that don't depend on the model being clever. Systems built that way tend to be better on the frontier too.
What has to be true
This is a bet, and it could fail in specific ways. If the knowledge base goes stale, retrieval feeds small models confident nonsense. If escalation is badly calibrated, the cheap path quietly produces worse work. And some problems (novel design, reasoning across large systems nobody has documented) will need frontier capability for a long time. Decentralized doesn't mean frontier-free. It means the frontier is a resource you call on when it's needed, not the default for everything.
Where the puck is going
None of these pieces is new on its own. Memory, harnessing, routing, skills, escalation and cheaper each exist today. What changes the picture is combining them, because they compound: memory makes small models capable, routing sends them the work they can do, escalation catches what they can't, and every solved problem goes back into memory for next time.
Run that loop long enough and the gap between a modest model you own and a frontier model you rent narrows for the work you actually do. That's where we think things are going, and it's what we're building toward.