Every useful session with a model ends the same way. The window closes and the reasoning is gone. The model worked out why the plugin system is shaped the way it is, which approach to caching you rejected and why, and what the deploy script assumes. Next week another session works it all out again, at the same cost, and sometimes gets a different answer.

We've come to think of this as a leak. are output. They cost compute, time and several rounds of correction to produce. A well-refined page of reasoning is worth something, and letting it evaporate means paying for the same intelligence twice.

MarkdownKB is the tool we built to stop the leak. It's a knowledge base over plain markdown files that people browse and query, and it runs on your own machine.

What it is

Point MarkdownKB at directories of markdown and it indexes them for hybrid search, combining keyword matching (BM25) with vector similarity, because each catches what the other misses. Keywords find the exact function name; vectors find the paragraph that describes the idea in different words. On top of search sit a conversational view for exploring what it found, a search history you can revisit, and an interface for agents.

It's built to be search-first. It isn't a general chat tool that happens to have your documents attached; the job is finding what you already know, and the model is there to help you read it.

Why markdown

Markdown is the closest thing there is to a native format shared by people and models. It's plain text, so it versions with git, cleanly and outlives every tool that edits it. It's what models produce when they structure their thinking (headings, lists, code blocks), so feeding it back in loses nothing. It reads fine unrendered, so one file serves a person browsing and an agent querying. And it chunks well: a section of a markdown file is still a meaningful unit, which matters when you can only hand a model part of a document.

Convert, don't connect

Many retrieval systems connect to where documents already live (a wiki, a drive, a chat workspace) and query them in place. MarkdownKB deliberately doesn't. You convert what's worth keeping into markdown files you own: upload a PDF or a slide deck, clip a web page or a video transcript, or pick files out of a repository. Each path ends in the same place, a plain file on your disk.

Connecting means your knowledge base is only as available as other people's platforms, and a canceled subscription deletes years of context. Converting means the files are yours. The friction is deliberate, too: choosing what to convert is what turns a pile of documents into a knowledge base worth querying. If something is worth querying for years, it's worth owning as a file.

Why local-first

Privacy. The documents most worth querying (unpublished research, a business's playbooks, hard-won architectural decisions) are the ones you least want to send to someone else's servers. A local model reading a local knowledge base keeps them at home.

A small trusted surface. Every runtime dependency is code that runs in the same process as your documents. So MarkdownKB keeps its dependency tree thin on purpose: use the platform before reaching for a library. The dashboard charts, for example, are plain SVG styled from the theme's own variables. A charting library would have animated more smoothly, and it would also have pulled in dozens of transitive packages to sit next to your private files. The charts don't animate.

Smaller models become enough. This is the part we find most interesting. A is disproportionately better at ambiguous, novel reasoning. But if a document has already resolved the ambiguity (here's the pattern, here's the constraint, here's what we decided and why), what's left is closer to following instructions. A small local model does well at “implement this plugin following this documented pattern” and badly at “work out how plugins are done here and design one.” Retrieval shrinks the problem before the model sees it.

How it fits with agents

The payoff comes when agents use it, which they do through an MCP interface: search, read, and the rest of what a person can do in the browser. The loop it enables looks like this:

  1. An agent works on something and solves a problem.
  2. It writes down what it built and why, in markdown, as part of the work rather than as a separate chore.
  3. The knowledge base indexes it.
  4. The next agent, on this project or another, queries the knowledge base before it acts and inherits the decision instead of re-deriving it.

What took a frontier model to work out the first time becomes something a smaller model can look up the second. Frontier models don't go away; they're what you reach for when the knowledge base has no answer yet. Over time, less of the work needs them.

Two features keep that loop honest. Buckets and scopes let a query run against exactly the right slice: one project, a temporary bucket holding a new paper you want to test against your existing notes, or everything. New material can be explored without polluting the permanent collection. And a pass reads the corpus for gaps, orphaned pages and contradictions between documents, and reports them without changing anything. Its report is itself indexed, so the knowledge base can tell you where it disagrees with itself.

The intensional angle

This site is about the gap between what gets measured and what was meant. Code, tests and outputs are the of a project: the artifacts you can point at. What usually goes missing is the , the reasons. Why this design and not the obvious one. What the naming convention signals. Which constraint the odd-looking check protects.

That's exactly what a new agent session can't recover from the code alone, and exactly what gets lost when a session ends. A knowledge base of decisions is, among other things, a way of keeping intent around long enough for the next agent to honor it.

What it doesn't do

It can't make a stale document true. The loop only compounds if what goes in is accurate, which means documentation gets updated as part of the work and superseded pages get removed. Retrieval quality follows curation quality. A local model on modest hardware is also slower than a hosted frontier model, and there are problems (reasoning across a large codebase nobody has documented yet) that no amount of retrieval fixes.

What it does do is make sure the next session starts from what the last one learned. That turns out to be most of the battle.