From code storage to code discovery
TL;DR: “Databasify software”. Agents scan code well. They do badly when they scan APIs they can’t see, and they waste effort scanning the same thing again and again. That is where software needs a database-like read side: capabilities indexed by intent, generated from the code and refreshed with every commit. The optimization that matters is how quickly and cheaply an agent finds the right thing to execute.
MCP (Model Context Protocol) may be the beginning of a broader shift: software and its APIs need to become queryable by AI agents, with minimal time and tokens spent searching for and evaluating capabilities before calling or testing one.
Indexes, not scans
This looks surprisingly like a database problem. A database doesn’t scan every record; it uses indexes to go straight to the data.
Agents mostly scan, and inside a repository that works well. Identifiers are precise keys, grep is fast, and nothing goes stale. Databases agree: on a small table, the query planner picks a full scan because it is cheaper.
The plan changes across an API boundary. An SDK consumer often has no source to scan, and the knowledge it needs most (what to avoid, which default applies, what comes next) is often not in code at all. Thousands of consumers also pay for the same exploration again and again. Many reads and few writes is the classic case for an index.
If MCP starts to address this for applications, what is the equivalent for libraries and SDKs?
The missing read side
We already have one half of a database for code. Version control systems like Git are the write side: the source of truth that preserves history, integrity and distributed changes. But Git is optimized for writing and storing code, less for finding and using it.
Software architecture has a name for this split. CQRS (Command Query Responsibility Segregation) separates the write model, built for consistent changes and transactions, from the read model, built for fast queries. For code, the read side exists only in fragments: code indexes like Sourcegraph’s SCIP, Aider’s repo map, generated wikis like DeepWiki, documentation servers like Context7, llms.txt and AGENTS.md. Most of them are organized by code structure or by page, not by intent. Each solves part of the problem, and their growing number shows the need is real. What’s missing is a common model: a standard read model for agents.
The closest thing we have to a read side today is documentation, but it is built like a write model. Reference docs are essentially normalized: each fact lives in one place, organized by code structure and connected by cross-references. That works well for humans learning a system.
For an agent solving a task, it means expensive joins: the method on one page, the parameter type on another, the coordinate system somewhere else. Every join costs tokens and time.
Samples are the exception. A good sample is already denormalized, built around a task, with everything needed in one place. But samples usually cover only the happy path of the most common tasks. The long tail of the API, where an agent needs help most, is left to the reference docs.
So reference docs have the coverage but the wrong shape, and samples have the right shape but not the coverage. An agent-facing read model needs both: denormalized like a sample, organized by intent rather than by class hierarchy, and as complete as a reference.
A single entry might look like this:
intent geocode parcel
description returns the cadastral parcel containing a point
call parcels.queryAt(lon, lat, crs)
crs default EPSG:4326, degrees
avoid for address search, use geocodeAddress
next parcels.getGeometry(id)
One open question is who owns the vocabulary. An agent may ask for “point in parcel” when the entry says “geocode parcel”. Intent isn’t an exact key, so this part of the read model behaves more like a search engine than a B-tree.
Materialized views
Denormalization has a well-known cost: duplicated data drifts out of sync with its source. Databases solve this with materialized views, derived from the source of truth and refreshed as it changes.
Code should do the same. The agent layer shouldn’t be yet another hand-written documentation set, but a view generated from more than signatures and refreshed with every commit. Intent, and when not to use an operation, live in tests, commit messages, issues and in how consumers actually use the API.
The view doesn’t have to be designed upfront. Databases know adaptive indexing: the index grows as a side effect of queries. Agents can do the same. Every successful scan becomes a candidate entry, kept only if its call still runs and passes a test. A generated summary can lie; an executable entry can be checked.
Agents play two roles here: producers write code into the write model, consumers query the read model to assemble solutions. Each role needs a different structure.

Figure 1: Two roles, two models. The same agent can play both.
From development to manufacturing
Engineers don’t design every screw. They select parts from catalogs with precise specifications: dimensions, tolerances, materials. A capability read model is that catalog for software.
A software factory has two wings. Before anything is manufactured, someone has to find out how to build it. That happens in a prototype workshop: the problem is explored, several hypotheses are tested in parallel as cheap prototypes, and a diagnostic loop refines them until one path is proven. Then the second wing takes over, and software manufacturing actually starts.
Takeaway
In the workshop, agents scan, write and test. They are producers, and the write model serves them well. In manufacturing, agents assemble proven parts. They are consumers, and they need a catalog: what each part does, where it fits, when not to use it. Today every agent rebuilds that catalog from scratch, for every task.
Git gave code its write model. Agents now need the read side.



























