From code storage to code discovery
TL;DR: Databasify software. Expose its capabilities through structures that agents can efficiently index, discover and query. The important optimization of the code is how quickly and cheaply an agent can find the right thing to execute.
MCP (Model Context Protocol) may be the beginning of a broader shift: software and its APIs need to become queryable by AI agents, with minimal time and tokens spent searching for and evaluating capabilities before it calls or test some API.
Indexes, not scans
This looks surprisingly like a database problem. A database doesn’t make you scan every record to find what you need. It gives you indexes and keys that lead you straight to the relevant data. Software needs indexes into its capabilities. Loading an entire API surface into an agent’s context for every task is the equivalent of SELECT * FROM API_table to fetch a single row.
If MCP addresses this for applications, what would be the equivalent for large codebases as libraries or SDK ?
The missing read side
We already have one half of a database for code. Version control systems like Git are the write side: the source of truth that preserves history, integrity and distributed changes. But Git is optimized for writing and storing code, less for finding and using it.
Software architecture has a name for this split. CQRS (Command Query Responsibility Segregation, an architectural design pattern) separates the write model, built for consistent changes or transaction, from the read model, built for fast queries. For code, the read side exists only in fragments: code search, codebase indexers, intellisense structure, documentation MCP servers, AGENTS.md. Each solves part of the problem, and their growing number shows the need is real. What’s missing is a common model: a standard read model for agents.
The closest thing we have to a read side today is documentation, however it is built like a write model. Reference docs are essentially normalized: each fact lives in one place, organized by code structure and connected by cross-references. That works well for humans learning a system.
For an agent solving a task, it means expensive joins: the method on one page, the parameter type on another, the coordinate system somewhere else. Every join, connecting API surface dots costs tokens and time.
Samples are the exception. A good sample is already denormalized, built around a task, with everything needed in one place. But samples usually cover only the happy path of the most common tasks. The long tail of the API, where an agent needs help most, is left to the reference docs.
So reference docs have the coverage but the wrong shape, and samples have the right shape but not the coverage. An agent-facing read model needs both: denormalized like a sample, organized by intent rather than by class hierarchy, and as complete as a reference.
A single entry might look like this:
intent geocode parcel
description returns the cadastral parcel containing a point
call parcels.queryAt(lon, lat, crs)
crs default EPSG:4326, degrees
avoid for address search, use geocodeAddress
next parcels.getGeometry(id)
Materialized views
Denormalization has a well-known cost: duplicated data drifts out of sync with its source. Databases solve this with materialized views, derived from the source of truth and refreshed as it changes.
Code should do the same. The agent layer shouldn’t be yet another hand-written documentation set, but a view generated from more than signatures and refreshed with every commit: intent, and when not to use an operation, live in tests, commit messages, issues and in how consumers actually use the API.
Agents play two roles here: producers write code into the write model, consumers query the read model to assemble solutions. Each role needs a different structure.

Figure 1: Two roles, two models. The same agent can play both.
From development to manufacturing
Engineers don’t design every screw. They select parts from catalogs with precise specifications: dimensions, tolerances, materials. A capability read model is that catalog for software.
A software factory has two wings. Before anything is manufactured, someone has to find out how to build it. That happens in a prototype workshop: the problem is explored, several hypotheses are tested in parallel as cheap prototypes, and a diagnostic loop refines them until one path is proven. Then the second wing takes over, and software manufacturing actually starts.