tsindex: Giving Coding Agents a Code Index Instead of a File Dump
Why we built a tree-sitter code index for coding agents, what real workloads taught us about tool design, and where our benchmarks show gains in speed, tokens, and cost.
By Esteban Torres · · 16 min read
We are open-sourcing tsindex, a small Rust CLI and MCP server that builds a tree-sitter index over one repo or a catalog of many, stores it in SQLite, and exposes a handful of read tools that let an LLM agent navigate code without reading whole files. This post covers why we built it, how it works, and the handful of engineering lessons that turned out to matter more than the indexing itself.
The problem: agents read too much
Watch a coding agent answer “who calls parse_file?” with the built-in tools and you see the same loop every time: grep for the name, Read each matching file, scroll, Read the next one. Every file it opens adds source text to the working context. On a broad question in a large repo the agent spends most of its budget carrying source text it looked at once and never needed again.
The information the agent actually wanted is structural: where is this symbol defined, what does this file contain, which lines reference this identifier, which function contains line 412. Text search approximates these questions, but a word match is not a structural reference. rg -w path can match comments, strings, and identifiers alike. Tree-sitter lets us distinguish syntactic identifier occurrences from surrounding text; resolving which declaration an identifier refers to still requires semantic analysis.
So tsindex does the boring thing: parse every file once with tree-sitter, store symbols and identifier occurrences in SQLite, and serve compact, bounded answers over MCP.
What it is
- A single binary,
tsindex, withinit,build,update,watch,symbol,outline,enclosing,refs,query,repos, andservesubcommands. - An index at
.tsindex/index.db: four tables (repos,files,symbols,refs) plus anindex_staterow, WAL mode, foreign keys on. - 22 languages via tree-sitter grammars: Python, JavaScript/TypeScript/TSX, HTML, CSS/SCSS, Rust, Go, Java, Kotlin, Swift, C/C++/C#, PHP, Ruby, Bash, Makefiles, JSON, YAML, and Markdown headings.
- Five MCP read tools:
get_symbol,list_file_outline,enclosing_symbol,find_references, and rawquery. One write tool over stdio only:replace_symbol. - A loopback-bound HTTP dashboard with read endpoints and repository-management routes.
- Multi-repo catalogs: register several repos, index them into one database, filter any query with
repo="...".
Each language is a LanguageSpec: file extensions, manifest files, a grammar constructor, and two tree-sitter queries as string literals. Adding a language is mostly writing those two queries. Here is Python’s, in full:
LanguageSpec {
id: "python",
manifests: &["pyproject.toml", "requirements.txt"],
extensions: &["py", "pyi"],
comment_prefixes: &["#"],
language: || tree_sitter_python::LANGUAGE.into(),
symbol_query: r#"
(import_statement) @import
(import_from_statement) @import
(class_definition name: (identifier) @class.name) @class.def
(function_definition name: (identifier) @function.name) @function.def
(assignment left: (identifier) @variable.name) @variable.def
"#,
ref_query: r#"(identifier) @ref"#,
decorator_wrappers: &["decorated_definition"],
}
The build walks each repo with the ignore crate (so .gitignore is honored), parses files in parallel with rayon, and drains results into SQLite in chunks on one writer thread. Unchanged files are skipped on language, mtime, and size without being read; SHA-256 is the slow-path fallback. watch and serve --mcp keep the index fresh from filesystem notifications.
None of that is novel. The interesting part is what happened when real agents started using it.
Lesson 1: the client has a hard cap, and you have to measure it
The first serious bug report was not about correctness. find_references on a common identifier returned a result the agent never saw. Claude Code rejects an oversized MCP tool result with “exceeds maximum allowed tokens” and spills it to a file. To the agent the call simply produced nothing useful.
We had sized page limits on a guess of four characters per token. The JSON we emit runs closer to 2.5. Measured on a 1,607-reference identifier:
| Page size | Serialized payload | Outcome |
|---|---|---|
| 200 refs | 25.7 KB | accepted |
| 300 refs | 39.3 KB | accepted |
| 500 refs | 67.3 KB | rejected |
The fix had two parts. Default and maximum result limits dropped to 200 and 300. More importantly, a row limit is not a byte limit: adding three lines of snippet roughly doubles bytes per reference, so a 300-row page can still blow the cap. find_references now serializes the page it intends to send and shrinks the kept count until it fits a max_response_chars budget (45,000 by default), deriving returned, truncated, and next_offset from what it actually kept. Pagination stays exact across variable-size pages. We verified by walking the full 1,607-reference set at snippet sizes of 3 and 10 lines: 12 and 27 pages, no gaps, no duplicates, total constant throughout.
The general lesson: a tool that talks to an agent has a consumer with a hard ceiling, and that ceiling is a property of the client, not the tool. Measure it, then bound against bytes.
Lesson 2: agents under-paginate, so make truncation loud and add aggregates
Bounding worked. Then a benchmark designed to overflow the budget showed a second failure mode. Asked “how many references to path, and which files have the most?”, the agent fetched one or two pages of a 662-reference result, extrapolated, and reported a wrong total. The truncated: true flag was in the response. It was simply not loud enough to change behavior.
Two changes followed, and each has a live A/B behind it:
- A
PARTIAL RESULTwarning in the response body when a page is truncated, telling the agent to follownext_offset. In the follow-up bench the agent paged through roughly eleven 3,000-character pages to the exact answer. The prior version of the tool got the right answer only because its unbounded response happened to fit. counts_only, an aggregate mode forfind_references. The question was an aggregation, and the index already has every reference in SQLite.SELECT file, COUNT(*) GROUP BY fileis one query. Returning that instead of 662 rows for the agent to count client-side cut the payload by roughly 23x and removed the pagination problem entirely.
The counts_only bench is the clearest illustration of why structural beats textual for this class of question:
| Identifier | tsindex | Ground truth | rg -w |
|---|---|---|---|
path | 304 | 304 | 440 |
repo | 187 | 187 | 370 |
name | 210 | 210 | 611 |
Text search was fastest and cheapest, and wrong on every total. With aggregates, the tsindex arm landed within a fraction of a cent of the text-search baseline while being exactly right.
The lesson we took: the optimization that matters is rarely a tighter wire format. It is answering the question the agent is actually asking with the smallest correct payload.
Lesson 3: memory is where “works on my repo” goes to die
tsindex ran fine on its own codebase. On our largest frontend monorepo the watch process sat idle at 4.8 GB of RSS.
The cause was not our code. The debouncer we use wraps notify with a default cache that stores a file ID for every path under the watched root so it can correlate renames. notify does not honor .gitignore, so on a repo with node_modules that was roughly 240,000 entries. We never use rename correlation, since each changed path is re-resolved independently. Switching to the no-cache debouncer took idle watch memory from 4.8 GB to 16 MB.
Two smaller fixes landed in the same pass:
- Tree-sitter
Querycompilation was happening once per file per rayon worker. Compiled queries are immutable andSync, so we compile each language’s queries once and share them viaArc. Peak build RSS on a 2,173-file repo dropped from 188 MB to 86 MB, and per-worker memory flattened from about 6 MB to 1.5 MB. - The build had been collecting every file job for the whole catalog into one
Vecbefore processing. It now streams through a fixed 128-file buffer, so peak memory is bounded by chunk size rather than repo size.
A related startup fix: serve --mcp used to run a synchronous refresh before answering the MCP initialize handshake. On a 41,000-file catalog that took 1 minute 52 seconds, past the client’s 30-second startup timeout, so every session failed. The refresh now runs on a background thread, and a redundant language-detection walk was removed. A no-op refresh of the same catalog takes about 3.3 seconds.
Lesson 4: an edit tool that knows what a symbol is
Editing with the built-in tools means quoting the exact old text to anchor a replacement. For a whole-function rewrite that is the function body twice: once as the anchor, once as the new text, plus a prior read of the file.
replace_symbol takes a file, a symbol name, and a new body. It re-parses the file live, so a stale index cannot send the edit to the wrong range. It splices the new body over the symbol’s lines, then re-parses the result and refuses the write if it introduced a syntax error the file did not already have. The write goes through a temp file and rename, so a crash mid-write cannot truncate the source. When two symbols share a name and kind, the tool returns both candidates and asks for a qualified name or a row.
We benchmarked it honestly and the result is worth stating plainly. On a one-line change in a small file, replace_symbol used about 2% more tokens and was about 45% slower than a plain edit. It ships the whole symbol body and adds a round-trip. Its payoff is reliability on whole-symbol rewrites in large files, where the alternative is a brittle exact-string match on hundreds of lines. It is not a general replacement for the editor and we do not present it as one.
Lesson 5: off-by-one errors survive unit tests
Two coordinate bugs shipped and were caught by live agents, not by the test suite.
First, rows were emitted 0-based. Every file:line citation the agent produced was one line off. Editors and grep are 1-based, so the wire format is now 1-based rows with 0-based UTF-8 byte columns, which is tree-sitter’s convention and differs from character counts on non-ASCII lines. That distinction is documented in every tool schema.
Second, a multi-line snippet returns a block with the reference in the middle, plus a single range. Nothing said which line of the block the range pointed at. An agent assumed the first line and reported three call sites two rows off. The index was correct. The response was ambiguous. Multi-line snippets now carry snippet_start_row.
The fix that stuck was a testing rule: the regression test asserts the invariant twice, once on the in-process struct and once on the serialized JSON. The in-process fields agreed with each other even while the wire conversion was missing. That is exactly how the bug got in.
Lesson 6: hardening a local tool that talks to the network
The HTTP server exists for a dashboard and a few repo-management endpoints. It binds 127.0.0.1 and rejects requests whose Host header is not a localhost form, which mitigates DNS rebinding. POST routes reject cross-origin requests using Origin and Sec-Fetch-Site. Requests have socket timeouts and a 64 KiB header cap. 5xx bodies are generic and the full error chain goes to stderr.
The index has a schema version stamped only after a completed build, so a half-migrated index keeps warning rather than serving as authoritative. Read tools return an explicit “index is not ready” error during the first build instead of an empty result that looks like truth. When a file’s source no longer matches its indexed SHA, find_references says so per row instead of slicing a snippet from the wrong text.
A final review pass before open-sourcing landed 40-odd findings of this kind in one PR, plus real-process integration tests that drive the built binary over stdio and TCP. The test suite is at 243 tests.
What the benchmarks say, and what they do not
Before publishing we ran a benchmark we could stand behind: two arms, real work, a target repo that is not tsindex itself, and a blind reviewer. The harness is bench/live-ab/run.py and the full report is bench/live-ab/RESULTS.md.
The setup:
- Target: a 170,000-line Kotlin payments service at a pinned commit. Neither arm had seen it in the session.
- Arms: the same model (Claude Opus) with its regular tools, plus the tsindex MCP server in one arm. Each run was a separate headless process in a fresh git worktree with no user or project settings, hooks, or project instructions loaded, no other MCP servers, and subagents disabled. A PATH shim made the tsindex CLI unreachable from the shell in both arms, so the only route to the index was MCP.
- Tasks: five pieces of real work. Audit every production usage of a class with exact lines. Find every throw site of an exception and how it maps to an HTTP status. Trace the async fan-out when an order is confirmed. Triage a five-frame stack trace to enclosing functions with exact ranges. Rename a method across an interface, implementation, callers, and tests in twelve files.
- Rounds: three per task, arm order alternating, for 30 agent runs.
- Reviewer: a separate Opus process in a clean worktree with no MCP, given both answers shuffled and unlabeled, instructed to establish ground truth from the source before grading. It classified the task, scored each answer on correctness, completeness, precision, and usefulness, and listed every verified error.
Means across all 30 runs:
| Metric | tsindex | Baseline | Difference |
|---|---|---|---|
| reviewer score /10 | 8.9 | 8.7 | +2.3% |
| duration | 80.3s | 94.1s | 14.6% less |
| total tokens processed | 357,712 | 403,002 | 11.2% less |
| cache-read tokens | 327,040 | 371,972 | 12.1% less |
| output tokens | 6,313 | 6,616 | 4.6% less |
| tool calls | 10.9 | 12.5 | 12.8% less |
| turns | 11.9 | 13.5 | 11.9% less |
| time to first token | 3.5s | 3.7s | 4.4% less |
| output tokens per second | 78.8 | 72.3 | 9.1% more |
| cost per run, list price | $0.57 | $0.60 | 5.1% less |
| cost per 1,000 tokens | $0.00175 | $0.00170 | 2.8% more |
The reviewer picked the tsindex answer in 8 of 15 pairs, the baseline in 4, and called 3 ties. tsindex was faster in 12 of 15 paired rounds and cheaper in 8.
What the per-task breakdown adds:
- The win scales with fan-out. On the order-confirmed trace, the largest task, tsindex used 22% fewer tokens and finished 23% sooner. On the stack-trace triage it was 36% faster and 27% cheaper, because
enclosing_symbolanswers five frames in one call. On the exception hunt the two arms were within a few percent of each other. - The rename was a wash. Both arms renamed all 26 occurrences in 12 files in every round with zero leftovers, in about six tool calls, mostly
rgandsed. The tsindex arm’s 46% higher average cost on this task comes from one round where the agent’s recursive grep matched the old name inside the SQLite index file sitting untracked in the worktree, and it stopped to investigate. That is a harness artifact we have since fixed by keeping the index outside the repo, but it is also a real footgun: gitignore.tsindex/. For a mechanical text rename, text tools are the right tool, and the results say so. - Price per token barely moved. The saving is in tokens processed, not in the rate. Both arms are mostly cache reads at the same price, so cost tracks total tokens.
- Precision errors were the common failure in both arms. The reviewer’s verified errors are mostly off-by-one line citations and overstated descriptions, split roughly evenly. One tsindex-specific pattern showed up:
enclosing_symbolreports a Kotlin function’s range from its annotation line, and an agent that passes that through without checking loses a point when the task defines the start as thefunkeyword. The tool’s convention is documented, and the agent did not read it.
Caveats that still apply: one repository, one language, one model, five task shapes, three rounds. Consecutive rounds within an arm share a prompt cache, which is why cache reads are reported separately from fresh input. The reviewer is a model and can be wrong, so its verified-error lists are in the report for anyone to re-check against the source. The total benchmark cost about $30 at list price and runs in under two hours with python3 bench/live-ab/run.py --rounds 3.
An earlier, smaller same-session A/B on tsindex’s own codebase reported larger gaps, including 31% fewer cache-read tokens and half the tool calls. That harness was not committed and its numbers are best treated as a historical data point. The dashboard’s “potential savings” estimate still uses that older rate and inherits its caveats.
Built with the tools it serves
Most commits in this repo are co-authored by Claude. tsindex was built by a small team using coding agents, and the agents were using tsindex to navigate tsindex for most of that time. The benchmarks in bench/ are transcripts of those agents doing the work. That loop is where the lessons above came from: the byte cap, the under-pagination, and the off-by-one errors were all found by an agent failing at a real task, not by a test we thought to write.
Try it
git clone https://github.com/legalzoom/tsindex.git tsindex
cd tsindex
cargo install --path . --locked
tsindex --root /path/to/repo init
tsindex --root /path/to/repo build
tsindex --root /path/to/repo symbol authenticate_user --include-body
Point your MCP client at:
tsindex --root /path/to/repo serve --mcp
Ready-made Codex skills live under skills/. The README covers multi-repo catalogs, server limits, and the HTTP API.
We would like help with three things: more language specs (each is two tree-sitter queries and a test fixture), reproducible benchmarks across languages and task shapes, and reports of index staleness or coordinate bugs from repos that do not look like ours. Issues and PRs are welcome on GitHub.
Written by
Esteban Torres Engineering at LegalZoomWe're building this — want in?
If shipping pragmatic, AI-native systems at the scale of millions of small businesses sounds like your kind of problem, we'd love to talk.
See open rolesMore in AI Platform & Infrastructure
tsindex on DeepSWE: Less Context, Lower Cost
A paired DeepSWE task run with GPT-5.5 shows how structural code navigation can reduce time, tokens, and cost—and what a single comparison can tell us.
Esteban Torres · · 7 min read
From Model Sprawl to a Shared ML Platform
How we consolidated dozens of one-off model deployments into a single FastAPI-based, Kubernetes-served paved path that carries everything from classic prediction services to LLM agents and shared MCP tool servers — without turning shared infrastructure into a bottleneck.
LegalZoom Engineering · · 7 min read
Putting Machine Learning in the Checkout Path Without Making Checkout Depend on It
How we run multiple production ML predictors inside a latency-sensitive purchase funnel — and why the funnel never blocks on inference.
LegalZoom Engineering · · 8 min read