tsindex on DeepSWE: Less Context, Lower Cost
A paired DeepSWE task run with GPT-5.5 shows how structural code navigation can reduce time, tokens, and cost—and what a single comparison can tell us.
By Esteban Torres · · 7 min read
How much does code navigation matter when the coding model is already capable? We tested that question by running the same DeepSWE task with and without tsindex, our tree-sitter-based code index for coding agents.
Both runs passed. The run with tsindex finished in 4 minutes 35 seconds, compared with 5 minutes 15 seconds without it. It also consumed 469,180 fewer input tokens and cost about $0.32 less.
This post adapts our June 2, 2026 benchmark report into a closer look at the setup, measurements, and limits of that comparison. It complements our introduction to tsindex, which covers the index architecture, tool design, and a separate set of benchmarks.
The task and the two runs
The task was tasks/true-myth-iterable-collection-combinators in the TypeScript repository true-myth/true-myth. It asked the agent to implement iterable-aware sequence, traverse, zip, filtering, and asynchronous task combinators across the Maybe, Result, Task, and toolbelt APIs.
The work spanned eight files: four source files and four test files. The goal was to follow the repository’s existing patterns while satisfying the task’s tests and coverage requirements. This was a feature implementation with changes across several APIs, giving the agent repeated reasons to inspect definitions and nearby patterns.
Both variants used gpt-5.5 through OpenCode v1.15.13, driven by the Pier/Harbor evaluation framework. The agents ran in Docker containers with outbound internet access disabled.
- With tsindex (A): the agent could query a host-side tree-sitter index through a remote MCP connection. It could request file outlines and specific symbols before reading the relevant code.
- Without tsindex (B): the agent used its ordinary search and file-reading tools, including
grep,glob, andRead, without a structural index.
The report records a single run for each variant. These are task-level observations, not averages across the full DeepSWE benchmark.
Results: the same outcome with less work
The table below uses the run without tsindex as the baseline. Reductions are calculated as (without − with) / without, so each percentage answers how much the indexed run saved compared with the fallback run.
| Metric | With tsindex (A) | Without tsindex (B) | Reduction vs. B |
|---|---|---|---|
| Task result | Passed | Passed | — |
| Execution time | 4 min 35 sec | 5 min 15 sec | 40 sec (12.7%) |
| Agent steps | 27 | 30 | 3 (10.0%) |
| Total input tokens | 3,394,777 | 3,863,957 | 469,180 (12.1%) |
| Output tokens | 10,007 | 10,808 | 801 (7.4%) |
| Peak context size | 161,098 | 177,169 | 16,071 (9.1%) |
| Total task cost (USD) | $2.810471 | $3.126283 | $0.315812 (10.1%) |
The original report expressed several differences in the other direction: the unindexed run took 14.5% longer, used 13.8% more input tokens, and cost 11.2% more than the indexed run. Those figures use A as the denominator. They describe the same absolute differences, but they are not the percentage reductions from baseline B.
Total input tokens accumulate across agent steps; they do not describe a single context window. Peak context size is the separate measure of the largest window reached during a run. Both measures were lower with tsindex.
Prompt caching helped both runs
The cost difference did not come from a higher cache-hit rate in the indexed run. The baseline’s rate was slightly higher.
| Input-token measure | With tsindex (A) | Without tsindex (B) |
|---|---|---|
| Cached input tokens | 3,231,232 | 3,685,376 |
| Cache-hit rate | 95.18% | 95.38% |
| Uncached input tokens | 163,545 | 178,581 |
Uncached counts are total input tokens minus cached input tokens. Even with roughly 95% of input tokens cached in each run, the indexed variant submitted fewer cached tokens, fewer uncached tokens, and fewer output tokens. Prompt caching and reducing the amount of context are complementary ways to lower cost.
What changed in the agent’s workflow
Structural queries avoided a sandbox search failure
In the baseline run, search attempts failed because ripgrep was not installed in the container. With outbound internet access disabled, the agent could not download it and had to fall back to other tools.
The tsindex run could reach the host-side index through MCP and retrieve structural information without depending on that missing binary. In this environment, the index was useful both as a navigation tool and as an available path around a tooling failure.
That detail matters to the interpretation. The comparison measures the two complete setups, including the missing search dependency. It does not isolate the benefit of structural navigation against a baseline with every search tool correctly provisioned. A follow-up should include that baseline to separate the two effects.
Outlines narrowed the code the agent needed to read
Without an index, locating definitions and patterns involved raw file reads. With tsindex, the agent could start with list_file_outline and get_symbol, then inspect the relevant code ranges.
An outline provides a map of declarations. A symbol lookup can narrow the next read to a definition and, when requested, return its body. Neither requires loading every surrounding line to answer a structural question.
The indexed run’s lower input-token total and smaller peak context are consistent with that workflow. The trace observations suggest a useful mechanism, but one pair of runs cannot assign an exact share of the savings to outlines, avoided search failures, or ordinary variation in model behavior.
Fewer steps meant less repeated context
The indexed run completed in 27 agent steps instead of 30. Each additional step can carry forward context already accumulated, so the cost of a broad read can extend beyond the call that first introduced it.
In this pair, fewer steps and narrower reads coincided with lower token usage. That is the practical reason to give an agent a compact map of the code: help it reach the next relevant definition with less material to carry along.
Leaderboard numbers are context, not a ranking
The June report also recorded the following DeepSWE leaderboard averages for gpt-5.5. They used a different agent framework, mini-swe-agent, and aggregate multiple tasks. Our result used OpenCode on one task. The report did not provide the leaderboard snapshot or the reasoning-effort setting for our run, so the figures below should be read only as context reported at that time.
| Measure | Leaderboard: xhigh | Leaderboard: medium | Our indexed task |
|---|---|---|---|
| Cost per task (USD) | $6.61 | ~$3.00 | $2.81 |
| Time per task | 21 min | Not reported | 4 min 35 sec |
| Output tokens | 47,000 | Not reported | 10,007 |
These values do not establish that tsindex is five times faster than the leaderboard or that it outperforms a particular reasoning-effort setting. Task difficulty, harness behavior, configuration, and aggregation all differ. The direct comparison we can make here is A versus B on the same task.
Extending the experiment
The next useful test is a broader set of paired runs with a fully provisioned baseline, repeated trials, and recorded model and harness settings. That would show whether the gains persist across tasks and how much run-to-run variation to expect.
The report’s proposed integration has three parts: keep an index server available, connect the sandboxed agent to it, and give the agent explicit instructions to use structural navigation.
Keep the index available during trials
The report used a host-side HTTP server on port 7337 and proposed keeping it running across trials:
tsindex --root ~/dev serve --http --port 7337
The host endpoint was http://127.0.0.1:7337; the Docker-facing address used the host.docker.internal hostname on the same port, as shown in the configuration below. Host-to-container reachability depends on the Docker and server configuration, so verify that connection in the benchmark environment before starting trials. Each trial also needs an index that corresponds to its own repository checkout.
Connect OpenCode through the harness
The proposed Pier agent configuration supplied the remote MCP mapping through OpenCode’s settings:
[agent.kwargs.opencode_config.mcp.tsindex]
type = "remote"
url = "http://host.docker.internal:7337"
This is the configuration recorded in the experiment; check it against the versions of the harness and tsindex used for a new run.
Make the navigation workflow explicit
Tool availability alone does not guarantee that an agent will use the index. The prompt should direct it to start structural investigations with an outline or symbol lookup, then read only the necessary bodies. Text search still has a role for literals, configuration keys, and prose.
For each pair, record pass/fail status alongside time, steps, cached and uncached input tokens, output tokens, peak context, and cost. The June run gives us a concrete starting point: the same task passed in both variants, with 40 seconds and roughly $0.32 saved by the indexed setup. Repeated measurements can tell us how broadly that result holds.
Written by
Esteban Torres Engineering at LegalZoomWe're building this — want in?
If shipping pragmatic, AI-native systems at the scale of millions of small businesses sounds like your kind of problem, we'd love to talk.
See open rolesMore in AI Platform & Infrastructure
tsindex: Giving Coding Agents a Code Index Instead of a File Dump
Why we built a tree-sitter code index for coding agents, what real workloads taught us about tool design, and where our benchmarks show gains in speed, tokens, and cost.
Esteban Torres · · 16 min read
From Model Sprawl to a Shared ML Platform
How we consolidated dozens of one-off model deployments into a single FastAPI-based, Kubernetes-served paved path that carries everything from classic prediction services to LLM agents and shared MCP tool servers — without turning shared infrastructure into a bottleneck.
LegalZoom Engineering · · 7 min read
Putting Machine Learning in the Checkout Path Without Making Checkout Depend on It
How we run multiple production ML predictors inside a latency-sensitive purchase funnel — and why the funnel never blocks on inference.
LegalZoom Engineering · · 8 min read