← All posts

tsindex on DeepSWE: Less Context, Lower Cost

A paired DeepSWE task run with GPT-5.5 shows how structural code navigation can reduce time, tokens, and cost—and what a single comparison can tell us.

By Esteban Torres · · 7 min read

How much does code navigation matter when the coding model is already capable? We tested that question by running the same DeepSWE task with and without tsindex, our tree-sitter-based code index for coding agents.

Both runs passed. The run with tsindex finished in 4 minutes 35 seconds, compared with 5 minutes 15 seconds without it. It also consumed 469,180 fewer input tokens and cost about $0.32 less.

This post adapts our June 2, 2026 benchmark report into a closer look at the setup, measurements, and limits of that comparison. It complements our introduction to tsindex, which covers the index architecture, tool design, and a separate set of benchmarks.

The task and the two runs

The task was tasks/true-myth-iterable-collection-combinators in the TypeScript repository true-myth/true-myth. It asked the agent to implement iterable-aware sequence, traverse, zip, filtering, and asynchronous task combinators across the Maybe, Result, Task, and toolbelt APIs.

The work spanned eight files: four source files and four test files. The goal was to follow the repository’s existing patterns while satisfying the task’s tests and coverage requirements. This was a feature implementation with changes across several APIs, giving the agent repeated reasons to inspect definitions and nearby patterns.

Both variants used gpt-5.5 through OpenCode v1.15.13, driven by the Pier/Harbor evaluation framework. The agents ran in Docker containers with outbound internet access disabled.

  • With tsindex (A): the agent could query a host-side tree-sitter index through a remote MCP connection. It could request file outlines and specific symbols before reading the relevant code.
  • Without tsindex (B): the agent used its ordinary search and file-reading tools, including grep, glob, and Read, without a structural index.

The report records a single run for each variant. These are task-level observations, not averages across the full DeepSWE benchmark.

Results: the same outcome with less work

The table below uses the run without tsindex as the baseline. Reductions are calculated as (without − with) / without, so each percentage answers how much the indexed run saved compared with the fallback run.

MetricWith tsindex (A)Without tsindex (B)Reduction vs. B
Task resultPassedPassed—
Execution time4 min 35 sec5 min 15 sec40 sec (12.7%)
Agent steps27303 (10.0%)
Total input tokens3,394,7773,863,957469,180 (12.1%)
Output tokens10,00710,808801 (7.4%)
Peak context size161,098177,16916,071 (9.1%)
Total task cost (USD)$2.810471$3.126283$0.315812 (10.1%)

The original report expressed several differences in the other direction: the unindexed run took 14.5% longer, used 13.8% more input tokens, and cost 11.2% more than the indexed run. Those figures use A as the denominator. They describe the same absolute differences, but they are not the percentage reductions from baseline B.

Total input tokens accumulate across agent steps; they do not describe a single context window. Peak context size is the separate measure of the largest window reached during a run. Both measures were lower with tsindex.

Prompt caching helped both runs

The cost difference did not come from a higher cache-hit rate in the indexed run. The baseline’s rate was slightly higher.

Input-token measureWith tsindex (A)Without tsindex (B)
Cached input tokens3,231,2323,685,376
Cache-hit rate95.18%95.38%
Uncached input tokens163,545178,581

Uncached counts are total input tokens minus cached input tokens. Even with roughly 95% of input tokens cached in each run, the indexed variant submitted fewer cached tokens, fewer uncached tokens, and fewer output tokens. Prompt caching and reducing the amount of context are complementary ways to lower cost.

What changed in the agent’s workflow

Structural queries avoided a sandbox search failure

In the baseline run, search attempts failed because ripgrep was not installed in the container. With outbound internet access disabled, the agent could not download it and had to fall back to other tools.

The tsindex run could reach the host-side index through MCP and retrieve structural information without depending on that missing binary. In this environment, the index was useful both as a navigation tool and as an available path around a tooling failure.

That detail matters to the interpretation. The comparison measures the two complete setups, including the missing search dependency. It does not isolate the benefit of structural navigation against a baseline with every search tool correctly provisioned. A follow-up should include that baseline to separate the two effects.

Outlines narrowed the code the agent needed to read

Without an index, locating definitions and patterns involved raw file reads. With tsindex, the agent could start with list_file_outline and get_symbol, then inspect the relevant code ranges.

An outline provides a map of declarations. A symbol lookup can narrow the next read to a definition and, when requested, return its body. Neither requires loading every surrounding line to answer a structural question.

The indexed run’s lower input-token total and smaller peak context are consistent with that workflow. The trace observations suggest a useful mechanism, but one pair of runs cannot assign an exact share of the savings to outlines, avoided search failures, or ordinary variation in model behavior.

Fewer steps meant less repeated context

The indexed run completed in 27 agent steps instead of 30. Each additional step can carry forward context already accumulated, so the cost of a broad read can extend beyond the call that first introduced it.

In this pair, fewer steps and narrower reads coincided with lower token usage. That is the practical reason to give an agent a compact map of the code: help it reach the next relevant definition with less material to carry along.

Leaderboard numbers are context, not a ranking

The June report also recorded the following DeepSWE leaderboard averages for gpt-5.5. They used a different agent framework, mini-swe-agent, and aggregate multiple tasks. Our result used OpenCode on one task. The report did not provide the leaderboard snapshot or the reasoning-effort setting for our run, so the figures below should be read only as context reported at that time.

MeasureLeaderboard: xhighLeaderboard: mediumOur indexed task
Cost per task (USD)$6.61~$3.00$2.81
Time per task21 minNot reported4 min 35 sec
Output tokens47,000Not reported10,007

These values do not establish that tsindex is five times faster than the leaderboard or that it outperforms a particular reasoning-effort setting. Task difficulty, harness behavior, configuration, and aggregation all differ. The direct comparison we can make here is A versus B on the same task.

Extending the experiment

The next useful test is a broader set of paired runs with a fully provisioned baseline, repeated trials, and recorded model and harness settings. That would show whether the gains persist across tasks and how much run-to-run variation to expect.

The report’s proposed integration has three parts: keep an index server available, connect the sandboxed agent to it, and give the agent explicit instructions to use structural navigation.

Keep the index available during trials

The report used a host-side HTTP server on port 7337 and proposed keeping it running across trials:

tsindex --root ~/dev serve --http --port 7337

The host endpoint was http://127.0.0.1:7337; the Docker-facing address used the host.docker.internal hostname on the same port, as shown in the configuration below. Host-to-container reachability depends on the Docker and server configuration, so verify that connection in the benchmark environment before starting trials. Each trial also needs an index that corresponds to its own repository checkout.

Connect OpenCode through the harness

The proposed Pier agent configuration supplied the remote MCP mapping through OpenCode’s settings:

[agent.kwargs.opencode_config.mcp.tsindex]
type = "remote"
url = "http://host.docker.internal:7337"

This is the configuration recorded in the experiment; check it against the versions of the harness and tsindex used for a new run.

Make the navigation workflow explicit

Tool availability alone does not guarantee that an agent will use the index. The prompt should direct it to start structural investigations with an outline or symbol lookup, then read only the necessary bodies. Text search still has a role for literals, configuration keys, and prose.

For each pair, record pass/fail status alongside time, steps, cached and uncached input tokens, output tokens, peak context, and cost. The June run gives us a concrete starting point: the same task passed in both variants, with 40 seconds and roughly $0.32 saved by the indexed setup. Repeated measurements can tell us how broadly that result holds.

tsindex coding agents benchmarks DeepSWE MCP tree-sitter

We're building this — want in?

If shipping pragmatic, AI-native systems at the scale of millions of small businesses sounds like your kind of problem, we'd love to talk.

See open roles

More in AI Platform & Infrastructure

AI Platform & Infrastructure

From Model Sprawl to a Shared ML Platform

How we consolidated dozens of one-off model deployments into a single FastAPI-based, Kubernetes-served paved path that carries everything from classic prediction services to LLM agents and shared MCP tool servers — without turning shared infrastructure into a bottleneck.

LegalZoom Engineering · · 7 min read