codingBy HowDoIUseAI Team

Why your coding agent burns tokens before it writes a single line

Most AI coding agent costs come from searching, not writing. Here's why that happens and how tools like Sonar Vortex cut it by up to 36%.

Ask a coding agent to rename a method across a mid-size codebase, and watch what happens before it touches a single character. It greps. It opens files. It greps again with slightly different terms. It reads more files to confirm what it just found. By the time it actually edits anything, you've burned through a huge chunk of your context window — and you paid for every token of it.

Here's the uncomfortable truth: most of what you pay a coding agent for isn't writing code. It's finding the right code to change in the first place.

Why does searching cost more than writing?

Writing code is cheap in token terms — a diff is usually small. Searching is expensive because it's a guessing game played with a blunt instrument: text search. An agent doesn't know your codebase's architecture, so it falls back on grep, keyword matching, and reading whole files to build a mental model from scratch, every single session.

This isn't just slow — it's a hidden cost multiplier. One breakdown of the problem found that every file an agent opens through text search persists in the conversation and gets re-billed on every subsequent turn through prompt caching, meaning when an agent reads a 600-line file on turn 40 of a 512-turn session, it does not pay for 600 lines once, it pays for 600 lines multiplied by roughly 470 remaining turns through prompt cache re-billing. That's a single unnecessary file read costing real money for the rest of the session.

What happens when the agent guesses wrong?

The token waste is the visible problem. The invisible one is worse: missed matches. If a function that needs updating is named differently than what the agent searches for, plain text search simply won't find it. The refactor looks complete, tests might even pass, and yet somewhere in the codebase a call site got skipped — because grep only finds the strings you thought to search for, not "the call that reaches your function through an interface, an alias, or another programming language."

That's the real cost of naive search: not just wasted tokens, but silent correctness gaps that surface later as production bugs.

What is Sonar Vortex and how is it different?

Sonar Vortex is built on top of SonarQube, a platform that already parses and analyzes entire codebases for quality and security. Vortex takes that existing understanding and exposes it directly to coding agents, so instead of guessing through text search, the agent can ask structural questions and get exact answers.

Sonar Vortex injects the right project context and constraints before the first line of code, then verifies every change in real time with SonarQube's algorithmic analysis, so agents produce better code with fewer tokens and less rework. In practice, that means the agent can ask things like which classes implement a given interface, or where a function is called from, and get a direct, precise answer instead of triggering a wave of file reads.

The mechanism underneath is a semantic graph called SemSitter. It's SonarSource's in-house semantic navigation engine, which keeps a local Unified Dependency Graph (UDG) of the repository updated instantly on every change. Every function, method, class, field, and parameter becomes a node in that graph, connected by typed relationships — calls, references, extends, and more.

Instead of the old pattern of grep, then a wider search, then an even wider one, the agent queries the graph for a specific node and gets back that node plus its typed relationships.

How fast does the graph actually update?

This is the part that matters most for anyone who's watched an agent work mid-edit. Code doesn't compile cleanly for most of a coding session — there are half-finished edits, broken syntax, files in flux. A tool that depends on a compiler or language server chokes on that.

Vortex sidesteps the problem entirely. The graph builds in seconds for roughly 1,000 source files and refreshes in about one millisecond after each edit. It requires no compiler, no language server, and no network calls — computation is local and in-process, adding zero tokens to the agent context. That's a meaningful design choice: the lookup itself never eats into your budget, only the answers it returns do.

How much does this actually save in tokens?

Sonar has published numbers from controlled testing, and it's worth separating the different studies since they measure slightly different things.

For refactoring-style tasks specifically — where locating every place code needs to change is most of the job — Sonar's research quantifies the efficiency leap: a study across six refactoring tasks and ten runs revealed token savings between 6% and 34%. A separate, broader benchmark comparing the graph engine against baseline grep-based navigation found up to 36% reduction in token consumption and cost per run when using the engine versus baseline grep-based navigation.

It's also worth noting the range isn't uniform across every task. One analysis of the same benchmark pointed out results ranging "down to roughly flat, depending on the task" — meaning the savings are largest specifically when navigation and cross-file tracing dominate the work, and smaller on tasks that are already localized to one or two files.

Beyond just navigation, Sonar's broader research also looked at what happens when the underlying codebase itself is cleaner and better maintained. Consistent quality, security, and dependency standards keep the code navigable, and a healthier codebase means every future agent run reads less and reasons less — Sonar research reported up to 8% fewer tokens with no drop in task completion. A more precise version of that same study found cleaner codebases used 7.2% fewer input tokens and 8.5% fewer output tokens with no drop in task completion across 540 runs. Context and code health compound — better navigation cuts search waste, and a cleaner codebase makes every future run cheaper too.

There's a quality angle here too, not just cost. In testing against a leading coding agent, Sonar Vortex reduced issues produced by 92% and lowered token consumption by up to 36%.

What languages and tools does it support right now?

Vortex's semantic navigation engine currently covers the languages most enterprise teams actually work in. The engine maintains an in-memory code graph across Java, Python, TypeScript, C#, and Rust, answering structural queries like type hierarchies, call graphs, and symbol references with exact file and line locations.

On the tooling side, it plugs into the agents developers are already using day to day. For Claude Code, Codex CLI, GitHub Copilot, Cursor, and Antigravity, the SonarQube agent plugin installs and runs the integration for you.

How do you actually set this up?

The primary entry point is the SonarQube CLI, which handles authentication and wiring everything into your agent of choice.

Step 1: Authenticate. Authenticate once with sonar auth login (browser flow; credentials stay in your OS keychain). The MCP server uses that login.

Step 2: Run the integration for your agent. The CLI has a dedicated subcommand per tool:

sonar integrate claude      # Claude Code: MCP, hooks, secrets scanning, Vortex
sonar integrate copilot     # GitHub Copilot CLI: MCP, hooks, secrets scanning
sonar integrate codex       # Codex: MCP, hooks, secrets scanning, Vortex analysis hook
sonar integrate cursor      # Cursor: MCP, secret-scanning hooks, Vortex analysis

Run these after sonar auth login. Use the /sonarqube:sonar-integrate skill if you prefer a guided flow (install/update CLI, login, then integrate).

Step 3: Let it discover and validate your project. The CLI handles this automatically — it locates your project's config using the project key auto-detection chain, or the explicit --project flag, and verifies the token. It calls SonarQube to confirm the token, organization, and project are valid.

Step 4: Restart your agent. Configuration is read at startup, so Claude reads its hook and MCP configuration at startup. Restart Claude Code for the integration to take effect.

If you're setting up Claude Code specifically, Sonar's own walkthrough for Claude Code covers the plugin path in more detail, including how the /sonarqube:sonar-integrate skill adds project-scoped resources like the Vortex context skill and MCP configuration.

For the full command reference, Sonar's official documentation on the AC/DC workflow breaks down exactly what context gets exposed to the agent and how the context and verification stages fit together.

One thing worth flagging up front: using these features requires a SonarQube Cloud Team (annual) or Enterprise plan, plus a separate subscription to the Sonar Agent Essentials product. This isn't a free tier feature — it's aimed at teams already running SonarQube for code quality who want to extend that investment into their agent workflows.

Is this worth it for your team?

If your team is running agents against small, single-file changes, the savings here will be modest — the benchmarks themselves show results trending toward flat on simpler tasks. But if you're doing cross-file refactors, dependency upgrades, or any change where "find everywhere this needs to happen" is the hard part, this is exactly the kind of overhead that compounds invisibly in your monthly token bill.

The bigger shift worth paying attention to isn't really about one product — it's the idea that token cost is fundamentally a search problem before it's a generation problem. Any team serious about running agents at scale should be asking not "which model is cheapest" but "how much of my context window is being spent on the agent figuring out where to look." That's the number actually worth optimizing.