
Same Bug, Same Model—The Bill Can Still Swing Thirtyfold
Across SWE-bench Verified agent runs, token spend swings wildly between tasks—and even between repeats of the same task.
If you have ever watched an agent burn through a token budget on a “simple” fix—or fail and still leave a large charge—you have already felt the industry pain: cost is opaque, outcomes are not guaranteed, and spend is hard to forecast. A 2026 multi-lab study on OpenHands and SWE-bench Verified puts numbers on that feeling.
Across tasks, the spread is enormous
When researchers sorted 500 verified tasks by average token use, the most expensive tasks consumed about 7 million more tokens than the cheapest. The higher the cost band, the larger the variance across runs. Expensive work was not just expensive—it was unstable.
Same task, four tries—still a moving target
To check whether a single issue has a “typical” bill, the team repeated each task four times with the same model and compared the most expensive run with the cheapest. On average, the costly run used roughly twice as many tokens as the cheap one. On some tasks, the gap reached about 30×.
- You cannot treat one past run as a reliable quote for the next
- A quiet success and a looping failure can look like the same product feature with two different invoices
Why the charge feels arbitrary
Agents explore. They open files, retry edits, run tests, and sometimes wander. Because each turn reloads growing context, small behavioral detours become large token swings. From the user’s seat, that looks like magic pricing: the same bug might cost 100,000 tokens one night and millions the next.
Exact pre-run forecasts may stay out of reach. What users and vendors can still do is treat volatility as a first-class product problem: budgets, hard stops, early warnings, and clear consent before an agent keeps spending. The uncomfortable truth in the data is simple—agent token use is not a stable unit price.