Toolathlon
Toolathlon evaluates agents on multi-step tasks inside real software environments, like email, files, calendars, and project trackers, exposed through MCP tool servers. The agent has to discover the right tools, chain many calls, and leave the environment in the state the task demands; grading checks the end state, not the transcript.
Tool-heavy work is where context grows fastest: every tool schema and every tool result is re-sent with each subsequent request. That makes Toolathlon a direct test of whether managing that context costs any task success.
Results
107 tasks, gpt-5.6-sol, three reasoning efforts. Each number is the change relative to the direct configuration; Induction costs include the proxy’s own optimization calls.
| Reasoning effort | Cost | Accuracy |
|---|---|---|
| medium | −40% | 0% |
| xhigh | −58% | −1% |
| max | −55% | +1% |
Each pair is one reasoning configuration: baseline and the same workload with Induction. Induction moves each configuration left at the same success rate; bars show 95% confidence intervals.
What to take from it
Cost drops 40–58% across the board. Pass rates stay within a point of direct at every effort.