Skip to Content
BenchmarksToolathlon

Toolathlon

Toolathlon  evaluates agents on multi-step tasks inside real software environments, like email, files, calendars, and project trackers, exposed through MCP tool servers. The agent has to discover the right tools, chain many calls, and leave the environment in the state the task demands; grading checks the end state, not the transcript.

Tool-heavy work is where context grows fastest: every tool schema and every tool result is re-sent with each subsequent request. That makes Toolathlon a direct test of whether managing that context costs any task success.

Results

107 tasks, gpt-5.6-sol, three reasoning efforts. Each number is the change relative to the direct configuration; Induction costs include the proxy’s own optimization calls.

Reasoning effortCostAccuracy
medium−40%0%
xhigh−58%−1%
max−55%+1%
All-in cost per benchmark run — gpt-5.6-sol
+ InductionBaseline
Reasoning
Pass rate
toolathlon
med
40%
$40$67
67.8%67.5%
xhigh
58%
$67$160
71.0%72.0%
max
55%
$105$234
76.4%75.9%
$0$50$100$150$200$250
Pass-rate differences are within measurement noise on all three configurations.
Task pass rate vs. all-in inference cost — toolathlon, gpt-5.6-sol
+ InductionBaseline

Each pair is one reasoning configuration: baseline and the same workload with Induction. Induction moves each configuration left at the same success rate; bars show 95% confidence intervals.

medxhighmax$0$250all-in inference cost80%40%

What to take from it

Cost drops 40–58% across the board. Pass rates stay within a point of direct at every effort.

Last updated on
Induction
Agent Optimization Layer

Request access

Leave your details and we’ll be in touch soon.