Skip to Content
Benchmarkstau3-banking

tau3-banking

tau3-bench  is Sierra Research’s benchmark for conversational agents. The agent plays a customer-service representative: it talks to a simulated user, follows a domain policy, and uses domain tools to resolve the request. A task passes only when the right actions were taken and the policy was respected.

It is a demanding setting for context management: conversations run long, tool results pile up, and the policy has to stay in force on every turn. That is the shape of workload where resending everything gets expensive.

Results

97 banking-domain tasks, gpt-5.6-sol, three reasoning efforts. Each number is the change relative to the direct configuration; Induction costs include the proxy’s own optimization calls.

Reasoning effortCostAccuracy
medium−37%+9%
xhigh−60%−1%
max−67%+7%
All-in cost per benchmark run — gpt-5.6-sol
+ InductionBaseline
Reasoning
Pass rate
tau3-banking
med
37%
$37$59
31.3%28.6%
xhigh
60%
$52$130
39.7%40.0%
max
67%
$74$225
47.1%43.8%
$0$50$100$150$200$250
Pass-rate differences are within measurement noise on all three configurations.
Task pass rate vs. all-in inference cost — tau3-banking, gpt-5.6-sol
+ InductionBaseline

Each pair is one reasoning configuration: baseline and the same workload with Induction. Induction moves each configuration left at the same success rate; bars show 95% confidence intervals.

medxhighmax$0$250all-in inference cost50%0%

What to take from it

Cost drops by a third to two-thirds at every effort, and pass rates do not pay for it: up 9% at medium, within a point at xhigh, and up 7% at max.

Last updated on
Induction
Agent Optimization Layer

Request access

Leave your details and we’ll be in touch soon.