AI is changing our world quickly, and there are many questions about how it will unfold. Will we reach AGI? Which labs will prevail? Will open or closed models win? But a few things seem clear: we are going to use an enormous amount of inference, it will cost a lot, and both are growing rapidly.
Against this backdrop, we are excited to introduce Induction AI. Our mission is to make AI agents more efficient.
Our first product, now available in technical preview, is a new kind of LLM endpoint optimized for agentic workloads. It works with the frontier models you already use and adds built-in context management, optimization, and routing to make these workloads more efficient. The high cost of running agents is a well-known bottleneck, and the problem will become more acute as agents take on longer and more complex tasks.
There are already a number of approaches to context optimization, but we have found that solving context management in a generic way is challenging. Techniques that look promising in isolation often fail to deliver meaningful gains under real workloads, or reduce task performance in the process. For this reason, we believe efficiency claims should be supported by benchmark results.
On tau3-banking with GPT-5.6-sol, Induction reduced costs by 67% at max reasoning. On Toolathlon we reduced costs by 55% with no loss in accuracy.
The benefits extend beyond lower cost. Greater efficiency opens up new possibilities for how inference is allocated: savings can be reinvested in higher reasoning levels, faster models, or other characteristics better suited to a particular workload, while still reducing overall cost.
Since Induction works on top of any frontier model, including future ones, these gains compound with improvements at the model layer.
We are starting with agent workloads in knowledge work, particularly those involving documents and tools. These represent a large share of applications beyond pure software development and, we believe, are poised to become one of the largest categories of agent usage. As agents take on longer, more difficult tasks, efficiency will become a critical part of the AI stack.
We are building Induction for that future. If you are working on ambitious AI agents, get in touch.