From our conversations with companies at every stage of their AI adoption journey, we've seen some common patterns. First, there is an exploration period as you bring on every new tool, dole out API keys freely, and let the tokens flow. Then, you converge on the canonical tools for your organization for agentic coding, for non-technical workflows, for running and deploying agents.
As companies formalize their AI adoption, they want to manage and oversee token spend for users, but budgets and rules only go so far. The best savings are the ones users never notice. Today, we are releasing Cloudflare's Auto Router in public beta, available through AI Gateway.
Set your model to cloudflare/auto and the Auto Router will automatically route each request to a model that is capable enough for the task, without requiring an end user to think about model selection. Our early results using the Auto Router internally through our OpenCode harness show a cost savings of up to 30% when compared to using only frontier models like OpenAI Sol and Anthropic Claude Opus.
Why we built this From our own experience tracking AI spend at Cloudflare, we’ve learned managing costs requires a multipronged approach. Previously, we talked about how to set budgets and limits around AI spend, and how to see who is spending across your organization by linking employees to their AI usage. In many harnesses, including OpenCode, Claude Code, and Codex, individual users still select models manually.
Of course, not all tasks are created equal, and often individuals end up using models that are overkill for their work. For example, you don't need Opus-level intelligence if you're looking to summarize an email or chat threads. However, you wouldn't want to block that model completely from your security engineering team.
Our goal is for AI Gateway to be the control plane for organizations deploying AI internally. Because every request from every user, agent, and tool already flows through it, AI Gateway is in a unique position to do more than observe and enforce. Budgets, spend limits, and identity-aware analytics give organizations visibility and guardrails, but they still rely on individuals to make cost-conscious choices request by request.
The next step is for the gateway itself to make intelligent decisions on a user's behalf: sending each request to a model that is capable enough for the task. That way, organizations reduce spend automatically, while users keep access to the most capable models when their work actually needs them. The results We use Auto Router internally at Cloudflare within our OpenCode deployment and within Cloudflare OS , our custom agent harness.
In our internal usage, we’ve seen results comparable with frontier models for coding tasks. Auto Router does best when used across a wide range of knowledge-work tasks, like those typically found in a large organization with work spanning both technical and non-technical teams. We evaluated cloudflare/auto against OpenAI’s GPT-6 Sol and Anthropic’s Claude Opus 5.
5 on our internal general knowledge work benchmark. The benchmark uses simulated workspace tools and covers common day-to-day workflows across email, calendars, Slack, files, travel and finance. Each task requires the model to use these tools to produce a verifiable answer or complete an action.
Model Successful Trials Success Rate Total Cost Cost per success cloudflare/auto 252/291 86. 6% (+6. 2/−6.
9 pp) $2. 10 $0. 0084 Anthropic Claude Opus 5.
5 281/291 96. 6% (+2. 7/−3.
8 pp) $5. 91 $0. 0210 OpenAI GPT-6 Sol 245/291 84.
2% (+6. 5/−6. 9 pp) $2.
64 $0. 0108 97 tasks with three samples per model per task. Parenthetical values show 95% confidence intervals estimated from 10,000 task-level bootstrap resamples, preserving all three repetitions within each task.
“pp” indicates percentage points. Our Auto Router delivered similar performance to other state-of-the-art daily-driver models, coming in at 80% the cost of Sol and 35% the cost of Opus. While that may initially seem surprising, one way to frame the problem a model router solves is through the “jagged frontier” across models.
The ability to solve a problem often exists somewhere in this portfolio of models; the router’s job is to choose the right model for each task while balancing quality and price. Savings come from not paying frontier rates for non- frontier work, and they grow with how much of that work you have. Another insight is that lower token prices do not always produce lower-cost outcomes.
A model that looks cheaper on paper may end up using disproportionately more tokens to solve a problem. A router should minimize predicted trajectory cost, not just load-balance by dollars per million tokens. This is already useful today, but it’s only the beginning of what the Auto Router can learn from Cloudflare’s position in the inference path.
How it works When you send a request to cloudflare/auto , AI Gateway first builds the pool of models that can actually serve it. It filters out models that do not support the request format or execution mode, and accounts for the credentials, billing configuration, access control policies, and spend limits attached to the gateway. It will also filter out unhealthy upstream providers or models during downtime and automatically bring them back into the pool after an outage.
For the remaining candidates, the router looks at a compact view of the conversation. It considers the most recent messages, prioritizing the newest turns. The conversation is then sent to a multi-head classification model running on Workers AI and deployed on GPUs across our edge network.
The classifier produces two sets of signals. First, it assigns probabilities across 14 task categories (like coding, planning, research, data analysis). It then rates the request across four dimensions on a scale from one to five: complexity, ambiguity, stakes, and dependence on earlier context.
Originally published at blog.cloudflare.com
