Agent Token Cost Routing: Cut LLM Spend with Model Gateways

Agent costs are dominated by input tokens re-sent every turn, and the fix is **routing**: send each step to the cheapest model that can handle it. In 2026 this drove the **AI gateway** wave (C1, LLM Gateway, Nutanix) — a proxy that routes, caches, and meters model calls. This guide explains token cost routing and how [Smoke Monkey Harness](/solutions/what-is-an-ai-agent-harness) does it in-process through a multi-provider API, with **Smoke Monkey Canvas** showing per-agent spend.
Agent Token Cost Routing: Cut LLM Spend with Model Gateways: Agent costs are dominated by input tokens re-sent every turn, and the fix is **routing**: send each step to the cheapest model that can handle it. In 2026 this drove the **AI gateway** wave (C1, LLM Gateway, Nutanix) — a proxy that routes, caches, and meters model calls. This guide explains token cost routing and how [Smoke Monkey Harness](/solutions/what-is-an-ai-agent-harness) does it in-process through a multi-provider API, with **Smoke Monkey Canvas** showing per-agent spend. Designed as a zero-dependency, open-source TypeScript architecture under the MIT License with native Model Context Protocol (MCP) support and deterministic phase state machines.
- Agent spend is mostly re-sent context; routing the right step to the right model is the biggest lever.
- An AI gateway centralizes routing, caching, and metering across providers.
- Smoke Monkey Harness routes in-process via a multi-provider API — no extra proxy hop required.
- Smoke Monkey Canvas reports per-agent token usage so you can catch a costly agent early.
import { createAgent } from 'smoke-monkey-harness';// Planner on a frontier model, executors on a cheap local modelconst planner = createAgent({ provider: 'anthropic', model: 'claude-opus-5', workspacePath: process.cwd() });const plan = await planner.run('Decompose the task into cheap, mechanical steps. Output JSON.');const steps = JSON.parse(plan.output).steps;const results = await Promise.all(steps.map((s) =>createAgent({provider: 'ollama', // local = near-zero marginal costmodel: 'qwen3:8b',workspacePath: process.cwd(),onUsage: (u) => cost.add(u),}).run(s.instruction)));console.log('spend', cost.total());// Per-agent spend on the canvas: npx @smoke-monkey/canvas start
Watch: Related Video Guides
Anthropic Just Built an Agentic OS — Open Source Harness Breakdown
Smoke Monkey
What is an AI Gateway and Why Do Your LLM Workloads Need One?
Nutanix | I/O
Where Agent Token Cost Comes From
A chat call sends one prompt. An agent sends the whole running history on every turn, so input tokens — not output — dominate the bill, and cost grows with every iteration. A 40-step run that re-reads a large context can cost far more than the task looks like it should. Two levers fix it: reduce what you send (context compaction) and send it to the right model (routing). The 2026 AI gateway trend formalized the second lever. Smoke Monkey Harness bakes both into the runtime, and Smoke Monkey Canvas shows the resulting spend per agent.
What an AI Gateway Does
An AI gateway — C1, LLM Gateway, and Nutanix's model-routing layer are 2026 examples — sits between your app and the providers. It routes a request to the best model for the job, caches identical calls, enforces rate limits, and meters usage. Centralizing that logic is useful, but a proxy is another hop to operate. Smoke Monkey takes a different route: the harness is [multi-provider by design](/solutions/multi-provider-agent-apis), so routing is a provider and model choice inside your code. You get gateway-style routing without deploying a gateway, and you can still put a gateway in front if your organization standardizes on one.
import { createAgent } from 'smoke-monkey-harness';// Route by task shape: cheap models for mechanical steps, frontier for hard onesconst runStep = (step) => {const hard = step.kind === 'plan' || step.kind === 'review';return createAgent({provider: hard ? 'anthropic' : 'ollama',model: hard ? 'claude-opus-5' : 'qwen3:8b',workspacePath: process.cwd(),onUsage: (u) => cost.add({ model: hard ? 'opus' : 'local', ...u }),}).run(step.instruction);};await Promise.all(steps.map(runStep));
Routing Strategy for Agents
The reliable policy is frontier for planning, cheap for execution. Planning happens once and drives everything, so it pays to use a strong model; the many mechanical steps — edits, test runs, summaries — go to a fast, cheap model. Add caching for identical tool calls and a fallback provider so a rate-limit spike does not stall the run. Combine this with token cost optimization and local models for the mechanical tier, and agent spend drops sharply. Smoke Monkey Canvas shows each agent's model and token usage, so you can see which route is actually paying off.
import { createAgent } from 'smoke-monkey-harness';// Caching + fallback completes the routing policyconst agent = createAgent({provider: 'anthropic',model: 'claude-sonnet-5',workspacePath: process.cwd(),cache: { toolResults: true }, // identical tool calls are not re-billedfallback: [{ provider: 'ollama', model: 'qwen3:8b' }],onUsage: (u) => cost.add(u),});await agent.run('Run the routine maintenance steps for this repo');
Metering and Watching Spend
Routing only works if you can measure it. Smoke Monkey exposes onUsage per turn, which you can aggregate per agent, per phase, or per model — exactly the metering a gateway provides, but in-process. Pipe it to your metrics backend, or just watch it on Smoke Monkey Canvas (npx @smoke-monkey/canvas start), where each agent card reports token spend across 300+ MCP tools. That feedback loop is what lets you tune the routing policy instead of guessing — multi-provider routing you can see, on the open Smoke Monkey Harness.
Frequently Asked Questions
Q:What is agent token cost routing?
It is sending each step of an agent run to the cheapest model that can handle it — frontier models for planning, cheap or local models for mechanical execution — to cut total spend.
Q:Do I need an AI gateway for routing?
Not necessarily. Smoke Monkey Harness is multi-provider in-process, so routing is just a provider and model choice in code. A gateway is still useful if your organization standardizes on one.
Q:How much can routing save on agent costs?
Savings depend on the mix, but moving the many mechanical steps off frontier models onto local or cheap models typically removes the majority of token spend, especially combined with context compaction.
Q:Can I see cost per agent?
Yes. Smoke Monkey reports token usage per turn via onUsage, and Smoke Monkey Canvas shows per-agent spend across the fleet, so you can spot an expensive route quickly.
Related Alternatives & Comparisons
Openai Codex Alternative Open Source: Free Open Source AI Agent & Runtime (2026)
9 Best Free Open Source ChatGPT Alternatives in 2026 (Self-Hosted & Local)
Grok Alternative: Free Open Source Local Agent (Ollama + Canvas) 2026
Related Architecture Guides
View all guidesBest Open Source Coding Agents in 2026: Free, Local & Fully Hackable Harnesses
MCP Server Security Best Practices: Hardening Model Context Protocol Agents in 2026
Open Source Coding Agent Harness: Build a Forkable, Local AI Engineering Runtime
Build with Smoke Monkey Harness
Zero dependencies. 24 built-in tools. Human-in-the-loop safety. 100% open source under the MIT License.