TokenSwitch: speed without losing control of cost, policy, or privacy<br>A switch on every<br>coding request,<br>flipped intelligently.<br>Stay in complete control while every request routes to the cheapest model capable of the task, across only the providers you allow.<br>Get startedHow it works<br>Connects to Claude Code, Cursor, and Codex.
your coding agentTokenSwitchclassify · routeDeepSeek V4 Flashtrivial · quick editsKimi K3economy · cheapest capableGLM 5.2balanced · most tasksSonnet 5strong · complex workClaude Opus 4.8frontier · escalate when needed
One-command setup<br>Implement with one command<br>No SDK, no code changes. One command points Claude Code or Codex at TokenSwitch, and every request is routed, capped, and logged from your next prompt on.
$ terminal
# Point Claude Code at TokenSwitch (Codex: add --tool codex)<br>npx tokenswitch-setup@latest ts_setup_your_code
✓ Redeemed your setup code and provisioned a token<br>✓ Claude Code now routes every request through TokenSwitch
Control without slowing teams<br>Governance for every<br>AI coding dollar<br>Set the rules once. Every developer’s agent inherits your policy for providers, budgets, privacy, and visibility.
See the controls →<br>providersBedrockOpenRouter<br>Data residency<br>Route through your own Bedrock or OpenRouter account, so data stays in a boundary you control.
Per-developer visibility<br>Spend, usage, and model mix per developer and repo, with no agent overhead.
promptmetadata<br>Privacy by default<br>Prompts and source are never stored. Only privacy-safe metadata leaves your environment.
cap<br>Budgets & spend caps<br>Hard caps per org, team, or developer. Over-budget requests are refused before a provider is called.
How it works<br>Control every dollar your coding agents spend<br>01<br>Classify<br>Every task is scored by complexity, cost sensitivity, and risk, so routing decisions are auditable, not a black box.
02<br>Route under policy<br>The request goes to the cheapest capable model your policy allows, across only the providers you approve, like OpenRouter or your own Amazon Bedrock account.
03<br>Escalate & log<br>If a cheaper model falls short, TokenSwitch escalates automatically, and every hop is logged and attributed to a developer, team, and repo.
Where TokenSwitch sends a coding workload<br>Only 1 in 5 requests needs a frontier model<br>per 100 requests
52 Free & open-source<br>28 Cheaper & mid-tier<br>20 Frontier
TokenSwitch escalates only when a task calls for it.
Speed for your team.<br>Control for you.<br>Point your team’s agents at TokenSwitch. Every AI coding dollar routed, capped, and logged under your policy.<br>Get started