Tools for agentic engineers, and the teams they form.
Own your tool definitions and agent skills. Point an agent at any repo. Route to any model. Run sub-agents and review loops in sandboxes with their own budgets and server-gated permissions.
Try Kimi K3 free
Become an alpha tester and get $15 in credit to spend on any model. No card, no commitment.
Claim your creditUse Switchboard
Switchboard sits between you and your inference providers, and hands control back. Route to the model you want. Enforce rate limits and spend caps per agent or per employee. Trigger actions on specific events to keep loops tight.
The platform
Every model your agents run.
Every model behind one API at its full native capability. Route by model id, cap what any agent can spend, and attribute every token per agent. No provider keys, no juggling accounts.
Explore Switchboard SwitchboardLocalThe same API, on-device.
Run open models locally through the exact Switchboard interface. Zero metering, no data leaving the machine, and one line to fall back to hosted models when you need them.
Explore SwitchboardLocal LinecardThe registry behind routing.
Pricing, capabilities, and provenance for every model. Human-verified before it goes live. Linecard is what Switchboard reads to route, so every model runs at full capability.
Browse LinecardManaged inference for agentic engineers, and the teams they form.
Here is what each capability does, and who it works hardest for.
Feature
Description
Target use
Routing
One request shape reaches every live model in the catalog. Change models by changing the model id. Switchboard speaks each provider's native format, so nothing is lost. Streaming, tool calls, and caching all work at full capability.
Your agent, your backend, and your app all send the format they already speak.
Metering
Every token measured as it streams and itemized at provider-exact rates. Input, output, and reasoning are all counted. Cached-prompt discounts are included.
You know what any request cost the moment it finishes.
Control
Spend caps, rate limits, and model policy set per user. Enforced at the router, not in your code.
A ceiling every autonomous run stays under, and the models it can and cannot reach.
Research
Capabilities, configuration, and per-token prices for every model. Human-verified in Linecard before they go live. Pick by what a model can do, not by its name.
Compare the new frontier model against your current one before routing a single request to it.
Analytics
Usage by model, by user, and over time. Typed error records in each user timeline.
You see who used what and what it cost. When a request failed, you see why.
Key management
Switchboard mints and manages keys for your application and every end user. Nothing ships in a binary, nothing gets pasted, and rotation happens over the wire.
No key sprawl across the team, no credentials inside the app you ship.
Attribution
Every provider on one balance. Every token attributed to the agent, member, or end user that spent it.
Your own runs or a thousand end users, you see exactly where inference went.
Automation
Tools advertised into your session let agents spawn independent sub-agents against your repositories, on any model in the catalog. Rules you set at the router steer a session the moment a matching tool call fires.
Fan work out to Kimi or DeepSeek from your own agent, and build review loops that run on your logic.
The same API, on-device.
Run open models locally through the exact Switchboard interface. Same request shape, same client, and nothing leaves the machine. When a task outgrows the local model, fall back to hosted models with one line.
SwitchboardLocal ships in the Swift SDK on macOS. The pi extension routes hosted models only.
Explore SwitchboardLocal-
Zero metering
Local tokens are free tokens. Nothing books to your balance.
-
Private by construction
Prompts and outputs never leave the machine.
-
One-line hosted fallback
Same client, so escalating to a frontier model is a model id change.
Nobody should hand-maintain a model catalog.
Maintaining it yourself
Every provider is different
Request quirks and tool-calling rules vary by provider. What works against one model quietly fails against the next.
Capabilities are a moving target
Vision and tools, caching and context limits. Every model supports a different set, documented in a different place.
Pricing drifts without notice
Per-token rates change under you, and the first place you find out is the bill.
New models mean new work
Every launch is another round of setup and testing before anyone on your side can use it.
With Switchboard
Instant models, on demand.
A new model shows up on your account the day it ships, already configured and priced. Pick it by what it can do, change the model id, and keep working. There is nothing to update and nothing to maintain.
- New models arrive ready to use
- Every model runs at full capability
- Prices are always current, on every invoice