One hotkey. One memory. Every model in harness.
Workers talk directly to providers using your API keys. The dashboard never sees them, and your prompts go straight to the provider — never proxied through us.
Direct API keys, no subscriptions required. Vendors change terms overnight; your fleet doesn't depend on any one of them staying put.
Hit a limit on Opus — the router fails over to Sonnet. Sonnet full? Ollama picks it up locally, on your own hardware, for free.
The Companion packs the local router, a fleet worker, and the Wall into one installer. It detects the AI tools you already have, runs offline by default, and only escalates to a frontier model when the question deserves it.
or via npm:npm install -g @mediagato/modelreins-worker
Roz's memory and your API keys stay on your hardware. Job prompts pass through our relay only when you dispatch to a remote worker — and you can run fully local.
When you install ModelReins, you get Roz — a local intelligence that lives on your hardware. Roz's memory, routing patterns, and learned behaviors stay with you, not with ModelReins. Our servers route metadata for dispatch and coordination; we don't query the content of your prompts, your outputs, or Roz's memory across tenants.
You stay because you want to, not because you have to. If you leave, Roz goes with you. Nothing held hostage.
Personal infrastructure. Built for you, not on you.
Claude Mythos Preview, frontier-model benchmarks.
Find the bugs. Write the code. Reason about the problem. Get better every quarter.
Route the right model to the right job. Detect caps and fail over. Enforce review gates. Keep an audit trail. Work offline when you need to.
Approve. Own the outcome. Sleep at night.
When a single frontier model scores 94% on SWE-bench Verified, the bottleneck isn't capability anymore. It's which agent handles which problem, who reviews, what happens when the rate limit hits, and whether anything you shipped can be audited tomorrow. That's the layer ModelReins lives in.
Music: "Vital" by Bensound
Download the Companion. The wizard finds your Ollama, LM Studio, and any AI tools you already use — then pulls a small local model if you don't have one yet.
Install the Saddle in VSCode. It finds the local Companion automatically over loopback. One keystroke to dispatch.
Type a prompt. The router picks the right worker — local first, frontier when you need it. Output streams back in real time.
Tag workers by capability. The router matches job complexity and type to the right worker automatically.
Rate limited on Opus? Router falls over to Sonnet. Sonnet full? Ollama picks it up locally. No third-party gateway in front of your prompts.
VSCode extension. One keystroke to dispatch from your editor. Output streams back inline. No context switch.
The control plane doesn't hold your API keys. Workers fetch credentials from your local config or your vault and talk directly to providers.
File, URL, or dead man's switch. Halt every worker from one place.
Every action HMAC-signed and logged. Verify integrity. Ship to your SIEM.
Fleet awareness, cost tracking, job scheduling, multi-tenant RBAC, and more in the full guide →
AI orchestration used to be a luxury reserved for teams with dedicated ML platform engineers. The free tier is our answer to that: your laptop, your local models, your workers. No account required to start.
Everything in Team — forever. All future features included. No renewals. No price creep. Limited time — may end without notice.
Start free. Upgrade when you need more workers.
Get Started Free