Take the reins.

Stop juggling AI tools.
Start orchestrating them.

One hotkey. One memory. Every model in harness.

YOU
your project, your rules
THE SADDLE
where you type · ctrl+enter to dispatch
THE DIRECTOR
local first, frontier when it's earned
Claude Code
Local Ollama
OpenAI

Your keys stay with you.

Workers talk directly to providers using your API keys. The dashboard never sees them, and your prompts go straight to the provider — never proxied through us.

No vendor lock-in.

Direct API keys, no subscriptions required. Vendors change terms overnight; your fleet doesn't depend on any one of them staying put.

Rate limited? Work never stops.

Hit a limit on Opus — the router fails over to Sonnet. Sonnet full? Ollama picks it up locally, on your own hardware, for free.

The Companion

Install everything in one click.

The Companion packs the local router, a fleet worker, and the Wall into one installer. It detects the AI tools you already have, runs offline by default, and only escalates to a frontier model when the question deserves it.

or via npm:
npm install -g @mediagato/modelreins-worker

The data difference

Your keys and your memory never leave your hardware.

Roz's memory and your API keys stay on your hardware. Job prompts pass through our relay only when you dispatch to a remote worker — and you can run fully local.

When you install ModelReins, you get Roz — a local intelligence that lives on your hardware. Roz's memory, routing patterns, and learned behaviors stay with you, not with ModelReins. Our servers route metadata for dispatch and coordination; we don't query the content of your prompts, your outputs, or Roz's memory across tenants.

You stay because you want to, not because you have to. If you leave, Roz goes with you. Nothing held hostage.

Personal infrastructure. Built for you, not on you.

Same problem, different rung
93.9%
SWE-bench Verified
82.0%
Terminal-Bench 2.0
83.1%
CyberGym

Claude Mythos Preview, frontier-model benchmarks.

Frontier models

Opus · GPT · Gemini

Find the bugs. Write the code. Reason about the problem. Get better every quarter.

ModelReins

the orchestration + governance layer

Route the right model to the right job. Detect caps and fail over. Enforce review gates. Keep an audit trail. Work offline when you need to.

Humans

you and your team

Approve. Own the outcome. Sleep at night.

When a single frontier model scores 94% on SWE-bench Verified, the bottleneck isn't capability anymore. It's which agent handles which problem, who reviews, what happens when the rate limit hits, and whether anything you shipped can be audited tomorrow. That's the layer ModelReins lives in.

What it looks like

One command. Your AI workforce is online.

worker — haiku-devbox
$ npx @mediagato/modelreins-worker Worker: haiku-devbox Provider: anthropic (haiku-4.5) Server: app.modelreins.com Tags: draft,triage,cheap,fast [20:24:01] Ready — waiting for jobs... [20:24:17] >>> Job No. 803 claimed Prompt: Write a product description for ModelReins... [20:24:22] <<< Job No. 803 complete (exit 0, 4.8s) [20:24:27] >>> Job No. 804 claimed Prompt: Triage this issue: auth middleware returns 403... [20:24:29] <<< Job No. 804 complete (exit 0, 1.2s) [20:24:34] Ready — waiting for jobs...

Music: "Vital" by Bensound

Three steps

How it works

01

Install

Download the Companion. The wizard finds your Ollama, LM Studio, and any AI tools you already use — then pulls a small local model if you don't have one yet.

02

Connect your editor

Install the Saddle in VSCode. It finds the local Companion automatically over loopback. One keystroke to dispatch.

03

Dispatch

Type a prompt. The router picks the right worker — local first, frontier when you need it. Output streams back in real time.

Capabilities

What's actually in the box

Smart Routing

Tag workers by capability. The router matches job complexity and type to the right worker automatically.

Auto-Failover

Rate limited on Opus? Router falls over to Sonnet. Sonnet full? Ollama picks it up locally. No third-party gateway in front of your prompts.

The Saddle

VSCode extension. One keystroke to dispatch from your editor. Output streams back inline. No context switch.

Keys Stay With You

The control plane doesn't hold your API keys. Workers fetch credentials from your local config or your vault and talk directly to providers.

Killswitch

File, URL, or dead man's switch. Halt every worker from one place.

Signed Audit Trail

Every action HMAC-signed and logged. Verify integrity. Ship to your SIEM.

Fleet awareness, cost tracking, job scheduling, multi-tenant RBAC, and more in the full guide →

Pricing

AI orchestration used to be a luxury reserved for teams with dedicated ML platform engineers. The free tier is our answer to that: your laptop, your local models, your workers. No account required to start.

Free

$0/mo
  • 2 workers
  • 20 jobs/hour · no monthly cap
  • Killswitch
  • Dashboard + streaming
  • Job chaining
  • Cost tracking
Start Free

Team

$79/mo
  • Everything in Pro
  • Team members
  • Multi-user RBAC
Upgrade
Early Adopter

Lifetime Team Access

Everything in Team — forever. All future features included. No renewals. No price creep. Limited time — may end without notice.

$999 once
pays back in 13 months vs Team
Lock It In →

Your agents. Your infrastructure. Your rules.

Start free. Upgrade when you need more workers.

Get Started Free

5 doors · one machine