Published on

9Router: one local endpoint in front of every AI coding account I have

9Router: one local endpoint in front of every AI coding account I have
Authors
On this page

This is one piece of my AI coding agent setup. It covers the gateway every agent talks to.

The problem

At some point I had more AI coding accounts than I could keep straight. Each agent CLI had its own idea of which key to use. When one account hit its limit, the agent stopped and waited for me. Switching meant editing config in three places.

What I wanted was one URL. Every tool points at it. Everything about accounts, keys, limits and failover lives behind it, in one place.

That is 9Router. It is an open source gateway that runs locally, speaks the OpenAI style API, and routes to real providers behind it.

9Router hub: Claude Code, Codex CLI, Pi, OpenCode and Hermes on the left all point at localhost:20128/v1; on the right, Antigravity, Codex Pro, Claude Max, OpenCode Go and LongCat

What sits behind it

ProviderHow manyWhat I use it for
Antigravity6 accountsBulk work and long agent loops
Codex Pro2 subscriptionsA second opinion with different blind spots
Claude Max1 subscriptionThe hard calls and final review
OpenCode Go1 subscriptionOpen models: GLM, Kimi, Qwen, DeepSeek, MiniMax
LongCat1Judge in fusion combos, cheap fallback

All of it is reachable at http://localhost:20128/v1. Today that endpoint lists a couple of hundred models.

Combos: the part that actually matters

An agent never asks for an account. It asks for a model name. In 9Router that name is often a combo: a named group of models with a strategy for picking between them. I have about twenty.

  • Fallback. Try the first model. If it errors or is rate limited, try the next. This is the default for anything I want to just work.
  • Round robin. Spread requests across accounts evenly, so six Antigravity accounts wear down together instead of one at a time.
  • Fusion. Send the same prompt to a small panel, then let a judge model pick or merge the best answer.
Three combo strategies: fallback skips failing models, round robin rotates accounts, fusion sends one prompt to a panel and a judge merges

When an account hits its limit, the next request goes to the next account. The agent never notices. Six Antigravity accounts means six chances before anything falls through to a paid fallback.

Fusion combos

Fusion is the strategy I like most, and the one I use most carefully.

Four fusion combos with their panels and judges: claude-fusion, gemini-fusion, longcat-fusion and sonnet-fusion
  • claude-fusion: Opus, Sonnet, Gemini Flash and LongCat answer, and the answers get merged.
  • gemini-fusion: LongCat, Gemini Flash and Sonnet, with Gemini Flash as judge.
  • longcat-fusion: Gemini Flash, GPT Luna and GLM Flash, with LongCat as judge.
  • sonnet-fusion: Gemini Flash, Sonnet and LongCat, with Sonnet as judge.

Different model families fail in different ways. A panel catches the confident wrong answer that one model alone would have given me. The cost is speed and quota, so fusion is never the default. I reach for it on architecture questions, on bugs that already fooled one model, and on reviews.

How agents reach it

Most agent CLIs can be pointed at a custom base URL. Claude Code needs a few environment variables, and Pi and Hermes read it from their config. I wrote 9agent so I never set those by hand: it reads the model list from 9Router, I pick one, and it launches the agent pointed at the gateway.

The most useful trick: Claude Code's Opus, Sonnet and Haiku slots can each map to a different 9Router model. The heavy slot gets a fusion combo, the everyday slot gets a fast model, the background slot gets a cheap one.

Keeping the model list fresh

New free models show up constantly. I do not track them by hand. Hermes runs a few scheduled jobs every morning that check for newly available models and add them to the right combos. By the time I sit down, the list is current.

What I would tell someone setting this up

  • Put the gateway in early. Once you have more than one tool and more than one account, configuring each tool separately becomes a mess fast.
  • Name combos by intent, not by provider. Agents ask for "the fusion one" or "the fast one". Which accounts serve that can change underneath without touching any agent config.
  • Keep a cheap fallback at the end of every chain. A slow answer beats an agent sitting idle on a rate limit.
  • Treat it as infrastructure. It runs on the always-on Windows PC, next to the agents, not on the laptop that sleeps.