OmniRoute
Local/remote AI gateway exposing OpenAI-compatible REST. One key, 207+ providers, auto-fallback, RTK token saver, MCP server, A2A agents.
Setup
export OMNIROUTE_URL="http://localhost:20128" # or VPS / tunnel URL
export OMNIROUTE_KEY="sk-..." # from Dashboard → API Keys
All requests: ${OMNIROUTE_URL}/v1/... with Authorization: Bearer ${OMNIROUTE_KEY}.
Verify: curl $OMNIROUTE_URL/api/health → {"ok":true}
Discover models
curl $OMNIROUTE_URL/v1/models # chat/LLM (default)
curl $OMNIROUTE_URL/v1/models/image # image-gen
curl $OMNIROUTE_URL/v1/models/tts # text-to-speech
curl $OMNIROUTE_URL/v1/models/embedding # embeddings
curl $OMNIROUTE_URL/v1/models/web # web search + fetch
curl $OMNIROUTE_URL/v1/models/stt # speech-to-text
Use data[].id as model field in requests. Combos appear with owned_by:"combo".
Capability skills
CLI skills (omniroute binary)
| Capability | Raw URL |
|---|---|
| CLI entry point | https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute-cli/SKILL.md |
| CLI admin & lifecycle | https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute-cli-admin/SKILL.md |
| CLI providers & keys | https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute-cli-providers/SKILL.md |
| CLI cloud agents | https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute-cli-cloud/SKILL.md |
| CLI evals & benchmarks | https://raw.githubusercontent.com/diegosouzapw/OmniRoute/main/skills/omniroute-cli-eval/SKILL.md |
Errors
401→ set/refreshOMNIROUTE_KEY(Dashboard → API Keys)400 Invalid model format→ checkmodelexists in/v1/models/<kind>503 Provider circuit open→ upstream provider down; retry afterRetry-Afterseconds429→ rate limited; honorRetry-After
Differentiators vs OpenAI direct
- Auto-fallback combos (14 strategies): never stop coding even if a provider rate-limits
- RTK token saver: tool_result compressed via 47 specialized filters (git-diff, test-jest, terraform-plan, docker-logs…) — 20-40% token reduction
- Caveman mode: optional terse system prompt injection (LITE/FULL/ULTRA) — 15-25% completion reduction
- MCP + A2A servers built-in (this is the only AI router that exposes both protocols)
- Memory with FTS5 + Qdrant for persistent agent context
- Guardrails for PII masking, prompt injection detection, vision policies