Routing
By default Thali picks the healthiest backend for the model you asked for. When that isn't what you want, say so.
Falling back to another model
{
"model": "openai/gpt-oss-20b:free",
"models": ["nvidia/nemotron-nano-9b-v2:free"],
"messages": [...]
}
Tried in order. The model that actually served is reported in the response
body's model field and in the X-Thali-Model header — never the one you asked
for. If a fallback fired, you will know.
A model in the list that doesn't exist is a 404, not a silent skip. A typo in a
fallback list would otherwise degrade quietly and permanently.
Fallback also respects your balance per model, at the moment each is reached. So
{"model": "paid-one", "models": ["free-one:free"]} on an empty balance serves
the free model rather than refusing the request.
Choosing a backend
{
"model": "openai/gpt-oss-20b:free",
"provider": {
"order": ["e2e-blr-1"],
"ignore": ["old-node"],
"sort": "latency",
"allow_fallbacks": true
},
"messages": [...]
}
| Field | Effect |
|---|---|
order |
Named backends first, in this order. Others follow unless allow_fallbacks is false. |
ignore |
Never use these. |
sort |
price, throughput, latency, or priority. |
allow_fallbacks |
false restricts the request to order only. |
Backends are named by label, which you'll find in GET /api/v1/status.
The default order is cost-aware
When you express no preference, healthy backends for a model are tried cheapest first — cheapest by what serving your request actually costs, which is how the same model gets served from whichever provider is most efficient at that moment without you doing anything. A backend whose cost is unknown sorts last; a provider that throttles or fails simply yields to the next one in line.
Writing "sort": "priority" explicitly is the opt-out: it restores the
operator's static ordering for that request.
Preferences narrow. They never widen.
This is the important rule. provider.order cannot reach a backend that is
unhealthy, licence-blocked, or sitting behind a tripped circuit breaker.
Preferences filter the set we were already willing to serve from — they are not
a way to request something we decided not to serve.
If your preferences exclude everything, you get a clean 503 no_healthy_backend
rather than a silent fallback to something you asked us not to use.
Variants
| Suffix | Meaning |
|---|---|
:free |
Part of the model's name, not a routing mode. |
:nitro |
Sort backends by measured throughput. |
:floor |
Sort candidate models by price, cheapest first. |
{"model": "openai/gpt-oss-20b:free:nitro"}
{"model": "expensive-model:floor", "models": ["cheaper-model"]}
An explicit provider.sort beats a variant suffix — you wrote it out in full, so
you meant it.
How latency and throughput are measured
Each gateway keeps a rolling average per backend, fed from completed requests. It is per instance and not shared, deliberately: latency is a property of the path between that gateway and that backend, and two instances in different racks legitimately disagree about which node is fastest. Averaging them across the fleet would produce a number true for neither.
A freshly started instance routes on priority until it has observed a few
requests, and a backend nobody has measured yet always sorts behind one that has.
Current measurements are visible at GET /api/v1/status.
What isn't here
No automatic model selection — no equivalent of an "auto" router that picks a
model for you. That is the ROI router, and it is deliberately not in this
release. The seam it will plug into is select_backend(); everything on this
page is built on the same one.
Your own MCP servers
Thali brokers tools the way it routes models. Register a server and its tools appear to any MCP client you connect, billed and rate-limited exactly like the built-in ones.
curl -X POST https://thaliai.in/api/v1/mcp/servers \
-H "Authorization: Bearer $THALI_API_KEY" \
-d '{
"label": "github",
"url": "https://mcp.example.com/mcp",
"auth_scheme": "bearer",
"auth_value": "ghp_..."
}'
The label namespaces that server's tools: its search becomes
github__search, so two servers can both expose a search and neither
collides with a built-in. Labels are lowercased, must be 3–32 characters of
letters, digits and hyphens, and are unique within your account — two
customers may both use github.
GET /api/v1/mcp/servers |
list yours |
GET /api/v1/mcp/servers/{id} |
one, with its tools |
PATCH /api/v1/mcp/servers/{id} |
change url, enabled, or auth |
POST /api/v1/mcp/servers/{id}/refresh |
re-probe and re-read tools/list |
DELETE /api/v1/mcp/servers/{id} |
remove it |
Things worth knowing
Tool lists are cached. Agents call tools/list constantly, so we serve each
server's tools from the last successful probe rather than fanning out to every
one of your servers on every call — otherwise your slowest server becomes
everyone's latency. Added a tool? Call refresh.
auth_value is write-only. It is sent once and never returned by any
endpoint; responses carry has_auth_value: true instead. Omitting it on a
PATCH keeps the stored value, so you can change a URL without re-sending a
secret you cannot read back.
Your server must be on the public internet. URLs resolving to loopback, private, link-local or reserved addresses are refused — at registration and again immediately before every call, so a hostname that changes where it points after registration does not get a free pass. Redirects are not followed.
A broken server is your server's outage, not ours. A failed call marks that
server unhealthy with the reason on last_error, returns an MCP tool error, and
leaves the rest of your session working.
Brokered calls count. They pass through the same quota and metering as the built-in tools — registering your own server is not a way around the rate limit.