Quickstart
Thali is an OpenAI-compatible API. If your code already talks to OpenAI, you change two lines: the base URL and the key.
Target: zero to your first completion in under five minutes.
1. Get a key
# Send yourself a verification code
curl -X POST https://api.thali.ai/api/v1/auth/otp/start \
-H 'content-type: application/json' \
-d '{"phone": "+919876543210"}'
# Exchange the code for an account and an API key
curl -X POST https://api.thali.ai/api/v1/auth/otp/verify \
-H 'content-type: application/json' \
-d '{"phone": "+919876543210", "code": "123456"}'
The response contains your key:
{
"account_id": "01J9…",
"tier": "free",
"new_account": true,
"api_key": "thali-sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx",
"key_id": "01J9…"
}
Copy it now. We store only a SHA-256 hash, so this is the only time the full
key exists anywhere outside your terminal. Lost it? Create another
(POST /api/v1/keys, up to 5 per account).
2. Your first completion
curl
curl https://api.thali.ai/api/v1/chat/completions \
-H "Authorization: Bearer $THALI_API_KEY" \
-H 'content-type: application/json' \
-d '{
"model": "openai/gpt-oss-20b:free",
"messages": [{"role": "user", "content": "Explain UPI in one sentence."}]
}'
Python (official openai SDK)
from openai import OpenAI
client = OpenAI(
base_url="https://api.thali.ai/api/v1", # <- the only two lines that change
api_key="thali-sk-...",
)
completion = client.chat.completions.create(
model="openai/gpt-oss-20b:free",
messages=[{"role": "user", "content": "Explain UPI in one sentence."}],
)
print(completion.choices[0].message.content)
TypeScript (official openai SDK)
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.thali.ai/api/v1",
apiKey: process.env.THALI_API_KEY,
});
const completion = await client.chat.completions.create({
model: "openai/gpt-oss-20b:free",
messages: [{ role: "user", content: "Explain UPI in one sentence." }],
});
console.log(completion.choices[0].message.content);
3. Streaming
Standard SSE, terminated by data: [DONE].
stream = client.chat.completions.create(
model="openai/gpt-oss-20b:free",
messages=[{"role": "user", "content": "Write a haiku about Chennai rain."}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
Want token counts on a stream? Ask for them:
stream = client.chat.completions.create(
...,
stream=True,
stream_options={"include_usage": True},
)
The final chunk then carries usage and an empty choices list. If you don't
ask, you won't get it — the stream stays byte-identical to what OpenAI sends.
4. Embeddings
result = client.embeddings.create(
model="nvidia/llama-nemotron-embed-vl-1b-v2:free",
input=["chennai", "madurai"],
)
print(len(result.data[0].embedding))
5. What's available
curl https://api.thali.ai/api/v1/models
No key required. Each entry carries the standard OpenAI fields plus:
| Field | Meaning |
|---|---|
pricing.prompt_inr_per_mtok |
₹ per million input tokens. 0 for :free. |
pricing.completion_inr_per_mtok |
₹ per million output tokens. |
license |
The model's actual licence. We only serve what we've cleared. |
data_policy |
zero_retention — we never store your prompts or completions. |
hosted_in |
in — served from Indian infrastructure. |
available |
Whether a healthy backend is serving it right now. |
context_length |
Maximum context window. |
Note on
pricing: OpenRouter returns strings of USD-per-token. We return numbers in ₹ per million tokens, because we bill in rupees. It's the one field where a port from OpenRouter needs a change.
6. Limits
See rate-limits.md. The short version: your quota is tokens per day, shared across every key on your account, and you can check it any time:
curl https://api.thali.ai/api/v1/usage -H "Authorization: Bearer $THALI_API_KEY"
7. Errors
Identical in shape to OpenAI's, so your SDK raises the exceptions it already knows:
{
"error": {
"message": "Daily token quota exhausted: 10,000 of 10,000 tokens used on tier 'free'. Resets in 41,203s at 00:00 UTC.",
"type": "rate_limit_error",
"param": null,
"code": "tokens_per_day"
}
}
| Status | When |
|---|---|
400 |
Bad request — including max_tokens above your tier's ceiling. We reject rather than silently truncate. |
401 |
Missing, malformed, unknown, or revoked key. |
404 |
Unknown model. The message names /api/v1/models. |
429 |
A limit was hit. code names which one; Retry-After says when. |
502 |
The upstream model rejected the request. |
503 |
No healthy backend, or free-tier budget exhausted for the month. |
Running it locally
cp .env.example .env
# set FREE_MONTHLY_TOKEN_BUDGET and *_TOKENS_PER_DAY -- they default to 0,
# which disables free inference on purpose
docker compose -f infra/docker-compose.yml --profile dev up --build
That brings up the gateway, Postgres, Redis, the metering worker, and a mock
backend that speaks the OpenAI protocol without needing a GPU. Then point any
OpenAI SDK at http://localhost:8000/api/v1.