Thali

← All models

Nemotron 3 Nano 30B A3B

nvidia/nemotron-3-nano-30b-a3b

available chat hosted: global zero_retention

NVIDIA's Nemotron line is tuned for efficient serving on NVIDIA hardware, with compact models that punch above their parameter count.

Context window 262,144 tokens
Input ₹6.00 per million tokens
Output ₹22.80 per million tokens
Licence varies verify before production use

Call this model

from openai import OpenAI client = OpenAI( base_url="https://thaliai.in/api/v1", api_key="thali-sk-...", ) completion = client.chat.completions.create( model="nvidia/nemotron-3-nano-30b-a3b", messages=[{"role": "user", "content": "Hello"}], )

Works with any OpenAI SDK. Streaming, fallback lists and provider preferences are documented in the routing guide.

This model is served via global infrastructure. For workloads that must stay in India, filter the catalog for hosted_in: "in".