Thali

← All models

Llama 3.1 8B Instruct

meta-llama/llama-3.1-8b-instruct

available chat hosted: global zero_retention

Meta's Llama models are the most widely deployed open-weight family. Broad ecosystem support, permissive licensing for most uses, and strong price-performance in the mid sizes.

Context window 16,384 tokens
Input ₹2.40 per million tokens
Output ₹6.00 per million tokens
Licence varies verify before production use

Call this model

from openai import OpenAI client = OpenAI( base_url="https://thaliai.in/api/v1", api_key="thali-sk-...", ) completion = client.chat.completions.create( model="meta-llama/llama-3.1-8b-instruct", messages=[{"role": "user", "content": "Hello"}], )

Works with any OpenAI SDK. Streaming, fallback lists and provider preferences are documented in the routing guide.

This model is served via global infrastructure. For workloads that must stay in India, filter the catalog for hosted_in: "in".