Llama 3.1 8B Instant API Pricing Calculator

Compare what your prompts cost across every major LLM — with exact token counts for OpenAI models, right in your browser.

Prices verified 2026-07-24 · per 1M tokens · sources linked per provider

Input / 1M tokens

$0.05

Cached input / 1M

Output / 1M tokens

$0.08

Context window: 131,072 tokensAPI model ID: llama-3.1-8b-instant

At Llama 3.1 8B Instant's rates ($0.05 per 1M input tokens, $0.08 per 1M output), a call with 1,000 input and 500 output tokens costs $0.00009. At 10,000 requests a month, that workload runs $0.9. A heavier call — a 100,000-token prompt returning 2,000 tokens — costs $0.00516.

Llama 3.1 8B Instant is the cheapest Groq model at this workload — Llama 3.3 70B Versatile costs 10.9× as much per call.

Llama 3.1 8B Instant's context window is 131,072 tokens.

ModelCost / callCost / month× best price
Llama 3.1 8B Instantcheapest$0.00013$0.13
GPT-OSS 120B$0.00075$0.755.8×
Llama 3.3 70B Versatile$0.00138$1.3810.6×

More Groq models: Llama 3.3 70B Versatile · GPT-OSS 120B — or all Groq pricing

Token counter

0 tokens

Token counts are exact for OpenAI models (tiktoken o200k_base, run in your browser). For other providers the count is an estimate — notably Claude Opus 4.7+/Sonnet 5 and similar newer models use a tokenizer that produces roughly 30% more tokens for the same text.

Related tools

Scaling AI content or ops? Elegant Atomics builds growth systems for B2B SaaS — Trevor’s company. Work with us →