Llama 3.1 8B Instruct
Updated Sep 11, 2026
llama131k context131k max outputreleased Jul 23, 2024Open weightsReasoningTool calling
Official
—
output per 1M tokens
Providers
15
offering this model
Spread
18.0×
max / min output
Offers
Metered per-token prices. Cheapest input and output are highlighted; † marks a context-tier surcharge.
| Provider | 1M input | 1M output | 1M cache read | 1M cache write | Context |
|---|---|---|---|---|---|
| Inferenceopen | $0.025 | Cheapest$0.025 | — | — | 16k |
| Kilo Gatewayopen | Cheapest$0.02 | $0.04 | — | — | 131k |
| Helicone | $0.02 | $0.05 | — | — | 16k |
| Abacusopen | $0.02 | $0.05 | — | — | 128k |
| NovitaAIopen | $0.02 | $0.05 | — | — | 16k |
| OpenRouteropen | $0.05 | $0.08 | $0.025 | — | 131k |
| Hugging Faceopen | $0.06 | $0.06 | — | — | 131k |
| NanoGPTopen | $0.0544 | $0.085 | $0.0272 | — | 131k |
| Cortecsopen | $0.167 | $0.167 | — | — | 128k |
| DigitalOceanopen | $0.198 | $0.198 | — | — | 131k |
| Pioneeropen | $0.2 | $0.2 | $0.2 | $0.2 | 128k |
| Amazon Bedrockopen | $0.22 | $0.22 | — | — | 128k |
| Amazon Bedrockusopen | $0.22 | $0.22 | — | — | 128k |
| Vercel AI Gateway | $0.22 | $0.22 | — | — | 128k |
| Neonopen | $0.15 | $0.45 | — | — | 131k |
| Nvidiafree | free | free | — | — | 16k |
Other llama models
- Hermes 2 Pro Llama 3 8B
- Llama 3.1 70B
- Llama 3.1 70B Instruct
- Llama 3.1 8B
- Llama 3.2 11B Instruct
- Llama 3.2 11B Vision Instruct
- Llama 3.2 1B Instruct
- Llama 3.2 3B
- Llama 3.2 3B Instruct
- Llama 3.2 90B Vision Instruct
- Llama 3.3 70B
- Llama-3.3-70B-Instruct
- Llama 3.3 70B Versatile
- Llama 3.3 Euryale 70B
- Llama 3 70B Instruct
- Llama 4 Maverick
- Llama 4 Maverick 17B 128E Instruct
- Llama 4 Maverick 17B 128E Instruct FP8
- Llama 4 Maverick 17B Instruct
- Llama 4 Scout
- Llama 4 Scout 17B 16E Instruct
- Llama 4 Scout 17B Instruct
- Llama-Guard-3-8B
- Llama Guard 4 12B
- Llama Prompt Guard 2 22M
- Magnum v4 72B
- MythoMax 13B
- ReMM SLERP 13B