The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks.
Pricing
- Input Tokens: $0.6 /M tokens
- Output Tokens: $0.6 /M tokens
Input Modalities
Providers
Groq groq-llama-3.3-70b-versatile
Pricing$0.649$0.869
Context131K
Max output32K
Latency-
Throughput-
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today
Deepinfra deepinfra-llama-3.3-70b-instant-turbo
Pricing$0.11$0.352
Context128K
Max output32K
Latency0.3S
Throughput13.0TPS
Uptime
100.00% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Azure aihubmix-Llama-3-3-70B-Instruct
Pricing$0.8$0.8
Context65K
Max output8K
Latency0.5S
Throughput57.2TPS
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today
Cerebras cerebras-llama-3.3-70b
Pricing$0.6$0.6
Context65K
Max output8K
Latency0.3S
Throughput3250.0TPS
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
0.00% uptime today
Performance for llama-3.3-70b
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python
