Llama Models
Usage
725M
22.3K
27 of 53
- llama-4-maverick
- llama-4-scout
- llama-3.3-70b
- llama-3.1-70b
- groq-llama-3.3-70b-versatile
- deepinfra-llama-4-scout-17b-16e-instruct
- llama-3.3-70b-instruct
- deepinfra-llama-4-maverick-17b-128e-instruct
- 6 more models
- llama-4-scout
- llama-4-maverick
- llama-3.3-70b
- llama-3.1-70b
- groq-llama-3.3-70b-versatile
- deepinfra-llama-4-scout-17b-16e-instruct
- deepinfra-llama-4-maverick-17b-128e-instruct
- qianfan-llama-vl-8b
- 19 more models
Which models that traffic went to
- Llama 4 Maverick72.8%528M
- Llama 4 Scout15.5%113M
- Llama 3.3 70B11.1%80.1M
- Llama 3.1 70B0.3%1.9M
- Groq Llama 3.3 70B Versatile0.2%1.5M
- Deepinfra Llama 4 Scout 17B 16e Instruct<0.1%359K
- Llama 3.3 70B Instruct<0.1%193K
- Deepinfra Llama 4 Maverick 17B 128e Instruct<0.1%45.6K
- 6 more models<0.1%9.5K
- Llama 4 Scout52.1%11.6K
- Llama 4 Maverick27.8%6.2K
- Llama 3.3 70B18.2%4K
- Llama 3.1 70B1.1%238
- Groq Llama 3.3 70B Versatile0.4%79
- Deepinfra Llama 4 Scout 17B 16e Instruct0.1%16
- Deepinfra Llama 4 Maverick 17B 128e Instruct<0.1%8
- Qianfan Llama VL 8B<0.1%8
- 19 more models0.4%79
All 53 Llama Models
Open in model listLlama on AIHubMix
Which Llama model should I start with?
llama-4-maverick at $0.2/M input — the cheapest entry here that declares tool calling, and it carries a 1.05M context. Move up to aihubmix-Llama-3-1-405B-Instruct when answer quality matters more than cost.
Why are there several entries for the same model?
Because each row is a route you can call, not a model release. Some IDs name an upstream (azure-, alicloud-, cc-), some are the open-weight repository form (meta-llama/…), and some differ only in capitalisation, kept so older integrations keep working.
The catalog does not carry a field saying which of those a given row is, so this page does not sort them into buckets it would have to invent. Every row shows that route’s own price, context and speed — compare those directly, and open a model to see the upstreams that serve it.
Do I need a separate Llama account?
No. One AIHubMix key covers every model on this page, and switching between them is a change to the model string — billing, rate limits, and logs stay in one place.
Start calling Llama in one line
One key, one endpoint, 880 models across 38 model authors.

