Llama 3.3 70B
Llama logo

Llama 3.3 70B

llama-3.3-70bllms.txt
Llama
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks.

Pricing

  • Input Tokens: $0.6 /M tokens
  • Output Tokens: $0.6 /M tokens

Input Modalities

    Providers

    Groq groq-llama-3.3-70b-versatile
    Pricing$0.649$0.869
    Context131K
    Max output32K
    Latency-
    Throughput-
    Uptime
    0.00% uptime 2 days ago
    0.00% uptime yesterday
    0.00% uptime today
    Deepinfra deepinfra-llama-3.3-70b-instant-turbo
    Pricing$0.11$0.352
    Context128K
    Max output32K
    Latency0.3S
    Throughput13.0TPS
    Uptime
    100.00% uptime 2 days ago
    100.00% uptime yesterday
    100.00% uptime today
    Azure aihubmix-Llama-3-3-70B-Instruct
    Pricing$0.8$0.8
    Context65K
    Max output8K
    Latency0.5S
    Throughput57.2TPS
    Uptime
    0.00% uptime 2 days ago
    0.00% uptime yesterday
    0.00% uptime today
    Cerebras cerebras-llama-3.3-70b
    Pricing$0.6$0.6
    Context65K
    Max output8K
    Latency0.3S
    Throughput3250.0TPS
    Uptime
    0.00% uptime 2 days ago
    0.00% uptime yesterday
    0.00% uptime today

    Performance for llama-3.3-70b

    Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

    Uptime
    Loading...
    Latency
    Loading...
    Throughput
    Loading...

    Try this model

    Python
    import os
    from openai import OpenAI
    
    client = OpenAI(
        api_key=os.environ["AIHUBMIX_API_KEY"],
        base_url="https://aihubmix.com/v1",
    )
    
    response = client.chat.completions.create(
        model="llama-3.3-70b",
        messages=[
          {
            "role": "user",
            "content": "Hello, how are you?"
          }
        ],
        max_tokens=1024,
        stream=False,
    )
    
    print(response.choices[0].message.content)

    Frequently asked questions

    What is Llama 3.3 70B?

    The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks.