ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE family. Built on a unified multimodal architecture, it jointly learns from text, images, audio, and video to deliver broad multimodal capabilities.
ERNIE 5.0 features significantly upgraded core capabilities and shows strong performance across benchmarks, with notable gains in multimodal understanding, instruction following, creative writing, factual accuracy, and agent planning with tool use.
Pricing
Input Modalities
- Text
- Vision
Output Modalities
- Text
Context length
- 119K tokens
Max output
- 65.5K tokens
Capabilities
- Thinking
- Streaming
- Tool calling
- Web search
- URL context
- Code interpreter
- Computer use
- File search
- Memory tool
- Structured outputs
- Citations
- Prompt caching
- Background mode
- Server-side sessions
Try this model
Python
