Hy4 Preview is Tencent Hunyuan’s large language model for agents, coding, office automation, and complex tool use.Hy4 preview has a total of 770B parameters and 49B active parameters, and is primarily optimized for agent, coding, and productivity scenarios. It strengthens understanding, planning, tool invocation, and sustained execution capabilities for complex tasks. Compared with the previous generation, Hy4 preview further improves multi-step agents, code development, and productivity tasks, offering better task decomposition, context continuity, instruction following, and long-horizon execution. In coding scenarios, it further enhances code understanding, generation, modification, and the handling of complex engineering tasks; in productivity scenarios, it focuses on improving document processing, information analysis, office automation, game development, webpage generation, and cross-tool collaboration. Hy4 preview is suitable for coding agents, complex tool invocation, and various agent workflows that require multi-step planning and sustained execution, providing more reliable task completion for complex real-world business scenarios.
Pricing
- Input Tokens: $0.845 /M tokens
- Output Tokens: $2.535 /M tokens
- Cache Read: $0.0423 /M tokens
Input Modalities
- Text
Output Modalities
- Text
Context length
- 1.05M tokens
Max output
- 64K tokens
Capabilities
- Thinking
- Streaming
- Tool calling
- Web search
- URL context
- Code interpreter
- Computer use
- File search
- Memory tool
- Structured outputs
- Citations
- Prompt caching
- Background mode
- Server-side sessions
Providers
Tencent tencent-hy4-preview
Pricing$0.845$2.535
Cache$0.0423
Context1M
Max output64K
Latency3.8S
Throughput22.4TPS
Uptime
100.00% uptime 2 days ago
100.00% uptime yesterday
100.00% uptime today
Performance for hy4-preview
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python
