Publication date: 01 September 2026
Tencent released and open-sourced Hy4 preview on 28 August 2026, a 770-billion-parameter model built for coding, office work and scientific research. It activates 49 billion parameters per token and supports a context window exceeding one million tokens.
The model is available as downloadable Apache 2.0 weights and through hosted APIs. For teams evaluating open models, the key question is whether Tencent’s large context and aggressive pricing translate into cheaper completed agent tasks.
What is Tencent Hy4 preview?
Hy4 preview is Tencent’s next-generation large language model for software engineering, office work, financial analysis, game development and scientific research. It uses a mixture-of-experts architecture, which means only a subset of its full parameter pool runs for each token.
Tencent lists 770 billion total parameters and 49 billion active parameters in the backbone, with a context window exceeding one million tokens. The model card adds a native multi-token prediction layer for speculative decoding and describes a sparse attention design intended to reduce long-context serving costs.
The generational jump is substantial on paper. Hy3, released in July, had 295 billion total parameters, 21 billion active parameters and a 256K context window. Hy4 preview therefore expands both capacity and context while keeping activation far below the full model size.
Hy4 preview pricing, access and key specifications
| Specification | Tencent Hy4 preview |
|---|---|
| Architecture | Mixture-of-Experts |
| Total parameters | 770B |
| Active parameters | 49B per token |
| Context | More than 1M tokens |
| Licence | Apache 2.0 |
| API price | $0.834 input / $2.501 output per 1M tokens |
| Access | Weights, Tencent Cloud TokenHub, OpenRouter |
Tencent published both full and FP8 weights on Hugging Face, with mirrors on ModelScope, GitCode and CNB. The Apache 2.0 licence permits commercial use, modification and redistribution without a custom model licence.
Hosted access is available through Tencent Cloud TokenHub and OpenRouter. Tencent lists $0.834 per million input tokens, $2.501 per million output tokens and $0.042 per million cached-input tokens. WorkBuddy and CodeBuddy also offer two weeks of free access from launch.
How does Hy4 preview compare with GLM-5.3 and Kimi K3?
Tencent’s most useful comparison is a blind internal evaluation covering 203 engineering tasks. A group of 163 Tencent experts rated Hy4 preview at 2.99 out of 4.00, compared with 2.92 for GLM-5.3 and 2.94 for Kimi K3.
Those numbers are vendor-reported and the evaluators were internal Tencent experts, so they should not be treated as an independent ranking. The margin is also small. The result is best read as evidence that Tencent believes Hy4 is competitive on practical engineering work, not proof that it is universally stronger.
Tencent also reports that Hy4 helped optimize parts of its own inference infrastructure, increasing end-to-end throughput by 31.8 percent against its baseline. The company does not provide enough external reproduction detail to treat that figure as a general serving-speed claim.
What is missing from the announcement?
Tencent is unusually explicit about two known limitations. The preview can spend longer than necessary reasoning through complex tasks and can over-verify its own work. That matters for coding agents because extra checking can increase latency and token consumption even when the final answer is correct.
Independent production measurements are still missing. A platform team considering self-hosting should verify memory requirements, throughput and total cost on its own hardware before treating 49B active parameters as a proxy for easy deployment. The full checkpoint remains very large.
What this means
Hy4 preview is worth testing now for teams evaluating open-weight coding or knowledge-work agents. The combination of Apache 2.0 weights, million-token context and comparatively low hosted pricing makes Hy4 preview easy to benchmark against existing Chinese and Western alternatives.
Start with long, multi-file tasks where the larger context window can matter. Measure completed-task cost, latency and unnecessary reasoning loops, not only benchmark scores. The preview label and Tencent’s own limitations make a controlled pilot more appropriate than an immediate production migration.
For more information, visit the official announcement of Hy4 preview on the Tencent website.