Loading...
Open-weight 295B (21B active) MoE with hybrid fast-and-slow thinking, strong on software development and productivity tasks
Take Hunyuan Hy3 apart, layer by layer. The 3D teardown below shows its 80 transformer layers (1 dense + 79 MoE), GQA attention, and expert routing, with values from the public model config.
Hunyuan is the foundation-model brand of Tencent Holdings, the Shenzhen company founded in 1998; the Hunyuan large language model was launched on 7 September 2023 at Tencent's Global Digital Ecosystem Summit and is served through Tencent Cloud and products such as WeChat. Most recent releases are open weights: Hunyuan-Large (November 2024), HunyuanVideo (December 2024), the Hunyuan3D 2.x series (2025), HunyuanImage 2.1 and 3.0 (September 2025) and the 295B Hy3 language model (July 2026, Apache 2.0). Hunyuan 3D 3.0, announced in September 2025, is served through Tencent's Hunyuan 3D platform and API.
more about Tencent Hunyuan โ4096-dim vectors over a 120,832-token vocabulary.
A deep pre-norm stack: only layer 0 uses a dense FFN, the remaining 79 are MoE. QK-Norm stabilises the attention scores across the deep stack.
64 query heads share 8 KV heads via GQA, with QK-Norm on the query/key projections for training stability.
A router scores 192 experts and activates 8 plus 1 shared (intermediate 1536), spending 21B of 295B per token. The relatively small expert FFN keeps each expert cheap.
Final RMSNorm then projection to 120,832 logits. Hy3 targets Tencent's product stack at 256K context.