ai-infra-jobs

tools/llm-vis / kimi-k3

Kimi K3

2.8T-parameter natively-multimodal MoE: a 3:1 hybrid of KDA linear attention and NoPE gated MLA over a 896-expert LatentMoE.

params=2.8Tactive=104Blayers=93ctx=1Mvocab=163,840

source: config.json ↗ · tech report ↗ snapshot 2026-08-24

attentionFFN / MoEnorm / residualembedding / headauxiliarydata flowresidualauxiliary
TRANSFORMER LAYER[× 93 layers]MULTI-HEAD LATENT ATTENTIONMIXTURE OF EXPERTS FFNinput_ids[B, T]Token embeddingV=163,840 → d=7168RMSNormeps=0.00001Kimi Delta Attention96 heads · d_h=128 · Δ-rule stateconv ×4 · forget gate · out gateQ low-rank projd=7168 → r=1536 → headsScaled dot-product attention96 headsNoPE · sigmoid output gateOutput projectionheads → d=7168KV latent compressiond=7168 → c_kv=512⊕ residual addRMSNormeps=0.00001SiTU-GLU MLPd=7168 → 33792 → 7168SiTU β=4/25Routersigmoid over 896 expertstop-16 · quantile balancingShared experts ×2SwiGLU d=7168 → 6144×2Routed expertsSiTU d=3584 → 3072×896 · 16 active… RMSNorm · MXFP4⊕ weighted sumΣ wᵢ·eᵢ(x)⊕ residual addRMSNormeps=0.00001LM headd=7168 → V=163,840next-token logits[B, T, 163840]MoonViT-V2 vision tower27 layers · patch 14 · d=1024→ d=7168 tokensBlock AttnRes8 blocks × 12 layers