ai-infra-jobs

tools/llm-vis / deepseek-v4-flash-dspark

DeepSeek-V4-Flash-DSpark

1M-context MoE with sparse compressed attention, hyper-connections, and the DSpark speculative-decoding module.

params=284Bactive=13Blayers=43ctx=1Mvocab=129,280

source: config.json ↗ · tech report ↗ snapshot 2026-08-24

attentionFFN / MoEnorm / residualembedding / headauxiliarydata flowresidualauxiliary
TRANSFORMER LAYER[× 43 layers]MULTI-HEAD LATENT ATTENTIONMIXTURE OF EXPERTS FFNinput_ids[B, T]Token embeddingV=129,280 → d=4096RMSNormeps=0.000001Q low-rank projd=4096 → r=1024 → headsScaled dot-product attention64 headsGrouped low-rank O-projr=1024 · 8 groupsShared KV projection1 shared KV head · d_h=512KV compressionfull / CSA ×4 / HCA ×128compress θ=160KDecoupled RoPEd_rope=64 · θ=10KYaRN ×16Lightning indexer64 idx heads · d=128top-512 tokensHyper-connections4 streams · Sinkhorn ×20RMSNormeps=0.000001Routersqrtsoftplus over 256 expertstop-6 · aux-loss-free · ×1.5Shared expertSwiGLU d=4096 → 2048Routed expertsSwiGLU d=4096 → 2048×256 · 6 active · FP4 · clip ±10⊕ weighted sumΣ wᵢ·eᵢ(x)×1.5Hyper-connections4 streams · Sinkhorn ×20RMSNormeps=0.000001LM headd=4096 → V=129,280next-token logits[B, T, 129280]Hash layers ×3token-id hash → expert×3Multi-Token Predictiondepth=1 · shares embedding & headDSpark draft moduleblock=5 · rank=256taps layers 40,41,42