N Neurarch Architectures Models Checks Data Docs Open the app

Models / kimi_linear

Kimi-Linear-48B-A3B-Instruct

Reconstructed from its own config.json with no weights read. 177K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
48.70B
48,702,058,240 parameters
In the published checkpoint
49.12B
49,122,681,728 scalars · safetensors.total, read 2025-12-16
Delta
-0.86%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `configuration_kimi.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
166
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$52.13
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsdoes not fit

Structure

168 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 512
2EmbeddingEmbedding1 × 512 × 2304
3RoPERoPE1 × 512 × 2304
4RMSNorm_1_1RMSNorm1 × 512 × 2304
5Attention_1Grouped Query Attn1 × 512 × 2304
6Add_1_attnAdd1 × 512 × 2304
7RMSNorm_1_2RMSNorm1 × 512 × 2304
8FFN_1SwiGLU1 × 512 × 2304
9Add_1_ffnAdd1 × 512 × 2304
10RMSNorm_2_1RMSNorm1 × 512 × 2304
11Attention_2Grouped Query Attn1 × 512 × 2304
12Add_2_attnAdd1 × 512 × 2304
13RMSNorm_2_2RMSNorm1 × 512 × 2304
14MoE_2Shared-Expert MoE1 × 512 × 2304
15Add_2_ffnAdd1 × 512 × 2304
16RMSNorm_3_1RMSNorm1 × 512 × 2304
17Attention_3Grouped Query Attn1 × 512 × 2304
18Add_3_attnAdd1 × 512 × 2304
19RMSNorm_3_2RMSNorm1 × 512 × 2304
20MoE_3Shared-Expert MoE1 × 512 × 2304
21Add_3_ffnAdd1 × 512 × 2304
22RMSNorm_4_1RMSNorm1 × 512 × 2304
23Attention_4Grouped Query Attn1 × 512 × 2304
24Add_4_attnAdd1 × 512 × 2304
25RMSNorm_4_2RMSNorm1 × 512 × 2304
26MoE_4Shared-Expert MoE1 × 512 × 2304
27Add_4_ffnAdd1 × 512 × 2304
28RMSNorm_5_1RMSNorm1 × 512 × 2304
29Attention_5Grouped Query Attn1 × 512 × 2304
30Add_5_attnAdd1 × 512 × 2304
31RMSNorm_5_2RMSNorm1 × 512 × 2304
32MoE_5Shared-Expert MoE1 × 512 × 2304
33Add_5_ffnAdd1 × 512 × 2304
34RMSNorm_6_1RMSNorm1 × 512 × 2304
35Attention_6Grouped Query Attn1 × 512 × 2304
36Add_6_attnAdd1 × 512 × 2304
37RMSNorm_6_2RMSNorm1 × 512 × 2304
38MoE_6Shared-Expert MoE1 × 512 × 2304
39Add_6_ffnAdd1 × 512 × 2304
40RMSNorm_7_1RMSNorm1 × 512 × 2304
41Attention_7Grouped Query Attn1 × 512 × 2304
42Add_7_attnAdd1 × 512 × 2304
43RMSNorm_7_2RMSNorm1 × 512 × 2304
44MoE_7Shared-Expert MoE1 × 512 × 2304
45Add_7_ffnAdd1 × 512 × 2304
46RMSNorm_8_1RMSNorm1 × 512 × 2304
47Attention_8Grouped Query Attn1 × 512 × 2304
48Add_8_attnAdd1 × 512 × 2304
49RMSNorm_8_2RMSNorm1 × 512 × 2304
50MoE_8Shared-Expert MoE1 × 512 × 2304
51Add_8_ffnAdd1 × 512 × 2304
52RMSNorm_9_1RMSNorm1 × 512 × 2304
53Attention_9Grouped Query Attn1 × 512 × 2304
54Add_9_attnAdd1 × 512 × 2304
55RMSNorm_9_2RMSNorm1 × 512 × 2304
56MoE_9Shared-Expert MoE1 × 512 × 2304
57Add_9_ffnAdd1 × 512 × 2304
58RMSNorm_10_1RMSNorm1 × 512 × 2304
59Attention_10Grouped Query Attn1 × 512 × 2304
60Add_10_attnAdd1 × 512 × 2304
61RMSNorm_10_2RMSNorm1 × 512 × 2304
62MoE_10Shared-Expert MoE1 × 512 × 2304
63Add_10_ffnAdd1 × 512 × 2304
64RMSNorm_11_1RMSNorm1 × 512 × 2304
65Attention_11Grouped Query Attn1 × 512 × 2304
66Add_11_attnAdd1 × 512 × 2304
67RMSNorm_11_2RMSNorm1 × 512 × 2304
68MoE_11Shared-Expert MoE1 × 512 × 2304
69Add_11_ffnAdd1 × 512 × 2304
70RMSNorm_12_1RMSNorm1 × 512 × 2304
71Attention_12Grouped Query Attn1 × 512 × 2304
72Add_12_attnAdd1 × 512 × 2304
73RMSNorm_12_2RMSNorm1 × 512 × 2304
74MoE_12Shared-Expert MoE1 × 512 × 2304
75Add_12_ffnAdd1 × 512 × 2304
76RMSNorm_13_1RMSNorm1 × 512 × 2304
77Attention_13Grouped Query Attn1 × 512 × 2304
78Add_13_attnAdd1 × 512 × 2304
79RMSNorm_13_2RMSNorm1 × 512 × 2304
80MoE_13Shared-Expert MoE1 × 512 × 2304
81Add_13_ffnAdd1 × 512 × 2304
82RMSNorm_14_1RMSNorm1 × 512 × 2304
83Attention_14Grouped Query Attn1 × 512 × 2304
84Add_14_attnAdd1 × 512 × 2304
85RMSNorm_14_2RMSNorm1 × 512 × 2304
86MoE_14Shared-Expert MoE1 × 512 × 2304
87Add_14_ffnAdd1 × 512 × 2304
88RMSNorm_15_1RMSNorm1 × 512 × 2304
89Attention_15Grouped Query Attn1 × 512 × 2304
90Add_15_attnAdd1 × 512 × 2304
91RMSNorm_15_2RMSNorm1 × 512 × 2304
92MoE_15Shared-Expert MoE1 × 512 × 2304
93Add_15_ffnAdd1 × 512 × 2304
94RMSNorm_16_1RMSNorm1 × 512 × 2304
95Attention_16Grouped Query Attn1 × 512 × 2304
96Add_16_attnAdd1 × 512 × 2304
97RMSNorm_16_2RMSNorm1 × 512 × 2304
98MoE_16Shared-Expert MoE1 × 512 × 2304
99Add_16_ffnAdd1 × 512 × 2304
100RMSNorm_17_1RMSNorm1 × 512 × 2304
101Attention_17Grouped Query Attn1 × 512 × 2304
102Add_17_attnAdd1 × 512 × 2304
103RMSNorm_17_2RMSNorm1 × 512 × 2304
104MoE_17Shared-Expert MoE1 × 512 × 2304
105Add_17_ffnAdd1 × 512 × 2304
106RMSNorm_18_1RMSNorm1 × 512 × 2304
107Attention_18Grouped Query Attn1 × 512 × 2304
108Add_18_attnAdd1 × 512 × 2304
109RMSNorm_18_2RMSNorm1 × 512 × 2304
110MoE_18Shared-Expert MoE1 × 512 × 2304
111Add_18_ffnAdd1 × 512 × 2304
112RMSNorm_19_1RMSNorm1 × 512 × 2304
113Attention_19Grouped Query Attn1 × 512 × 2304
114Add_19_attnAdd1 × 512 × 2304
115RMSNorm_19_2RMSNorm1 × 512 × 2304
116MoE_19Shared-Expert MoE1 × 512 × 2304
117Add_19_ffnAdd1 × 512 × 2304
118RMSNorm_20_1RMSNorm1 × 512 × 2304
119Attention_20Grouped Query Attn1 × 512 × 2304
120Add_20_attnAdd1 × 512 × 2304
121RMSNorm_20_2RMSNorm1 × 512 × 2304
122MoE_20Shared-Expert MoE1 × 512 × 2304
123Add_20_ffnAdd1 × 512 × 2304
124RMSNorm_21_1RMSNorm1 × 512 × 2304
125Attention_21Grouped Query Attn1 × 512 × 2304
126Add_21_attnAdd1 × 512 × 2304
127RMSNorm_21_2RMSNorm1 × 512 × 2304
128MoE_21Shared-Expert MoE1 × 512 × 2304
129Add_21_ffnAdd1 × 512 × 2304
130RMSNorm_22_1RMSNorm1 × 512 × 2304
131Attention_22Grouped Query Attn1 × 512 × 2304
132Add_22_attnAdd1 × 512 × 2304
133RMSNorm_22_2RMSNorm1 × 512 × 2304
134MoE_22Shared-Expert MoE1 × 512 × 2304
135Add_22_ffnAdd1 × 512 × 2304
136RMSNorm_23_1RMSNorm1 × 512 × 2304
137Attention_23Grouped Query Attn1 × 512 × 2304
138Add_23_attnAdd1 × 512 × 2304
139RMSNorm_23_2RMSNorm1 × 512 × 2304
140MoE_23Shared-Expert MoE1 × 512 × 2304
141Add_23_ffnAdd1 × 512 × 2304
142RMSNorm_24_1RMSNorm1 × 512 × 2304
143Attention_24Grouped Query Attn1 × 512 × 2304
144Add_24_attnAdd1 × 512 × 2304
145RMSNorm_24_2RMSNorm1 × 512 × 2304
146MoE_24Shared-Expert MoE1 × 512 × 2304
147Add_24_ffnAdd1 × 512 × 2304
148RMSNorm_25_1RMSNorm1 × 512 × 2304
149Attention_25Grouped Query Attn1 × 512 × 2304
150Add_25_attnAdd1 × 512 × 2304
151RMSNorm_25_2RMSNorm1 × 512 × 2304
152MoE_25Shared-Expert MoE1 × 512 × 2304
153Add_25_ffnAdd1 × 512 × 2304
154RMSNorm_26_1RMSNorm1 × 512 × 2304
155Attention_26Grouped Query Attn1 × 512 × 2304
156Add_26_attnAdd1 × 512 × 2304
157RMSNorm_26_2RMSNorm1 × 512 × 2304
158MoE_26Shared-Expert MoE1 × 512 × 2304
159Add_26_ffnAdd1 × 512 × 2304
160RMSNorm_27_1RMSNorm1 × 512 × 2304
161Attention_27Grouped Query Attn1 × 512 × 2304
162Add_27_attnAdd1 × 512 × 2304
163RMSNorm_27_2RMSNorm1 × 512 × 2304
164MoE_27Shared-Expert MoE1 × 512 × 2304
165Add_27_ffnAdd1 × 512 × 2304
166Final_RMSNormRMSNorm1 × 512 × 2304
167LM_HeadLinear1 × 512 × 163840
168OutputOutput1 × 512 × 163840

What the verifier says

info27 attention layers at embedDim 2304 cache full per-head K/V: about 243 KB per token at fp16, which dominates memory at long context. Grouped-query attention (e.g. 8:1) would cut this ~8×; multi-head latent attention (MLA) shrinks it ~10× or more. This is the move production LLMs make; it does not change the parameter count. Fix: Switch attention to groupedQueryAttention (set numKVHeads below numHeads, e.g. numHeads/4) or mla (a low-rank cached latent).
full-mha-serving-cost
infoAt 27 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace moonshotai/Kimi-Linear-48B-A3B-Instruct --plan --share