N Neurarch Architectures Models Checks Data Docs Open the app

Models / gpt_oss

gpt-oss-20b

Reconstructed from its own config.json with no weights read. 6.4M downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
20.91B
20,908,128,448 parameters
In the published checkpoint
20.91B
20,914,757,184 scalars · safetensors.total, read 2026-09-06
Delta
-0.03%

quantized This checkpoint is stored quantized (mxfp4). The tensor count in the file counts stored elements under a packing scheme, not logical parameters, so the two numbers below are not measuring the same thing in either direction.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
148
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$80906.55
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsdoes not fit

Structure

150 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 131072
2EmbeddingEmbedding1 × 131072 × 2880
3RoPERoPE1 × 131072 × 2880
4RMSNorm_1_1RMSNorm1 × 131072 × 2880
5Attention_1Grouped Query Attn1 × 131072 × 2880
6Add_1_attnAdd1 × 131072 × 2880
7RMSNorm_1_2RMSNorm1 × 131072 × 2880
8MoE_1MoE Layer1 × 131072 × 2880
9Add_1_ffnAdd1 × 131072 × 2880
10RMSNorm_2_1RMSNorm1 × 131072 × 2880
11Attention_2Grouped Query Attn1 × 131072 × 2880
12Add_2_attnAdd1 × 131072 × 2880
13RMSNorm_2_2RMSNorm1 × 131072 × 2880
14MoE_2MoE Layer1 × 131072 × 2880
15Add_2_ffnAdd1 × 131072 × 2880
16RMSNorm_3_1RMSNorm1 × 131072 × 2880
17Attention_3Grouped Query Attn1 × 131072 × 2880
18Add_3_attnAdd1 × 131072 × 2880
19RMSNorm_3_2RMSNorm1 × 131072 × 2880
20MoE_3MoE Layer1 × 131072 × 2880
21Add_3_ffnAdd1 × 131072 × 2880
22RMSNorm_4_1RMSNorm1 × 131072 × 2880
23Attention_4Grouped Query Attn1 × 131072 × 2880
24Add_4_attnAdd1 × 131072 × 2880
25RMSNorm_4_2RMSNorm1 × 131072 × 2880
26MoE_4MoE Layer1 × 131072 × 2880
27Add_4_ffnAdd1 × 131072 × 2880
28RMSNorm_5_1RMSNorm1 × 131072 × 2880
29Attention_5Grouped Query Attn1 × 131072 × 2880
30Add_5_attnAdd1 × 131072 × 2880
31RMSNorm_5_2RMSNorm1 × 131072 × 2880
32MoE_5MoE Layer1 × 131072 × 2880
33Add_5_ffnAdd1 × 131072 × 2880
34RMSNorm_6_1RMSNorm1 × 131072 × 2880
35Attention_6Grouped Query Attn1 × 131072 × 2880
36Add_6_attnAdd1 × 131072 × 2880
37RMSNorm_6_2RMSNorm1 × 131072 × 2880
38MoE_6MoE Layer1 × 131072 × 2880
39Add_6_ffnAdd1 × 131072 × 2880
40RMSNorm_7_1RMSNorm1 × 131072 × 2880
41Attention_7Grouped Query Attn1 × 131072 × 2880
42Add_7_attnAdd1 × 131072 × 2880
43RMSNorm_7_2RMSNorm1 × 131072 × 2880
44MoE_7MoE Layer1 × 131072 × 2880
45Add_7_ffnAdd1 × 131072 × 2880
46RMSNorm_8_1RMSNorm1 × 131072 × 2880
47Attention_8Grouped Query Attn1 × 131072 × 2880
48Add_8_attnAdd1 × 131072 × 2880
49RMSNorm_8_2RMSNorm1 × 131072 × 2880
50MoE_8MoE Layer1 × 131072 × 2880
51Add_8_ffnAdd1 × 131072 × 2880
52RMSNorm_9_1RMSNorm1 × 131072 × 2880
53Attention_9Grouped Query Attn1 × 131072 × 2880
54Add_9_attnAdd1 × 131072 × 2880
55RMSNorm_9_2RMSNorm1 × 131072 × 2880
56MoE_9MoE Layer1 × 131072 × 2880
57Add_9_ffnAdd1 × 131072 × 2880
58RMSNorm_10_1RMSNorm1 × 131072 × 2880
59Attention_10Grouped Query Attn1 × 131072 × 2880
60Add_10_attnAdd1 × 131072 × 2880
61RMSNorm_10_2RMSNorm1 × 131072 × 2880
62MoE_10MoE Layer1 × 131072 × 2880
63Add_10_ffnAdd1 × 131072 × 2880
64RMSNorm_11_1RMSNorm1 × 131072 × 2880
65Attention_11Grouped Query Attn1 × 131072 × 2880
66Add_11_attnAdd1 × 131072 × 2880
67RMSNorm_11_2RMSNorm1 × 131072 × 2880
68MoE_11MoE Layer1 × 131072 × 2880
69Add_11_ffnAdd1 × 131072 × 2880
70RMSNorm_12_1RMSNorm1 × 131072 × 2880
71Attention_12Grouped Query Attn1 × 131072 × 2880
72Add_12_attnAdd1 × 131072 × 2880
73RMSNorm_12_2RMSNorm1 × 131072 × 2880
74MoE_12MoE Layer1 × 131072 × 2880
75Add_12_ffnAdd1 × 131072 × 2880
76RMSNorm_13_1RMSNorm1 × 131072 × 2880
77Attention_13Grouped Query Attn1 × 131072 × 2880
78Add_13_attnAdd1 × 131072 × 2880
79RMSNorm_13_2RMSNorm1 × 131072 × 2880
80MoE_13MoE Layer1 × 131072 × 2880
81Add_13_ffnAdd1 × 131072 × 2880
82RMSNorm_14_1RMSNorm1 × 131072 × 2880
83Attention_14Grouped Query Attn1 × 131072 × 2880
84Add_14_attnAdd1 × 131072 × 2880
85RMSNorm_14_2RMSNorm1 × 131072 × 2880
86MoE_14MoE Layer1 × 131072 × 2880
87Add_14_ffnAdd1 × 131072 × 2880
88RMSNorm_15_1RMSNorm1 × 131072 × 2880
89Attention_15Grouped Query Attn1 × 131072 × 2880
90Add_15_attnAdd1 × 131072 × 2880
91RMSNorm_15_2RMSNorm1 × 131072 × 2880
92MoE_15MoE Layer1 × 131072 × 2880
93Add_15_ffnAdd1 × 131072 × 2880
94RMSNorm_16_1RMSNorm1 × 131072 × 2880
95Attention_16Grouped Query Attn1 × 131072 × 2880
96Add_16_attnAdd1 × 131072 × 2880
97RMSNorm_16_2RMSNorm1 × 131072 × 2880
98MoE_16MoE Layer1 × 131072 × 2880
99Add_16_ffnAdd1 × 131072 × 2880
100RMSNorm_17_1RMSNorm1 × 131072 × 2880
101Attention_17Grouped Query Attn1 × 131072 × 2880
102Add_17_attnAdd1 × 131072 × 2880
103RMSNorm_17_2RMSNorm1 × 131072 × 2880
104MoE_17MoE Layer1 × 131072 × 2880
105Add_17_ffnAdd1 × 131072 × 2880
106RMSNorm_18_1RMSNorm1 × 131072 × 2880
107Attention_18Grouped Query Attn1 × 131072 × 2880
108Add_18_attnAdd1 × 131072 × 2880
109RMSNorm_18_2RMSNorm1 × 131072 × 2880
110MoE_18MoE Layer1 × 131072 × 2880
111Add_18_ffnAdd1 × 131072 × 2880
112RMSNorm_19_1RMSNorm1 × 131072 × 2880
113Attention_19Grouped Query Attn1 × 131072 × 2880
114Add_19_attnAdd1 × 131072 × 2880
115RMSNorm_19_2RMSNorm1 × 131072 × 2880
116MoE_19MoE Layer1 × 131072 × 2880
117Add_19_ffnAdd1 × 131072 × 2880
118RMSNorm_20_1RMSNorm1 × 131072 × 2880
119Attention_20Grouped Query Attn1 × 131072 × 2880
120Add_20_attnAdd1 × 131072 × 2880
121RMSNorm_20_2RMSNorm1 × 131072 × 2880
122MoE_20MoE Layer1 × 131072 × 2880
123Add_20_ffnAdd1 × 131072 × 2880
124RMSNorm_21_1RMSNorm1 × 131072 × 2880
125Attention_21Grouped Query Attn1 × 131072 × 2880
126Add_21_attnAdd1 × 131072 × 2880
127RMSNorm_21_2RMSNorm1 × 131072 × 2880
128MoE_21MoE Layer1 × 131072 × 2880
129Add_21_ffnAdd1 × 131072 × 2880
130RMSNorm_22_1RMSNorm1 × 131072 × 2880
131Attention_22Grouped Query Attn1 × 131072 × 2880
132Add_22_attnAdd1 × 131072 × 2880
133RMSNorm_22_2RMSNorm1 × 131072 × 2880
134MoE_22MoE Layer1 × 131072 × 2880
135Add_22_ffnAdd1 × 131072 × 2880
136RMSNorm_23_1RMSNorm1 × 131072 × 2880
137Attention_23Grouped Query Attn1 × 131072 × 2880
138Add_23_attnAdd1 × 131072 × 2880
139RMSNorm_23_2RMSNorm1 × 131072 × 2880
140MoE_23MoE Layer1 × 131072 × 2880
141Add_23_ffnAdd1 × 131072 × 2880
142RMSNorm_24_1RMSNorm1 × 131072 × 2880
143Attention_24Grouped Query Attn1 × 131072 × 2880
144Add_24_attnAdd1 × 131072 × 2880
145RMSNorm_24_2RMSNorm1 × 131072 × 2880
146MoE_24MoE Layer1 × 131072 × 2880
147Add_24_ffnAdd1 × 131072 × 2880
148Final_RMSNormRMSNorm1 × 131072 × 2880
149LM_HeadLinear1 × 131072 × 201088
150OutputOutput1 × 131072 × 201088

What the verifier says

infoMoE layers require an auxiliary router z-loss + load-balance loss during training to prevent expert collapse. This is not visible in the architecture diagram but must be in the training loop. Applies to all 24: MoE_1, MoE_2, MoE_3, MoE_4, MoE_5, MoE_6, MoE_7, MoE_8, MoE_9, MoE_10, MoE_11, MoE_12, MoE_13, MoE_14, MoE_15, MoE_16, MoE_17, MoE_18, MoE_19, MoE_20, MoE_21, MoE_22, MoE_23, MoE_24. Fix: Add a note on these layers. Typical aux_loss coefficient: 1e-2 (Mixtral/Switch Transformer).
moe-no-aux-loss
infoAt 24 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace openai/gpt-oss-20b --plan --share