N Neurarch Architectures Models Checks Data Docs Open the app

Models / qwen2

Qwen2.5-7B-Instruct

Reconstructed from its own config.json with no weights read. 11.5M downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
7.62B
7,615,639,552 parameters
In the published checkpoint
7.62B
7,615,616,512 scalars · safetensors.total, read 2026-09-06
Delta
+0.00%

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
172
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$12705.07
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsdoes not fit

Structure

174 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 32768
2EmbeddingEmbedding1 × 32768 × 3584
3RoPERoPE1 × 32768 × 3584
4RMSNorm_1_1RMSNorm1 × 32768 × 3584
5Attention_1Grouped Query Attn1 × 32768 × 3584
6Add_1_attnAdd1 × 32768 × 3584
7RMSNorm_1_2RMSNorm1 × 32768 × 3584
8FFN_1SwiGLU1 × 32768 × 3584
9Add_1_ffnAdd1 × 32768 × 3584
10RMSNorm_2_1RMSNorm1 × 32768 × 3584
11Attention_2Grouped Query Attn1 × 32768 × 3584
12Add_2_attnAdd1 × 32768 × 3584
13RMSNorm_2_2RMSNorm1 × 32768 × 3584
14FFN_2SwiGLU1 × 32768 × 3584
15Add_2_ffnAdd1 × 32768 × 3584
16RMSNorm_3_1RMSNorm1 × 32768 × 3584
17Attention_3Grouped Query Attn1 × 32768 × 3584
18Add_3_attnAdd1 × 32768 × 3584
19RMSNorm_3_2RMSNorm1 × 32768 × 3584
20FFN_3SwiGLU1 × 32768 × 3584
21Add_3_ffnAdd1 × 32768 × 3584
22RMSNorm_4_1RMSNorm1 × 32768 × 3584
23Attention_4Grouped Query Attn1 × 32768 × 3584
24Add_4_attnAdd1 × 32768 × 3584
25RMSNorm_4_2RMSNorm1 × 32768 × 3584
26FFN_4SwiGLU1 × 32768 × 3584
27Add_4_ffnAdd1 × 32768 × 3584
28RMSNorm_5_1RMSNorm1 × 32768 × 3584
29Attention_5Grouped Query Attn1 × 32768 × 3584
30Add_5_attnAdd1 × 32768 × 3584
31RMSNorm_5_2RMSNorm1 × 32768 × 3584
32FFN_5SwiGLU1 × 32768 × 3584
33Add_5_ffnAdd1 × 32768 × 3584
34RMSNorm_6_1RMSNorm1 × 32768 × 3584
35Attention_6Grouped Query Attn1 × 32768 × 3584
36Add_6_attnAdd1 × 32768 × 3584
37RMSNorm_6_2RMSNorm1 × 32768 × 3584
38FFN_6SwiGLU1 × 32768 × 3584
39Add_6_ffnAdd1 × 32768 × 3584
40RMSNorm_7_1RMSNorm1 × 32768 × 3584
41Attention_7Grouped Query Attn1 × 32768 × 3584
42Add_7_attnAdd1 × 32768 × 3584
43RMSNorm_7_2RMSNorm1 × 32768 × 3584
44FFN_7SwiGLU1 × 32768 × 3584
45Add_7_ffnAdd1 × 32768 × 3584
46RMSNorm_8_1RMSNorm1 × 32768 × 3584
47Attention_8Grouped Query Attn1 × 32768 × 3584
48Add_8_attnAdd1 × 32768 × 3584
49RMSNorm_8_2RMSNorm1 × 32768 × 3584
50FFN_8SwiGLU1 × 32768 × 3584
51Add_8_ffnAdd1 × 32768 × 3584
52RMSNorm_9_1RMSNorm1 × 32768 × 3584
53Attention_9Grouped Query Attn1 × 32768 × 3584
54Add_9_attnAdd1 × 32768 × 3584
55RMSNorm_9_2RMSNorm1 × 32768 × 3584
56FFN_9SwiGLU1 × 32768 × 3584
57Add_9_ffnAdd1 × 32768 × 3584
58RMSNorm_10_1RMSNorm1 × 32768 × 3584
59Attention_10Grouped Query Attn1 × 32768 × 3584
60Add_10_attnAdd1 × 32768 × 3584
61RMSNorm_10_2RMSNorm1 × 32768 × 3584
62FFN_10SwiGLU1 × 32768 × 3584
63Add_10_ffnAdd1 × 32768 × 3584
64RMSNorm_11_1RMSNorm1 × 32768 × 3584
65Attention_11Grouped Query Attn1 × 32768 × 3584
66Add_11_attnAdd1 × 32768 × 3584
67RMSNorm_11_2RMSNorm1 × 32768 × 3584
68FFN_11SwiGLU1 × 32768 × 3584
69Add_11_ffnAdd1 × 32768 × 3584
70RMSNorm_12_1RMSNorm1 × 32768 × 3584
71Attention_12Grouped Query Attn1 × 32768 × 3584
72Add_12_attnAdd1 × 32768 × 3584
73RMSNorm_12_2RMSNorm1 × 32768 × 3584
74FFN_12SwiGLU1 × 32768 × 3584
75Add_12_ffnAdd1 × 32768 × 3584
76RMSNorm_13_1RMSNorm1 × 32768 × 3584
77Attention_13Grouped Query Attn1 × 32768 × 3584
78Add_13_attnAdd1 × 32768 × 3584
79RMSNorm_13_2RMSNorm1 × 32768 × 3584
80FFN_13SwiGLU1 × 32768 × 3584
81Add_13_ffnAdd1 × 32768 × 3584
82RMSNorm_14_1RMSNorm1 × 32768 × 3584
83Attention_14Grouped Query Attn1 × 32768 × 3584
84Add_14_attnAdd1 × 32768 × 3584
85RMSNorm_14_2RMSNorm1 × 32768 × 3584
86FFN_14SwiGLU1 × 32768 × 3584
87Add_14_ffnAdd1 × 32768 × 3584
88RMSNorm_15_1RMSNorm1 × 32768 × 3584
89Attention_15Grouped Query Attn1 × 32768 × 3584
90Add_15_attnAdd1 × 32768 × 3584
91RMSNorm_15_2RMSNorm1 × 32768 × 3584
92FFN_15SwiGLU1 × 32768 × 3584
93Add_15_ffnAdd1 × 32768 × 3584
94RMSNorm_16_1RMSNorm1 × 32768 × 3584
95Attention_16Grouped Query Attn1 × 32768 × 3584
96Add_16_attnAdd1 × 32768 × 3584
97RMSNorm_16_2RMSNorm1 × 32768 × 3584
98FFN_16SwiGLU1 × 32768 × 3584
99Add_16_ffnAdd1 × 32768 × 3584
100RMSNorm_17_1RMSNorm1 × 32768 × 3584
101Attention_17Grouped Query Attn1 × 32768 × 3584
102Add_17_attnAdd1 × 32768 × 3584
103RMSNorm_17_2RMSNorm1 × 32768 × 3584
104FFN_17SwiGLU1 × 32768 × 3584
105Add_17_ffnAdd1 × 32768 × 3584
106RMSNorm_18_1RMSNorm1 × 32768 × 3584
107Attention_18Grouped Query Attn1 × 32768 × 3584
108Add_18_attnAdd1 × 32768 × 3584
109RMSNorm_18_2RMSNorm1 × 32768 × 3584
110FFN_18SwiGLU1 × 32768 × 3584
111Add_18_ffnAdd1 × 32768 × 3584
112RMSNorm_19_1RMSNorm1 × 32768 × 3584
113Attention_19Grouped Query Attn1 × 32768 × 3584
114Add_19_attnAdd1 × 32768 × 3584
115RMSNorm_19_2RMSNorm1 × 32768 × 3584
116FFN_19SwiGLU1 × 32768 × 3584
117Add_19_ffnAdd1 × 32768 × 3584
118RMSNorm_20_1RMSNorm1 × 32768 × 3584
119Attention_20Grouped Query Attn1 × 32768 × 3584
120Add_20_attnAdd1 × 32768 × 3584
121RMSNorm_20_2RMSNorm1 × 32768 × 3584
122FFN_20SwiGLU1 × 32768 × 3584
123Add_20_ffnAdd1 × 32768 × 3584
124RMSNorm_21_1RMSNorm1 × 32768 × 3584
125Attention_21Grouped Query Attn1 × 32768 × 3584
126Add_21_attnAdd1 × 32768 × 3584
127RMSNorm_21_2RMSNorm1 × 32768 × 3584
128FFN_21SwiGLU1 × 32768 × 3584
129Add_21_ffnAdd1 × 32768 × 3584
130RMSNorm_22_1RMSNorm1 × 32768 × 3584
131Attention_22Grouped Query Attn1 × 32768 × 3584
132Add_22_attnAdd1 × 32768 × 3584
133RMSNorm_22_2RMSNorm1 × 32768 × 3584
134FFN_22SwiGLU1 × 32768 × 3584
135Add_22_ffnAdd1 × 32768 × 3584
136RMSNorm_23_1RMSNorm1 × 32768 × 3584
137Attention_23Grouped Query Attn1 × 32768 × 3584
138Add_23_attnAdd1 × 32768 × 3584
139RMSNorm_23_2RMSNorm1 × 32768 × 3584
140FFN_23SwiGLU1 × 32768 × 3584
141Add_23_ffnAdd1 × 32768 × 3584
142RMSNorm_24_1RMSNorm1 × 32768 × 3584
143Attention_24Grouped Query Attn1 × 32768 × 3584
144Add_24_attnAdd1 × 32768 × 3584
145RMSNorm_24_2RMSNorm1 × 32768 × 3584
146FFN_24SwiGLU1 × 32768 × 3584
147Add_24_ffnAdd1 × 32768 × 3584
148RMSNorm_25_1RMSNorm1 × 32768 × 3584
149Attention_25Grouped Query Attn1 × 32768 × 3584
150Add_25_attnAdd1 × 32768 × 3584
151RMSNorm_25_2RMSNorm1 × 32768 × 3584
152FFN_25SwiGLU1 × 32768 × 3584
153Add_25_ffnAdd1 × 32768 × 3584
154RMSNorm_26_1RMSNorm1 × 32768 × 3584
155Attention_26Grouped Query Attn1 × 32768 × 3584
156Add_26_attnAdd1 × 32768 × 3584
157RMSNorm_26_2RMSNorm1 × 32768 × 3584
158FFN_26SwiGLU1 × 32768 × 3584
159Add_26_ffnAdd1 × 32768 × 3584
160RMSNorm_27_1RMSNorm1 × 32768 × 3584
161Attention_27Grouped Query Attn1 × 32768 × 3584
162Add_27_attnAdd1 × 32768 × 3584
163RMSNorm_27_2RMSNorm1 × 32768 × 3584
164FFN_27SwiGLU1 × 32768 × 3584
165Add_27_ffnAdd1 × 32768 × 3584
166RMSNorm_28_1RMSNorm1 × 32768 × 3584
167Attention_28Grouped Query Attn1 × 32768 × 3584
168Add_28_attnAdd1 × 32768 × 3584
169RMSNorm_28_2RMSNorm1 × 32768 × 3584
170FFN_28SwiGLU1 × 32768 × 3584
171Add_28_ffnAdd1 × 32768 × 3584
172Final_RMSNormRMSNorm1 × 32768 × 3584
173LM_HeadLinear1 × 32768 × 152064
174OutputOutput1 × 32768 × 152064

What the verifier says

infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoAt 28 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace Qwen/Qwen2.5-7B-Instruct --plan --share