N Neurarch Architectures Models Checks Data Docs Open the app

Models / internvl_chat

InternVL3-8B

Reconstructed from its own config.json with no weights read. 103K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
7.92B
7,920,430,202 parameters
In the published checkpoint
7.94B
7,944,373,760 scalars · safetensors.total, read 2025-09-11
Delta
-0.30%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `configuration_internvl_chat.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
273
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$13242.86
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsdoes not fit

Structure

276 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 32768
2EmbeddingEmbedding1 × 32768 × 3584
3RoPERoPE1 × 32768 × 3584
4Vision inputInput3 × 448 × 448
5PatchEmbedPatch Embed1024 × 1024
6Patch_Position_EmbeddingLearned Pos Embed1024 × 1024
7Vision_LN_1LayerNorm1024 × 1024
8Vision_Attn_1Multi-Head Attention1024 × 1024
9Vision_Add_1Add1024 × 1024
10Vision_FFN_1Feed Forward1024 × 1024
11Vision_LN_2LayerNorm1024 × 1024
12Vision_Attn_2Multi-Head Attention1024 × 1024
13Vision_Add_2Add1024 × 1024
14Vision_FFN_2Feed Forward1024 × 1024
15Vision_LN_3LayerNorm1024 × 1024
16Vision_Attn_3Multi-Head Attention1024 × 1024
17Vision_Add_3Add1024 × 1024
18Vision_FFN_3Feed Forward1024 × 1024
19Vision_LN_4LayerNorm1024 × 1024
20Vision_Attn_4Multi-Head Attention1024 × 1024
21Vision_Add_4Add1024 × 1024
22Vision_FFN_4Feed Forward1024 × 1024
23Vision_LN_5LayerNorm1024 × 1024
24Vision_Attn_5Multi-Head Attention1024 × 1024
25Vision_Add_5Add1024 × 1024
26Vision_FFN_5Feed Forward1024 × 1024
27Vision_LN_6LayerNorm1024 × 1024
28Vision_Attn_6Multi-Head Attention1024 × 1024
29Vision_Add_6Add1024 × 1024
30Vision_FFN_6Feed Forward1024 × 1024
31Vision_LN_7LayerNorm1024 × 1024
32Vision_Attn_7Multi-Head Attention1024 × 1024
33Vision_Add_7Add1024 × 1024
34Vision_FFN_7Feed Forward1024 × 1024
35Vision_LN_8LayerNorm1024 × 1024
36Vision_Attn_8Multi-Head Attention1024 × 1024
37Vision_Add_8Add1024 × 1024
38Vision_FFN_8Feed Forward1024 × 1024
39Vision_LN_9LayerNorm1024 × 1024
40Vision_Attn_9Multi-Head Attention1024 × 1024
41Vision_Add_9Add1024 × 1024
42Vision_FFN_9Feed Forward1024 × 1024
43Vision_LN_10LayerNorm1024 × 1024
44Vision_Attn_10Multi-Head Attention1024 × 1024
45Vision_Add_10Add1024 × 1024
46Vision_FFN_10Feed Forward1024 × 1024
47Vision_LN_11LayerNorm1024 × 1024
48Vision_Attn_11Multi-Head Attention1024 × 1024
49Vision_Add_11Add1024 × 1024
50Vision_FFN_11Feed Forward1024 × 1024
51Vision_LN_12LayerNorm1024 × 1024
52Vision_Attn_12Multi-Head Attention1024 × 1024
53Vision_Add_12Add1024 × 1024
54Vision_FFN_12Feed Forward1024 × 1024
55Vision_LN_13LayerNorm1024 × 1024
56Vision_Attn_13Multi-Head Attention1024 × 1024
57Vision_Add_13Add1024 × 1024
58Vision_FFN_13Feed Forward1024 × 1024
59Vision_LN_14LayerNorm1024 × 1024
60Vision_Attn_14Multi-Head Attention1024 × 1024
61Vision_Add_14Add1024 × 1024
62Vision_FFN_14Feed Forward1024 × 1024
63Vision_LN_15LayerNorm1024 × 1024
64Vision_Attn_15Multi-Head Attention1024 × 1024
65Vision_Add_15Add1024 × 1024
66Vision_FFN_15Feed Forward1024 × 1024
67Vision_LN_16LayerNorm1024 × 1024
68Vision_Attn_16Multi-Head Attention1024 × 1024
69Vision_Add_16Add1024 × 1024
70Vision_FFN_16Feed Forward1024 × 1024
71Vision_LN_17LayerNorm1024 × 1024
72Vision_Attn_17Multi-Head Attention1024 × 1024
73Vision_Add_17Add1024 × 1024
74Vision_FFN_17Feed Forward1024 × 1024
75Vision_LN_18LayerNorm1024 × 1024
76Vision_Attn_18Multi-Head Attention1024 × 1024
77Vision_Add_18Add1024 × 1024
78Vision_FFN_18Feed Forward1024 × 1024
79Vision_LN_19LayerNorm1024 × 1024
80Vision_Attn_19Multi-Head Attention1024 × 1024
81Vision_Add_19Add1024 × 1024
82Vision_FFN_19Feed Forward1024 × 1024
83Vision_LN_20LayerNorm1024 × 1024
84Vision_Attn_20Multi-Head Attention1024 × 1024
85Vision_Add_20Add1024 × 1024
86Vision_FFN_20Feed Forward1024 × 1024
87Vision_LN_21LayerNorm1024 × 1024
88Vision_Attn_21Multi-Head Attention1024 × 1024
89Vision_Add_21Add1024 × 1024
90Vision_FFN_21Feed Forward1024 × 1024
91Vision_LN_22LayerNorm1024 × 1024
92Vision_Attn_22Multi-Head Attention1024 × 1024
93Vision_Add_22Add1024 × 1024
94Vision_FFN_22Feed Forward1024 × 1024
95Vision_LN_23LayerNorm1024 × 1024
96Vision_Attn_23Multi-Head Attention1024 × 1024
97Vision_Add_23Add1024 × 1024
98Vision_FFN_23Feed Forward1024 × 1024
99Vision_LN_24LayerNorm1024 × 1024
100Vision_Attn_24Multi-Head Attention1024 × 1024
101Vision_Add_24Add1024 × 1024
102Vision_FFN_24Feed Forward1024 × 1024
103Vision projectorProjection1024 × 3584
104Vision tokensReshape1 × 1024 × 3584
105Multimodal fusion (concat tokens)Concatenate1 × 33792 × 3584
106RMSNorm_1_1RMSNorm1 × 33792 × 3584
107Attention_1Grouped Query Attn1 × 33792 × 3584
108Add_1_attnAdd1 × 33792 × 3584
109RMSNorm_1_2RMSNorm1 × 33792 × 3584
110FFN_1SwiGLU1 × 33792 × 3584
111Add_1_ffnAdd1 × 33792 × 3584
112RMSNorm_2_1RMSNorm1 × 33792 × 3584
113Attention_2Grouped Query Attn1 × 33792 × 3584
114Add_2_attnAdd1 × 33792 × 3584
115RMSNorm_2_2RMSNorm1 × 33792 × 3584
116FFN_2SwiGLU1 × 33792 × 3584
117Add_2_ffnAdd1 × 33792 × 3584
118RMSNorm_3_1RMSNorm1 × 33792 × 3584
119Attention_3Grouped Query Attn1 × 33792 × 3584
120Add_3_attnAdd1 × 33792 × 3584
121RMSNorm_3_2RMSNorm1 × 33792 × 3584
122FFN_3SwiGLU1 × 33792 × 3584
123Add_3_ffnAdd1 × 33792 × 3584
124RMSNorm_4_1RMSNorm1 × 33792 × 3584
125Attention_4Grouped Query Attn1 × 33792 × 3584
126Add_4_attnAdd1 × 33792 × 3584
127RMSNorm_4_2RMSNorm1 × 33792 × 3584
128FFN_4SwiGLU1 × 33792 × 3584
129Add_4_ffnAdd1 × 33792 × 3584
130RMSNorm_5_1RMSNorm1 × 33792 × 3584
131Attention_5Grouped Query Attn1 × 33792 × 3584
132Add_5_attnAdd1 × 33792 × 3584
133RMSNorm_5_2RMSNorm1 × 33792 × 3584
134FFN_5SwiGLU1 × 33792 × 3584
135Add_5_ffnAdd1 × 33792 × 3584
136RMSNorm_6_1RMSNorm1 × 33792 × 3584
137Attention_6Grouped Query Attn1 × 33792 × 3584
138Add_6_attnAdd1 × 33792 × 3584
139RMSNorm_6_2RMSNorm1 × 33792 × 3584
140FFN_6SwiGLU1 × 33792 × 3584
141Add_6_ffnAdd1 × 33792 × 3584
142RMSNorm_7_1RMSNorm1 × 33792 × 3584
143Attention_7Grouped Query Attn1 × 33792 × 3584
144Add_7_attnAdd1 × 33792 × 3584
145RMSNorm_7_2RMSNorm1 × 33792 × 3584
146FFN_7SwiGLU1 × 33792 × 3584
147Add_7_ffnAdd1 × 33792 × 3584
148RMSNorm_8_1RMSNorm1 × 33792 × 3584
149Attention_8Grouped Query Attn1 × 33792 × 3584
150Add_8_attnAdd1 × 33792 × 3584
151RMSNorm_8_2RMSNorm1 × 33792 × 3584
152FFN_8SwiGLU1 × 33792 × 3584
153Add_8_ffnAdd1 × 33792 × 3584
154RMSNorm_9_1RMSNorm1 × 33792 × 3584
155Attention_9Grouped Query Attn1 × 33792 × 3584
156Add_9_attnAdd1 × 33792 × 3584
157RMSNorm_9_2RMSNorm1 × 33792 × 3584
158FFN_9SwiGLU1 × 33792 × 3584
159Add_9_ffnAdd1 × 33792 × 3584
160RMSNorm_10_1RMSNorm1 × 33792 × 3584
161Attention_10Grouped Query Attn1 × 33792 × 3584
162Add_10_attnAdd1 × 33792 × 3584
163RMSNorm_10_2RMSNorm1 × 33792 × 3584
164FFN_10SwiGLU1 × 33792 × 3584
165Add_10_ffnAdd1 × 33792 × 3584
166RMSNorm_11_1RMSNorm1 × 33792 × 3584
167Attention_11Grouped Query Attn1 × 33792 × 3584
168Add_11_attnAdd1 × 33792 × 3584
169RMSNorm_11_2RMSNorm1 × 33792 × 3584
170FFN_11SwiGLU1 × 33792 × 3584
171Add_11_ffnAdd1 × 33792 × 3584
172RMSNorm_12_1RMSNorm1 × 33792 × 3584
173Attention_12Grouped Query Attn1 × 33792 × 3584
174Add_12_attnAdd1 × 33792 × 3584
175RMSNorm_12_2RMSNorm1 × 33792 × 3584
176FFN_12SwiGLU1 × 33792 × 3584
177Add_12_ffnAdd1 × 33792 × 3584
178RMSNorm_13_1RMSNorm1 × 33792 × 3584
179Attention_13Grouped Query Attn1 × 33792 × 3584
180Add_13_attnAdd1 × 33792 × 3584
181RMSNorm_13_2RMSNorm1 × 33792 × 3584
182FFN_13SwiGLU1 × 33792 × 3584
183Add_13_ffnAdd1 × 33792 × 3584
184RMSNorm_14_1RMSNorm1 × 33792 × 3584
185Attention_14Grouped Query Attn1 × 33792 × 3584
186Add_14_attnAdd1 × 33792 × 3584
187RMSNorm_14_2RMSNorm1 × 33792 × 3584
188FFN_14SwiGLU1 × 33792 × 3584
189Add_14_ffnAdd1 × 33792 × 3584
190RMSNorm_15_1RMSNorm1 × 33792 × 3584
191Attention_15Grouped Query Attn1 × 33792 × 3584
192Add_15_attnAdd1 × 33792 × 3584
193RMSNorm_15_2RMSNorm1 × 33792 × 3584
194FFN_15SwiGLU1 × 33792 × 3584
195Add_15_ffnAdd1 × 33792 × 3584
196RMSNorm_16_1RMSNorm1 × 33792 × 3584
197Attention_16Grouped Query Attn1 × 33792 × 3584
198Add_16_attnAdd1 × 33792 × 3584
199RMSNorm_16_2RMSNorm1 × 33792 × 3584
200FFN_16SwiGLU1 × 33792 × 3584
201Add_16_ffnAdd1 × 33792 × 3584
202RMSNorm_17_1RMSNorm1 × 33792 × 3584
203Attention_17Grouped Query Attn1 × 33792 × 3584
204Add_17_attnAdd1 × 33792 × 3584
205RMSNorm_17_2RMSNorm1 × 33792 × 3584
206FFN_17SwiGLU1 × 33792 × 3584
207Add_17_ffnAdd1 × 33792 × 3584
208RMSNorm_18_1RMSNorm1 × 33792 × 3584
209Attention_18Grouped Query Attn1 × 33792 × 3584
210Add_18_attnAdd1 × 33792 × 3584
211RMSNorm_18_2RMSNorm1 × 33792 × 3584
212FFN_18SwiGLU1 × 33792 × 3584
213Add_18_ffnAdd1 × 33792 × 3584
214RMSNorm_19_1RMSNorm1 × 33792 × 3584
215Attention_19Grouped Query Attn1 × 33792 × 3584
216Add_19_attnAdd1 × 33792 × 3584
217RMSNorm_19_2RMSNorm1 × 33792 × 3584
218FFN_19SwiGLU1 × 33792 × 3584
219Add_19_ffnAdd1 × 33792 × 3584
220RMSNorm_20_1RMSNorm1 × 33792 × 3584
221Attention_20Grouped Query Attn1 × 33792 × 3584
222Add_20_attnAdd1 × 33792 × 3584
223RMSNorm_20_2RMSNorm1 × 33792 × 3584
224FFN_20SwiGLU1 × 33792 × 3584
225Add_20_ffnAdd1 × 33792 × 3584
226RMSNorm_21_1RMSNorm1 × 33792 × 3584
227Attention_21Grouped Query Attn1 × 33792 × 3584
228Add_21_attnAdd1 × 33792 × 3584
229RMSNorm_21_2RMSNorm1 × 33792 × 3584
230FFN_21SwiGLU1 × 33792 × 3584
231Add_21_ffnAdd1 × 33792 × 3584
232RMSNorm_22_1RMSNorm1 × 33792 × 3584
233Attention_22Grouped Query Attn1 × 33792 × 3584
234Add_22_attnAdd1 × 33792 × 3584
235RMSNorm_22_2RMSNorm1 × 33792 × 3584
236FFN_22SwiGLU1 × 33792 × 3584
237Add_22_ffnAdd1 × 33792 × 3584
238RMSNorm_23_1RMSNorm1 × 33792 × 3584
239Attention_23Grouped Query Attn1 × 33792 × 3584
240Add_23_attnAdd1 × 33792 × 3584
241RMSNorm_23_2RMSNorm1 × 33792 × 3584
242FFN_23SwiGLU1 × 33792 × 3584
243Add_23_ffnAdd1 × 33792 × 3584
244RMSNorm_24_1RMSNorm1 × 33792 × 3584
245Attention_24Grouped Query Attn1 × 33792 × 3584
246Add_24_attnAdd1 × 33792 × 3584
247RMSNorm_24_2RMSNorm1 × 33792 × 3584
248FFN_24SwiGLU1 × 33792 × 3584
249Add_24_ffnAdd1 × 33792 × 3584
250RMSNorm_25_1RMSNorm1 × 33792 × 3584
251Attention_25Grouped Query Attn1 × 33792 × 3584
252Add_25_attnAdd1 × 33792 × 3584
253RMSNorm_25_2RMSNorm1 × 33792 × 3584
254FFN_25SwiGLU1 × 33792 × 3584
255Add_25_ffnAdd1 × 33792 × 3584
256RMSNorm_26_1RMSNorm1 × 33792 × 3584
257Attention_26Grouped Query Attn1 × 33792 × 3584
258Add_26_attnAdd1 × 33792 × 3584
259RMSNorm_26_2RMSNorm1 × 33792 × 3584
260FFN_26SwiGLU1 × 33792 × 3584
261Add_26_ffnAdd1 × 33792 × 3584
262RMSNorm_27_1RMSNorm1 × 33792 × 3584
263Attention_27Grouped Query Attn1 × 33792 × 3584
264Add_27_attnAdd1 × 33792 × 3584
265RMSNorm_27_2RMSNorm1 × 33792 × 3584
266FFN_27SwiGLU1 × 33792 × 3584
267Add_27_ffnAdd1 × 33792 × 3584
268RMSNorm_28_1RMSNorm1 × 33792 × 3584
269Attention_28Grouped Query Attn1 × 33792 × 3584
270Add_28_attnAdd1 × 33792 × 3584
271RMSNorm_28_2RMSNorm1 × 33792 × 3584
272FFN_28SwiGLU1 × 33792 × 3584
273Add_28_ffnAdd1 × 33792 × 3584
274Final_RMSNormRMSNorm1 × 33792 × 3584
275LM_HeadLinear1 × 33792 × 151674
276OutputOutput1 × 33792 × 151674

What the verifier says

warn"RoPE" receives input but its output is not connected. This layer will be unreachable in the forward pass. Fix: Connect the output forward, or add an Output node if this is the final layer.
dead-end
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 18944 (5.29× embedDim). Expected: ~9472. Fix: Set intermediateSize to 9472 for embedDim=3584.
swiglu-dim-convention
infoAt 52 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace OpenGVLab/InternVL3-8B --plan --share

Other internvl_chat checkpoints

InternVL2-1B
935M derived · -0.38% against the checkpoint
InternVL2-26B
25.42B derived · -0.38% against the checkpoint
InternVL2-2B
2.20B derived · -0.48% against the checkpoint
InternVL2_5-4B
3.70B derived · -0.28% against the checkpoint