N Neurarch Architectures Models Checks Data Docs Open the app

Models / siglip

siglip2-giant-opt-patch16-384

Reconstructed from its own config.json with no weights read. 2.5M downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
1.84B
1,844,534,256 parameters
In the published checkpoint
1.87B
1,871,885,426 scalars · safetensors.total, read 2026-09-06
Delta
-1.46%

multi-tower The config declares 2 sub-models (text_config, vision_config). The published checkpoint carries all of them; the graph below carries the towers the reader reconstructs. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
329
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$43.37
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

332 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 512
2EmbeddingEmbedding1 × 512 × 1152
3Positional_EmbeddingLearned Pos Embed1 × 512 × 1152
4Vision inputInput3 × 384 × 384
5PatchEmbedPatch Embed576 × 1536
6Patch_Position_EmbeddingLearned Pos Embed576 × 1536
7Vision_LN_1LayerNorm576 × 1536
8Vision_Attn_1Multi-Head Attention576 × 1536
9Vision_Add_1Add576 × 1536
10Vision_FFN_1Feed Forward576 × 1536
11Vision_LN_2LayerNorm576 × 1536
12Vision_Attn_2Multi-Head Attention576 × 1536
13Vision_Add_2Add576 × 1536
14Vision_FFN_2Feed Forward576 × 1536
15Vision_LN_3LayerNorm576 × 1536
16Vision_Attn_3Multi-Head Attention576 × 1536
17Vision_Add_3Add576 × 1536
18Vision_FFN_3Feed Forward576 × 1536
19Vision_LN_4LayerNorm576 × 1536
20Vision_Attn_4Multi-Head Attention576 × 1536
21Vision_Add_4Add576 × 1536
22Vision_FFN_4Feed Forward576 × 1536
23Vision_LN_5LayerNorm576 × 1536
24Vision_Attn_5Multi-Head Attention576 × 1536
25Vision_Add_5Add576 × 1536
26Vision_FFN_5Feed Forward576 × 1536
27Vision_LN_6LayerNorm576 × 1536
28Vision_Attn_6Multi-Head Attention576 × 1536
29Vision_Add_6Add576 × 1536
30Vision_FFN_6Feed Forward576 × 1536
31Vision_LN_7LayerNorm576 × 1536
32Vision_Attn_7Multi-Head Attention576 × 1536
33Vision_Add_7Add576 × 1536
34Vision_FFN_7Feed Forward576 × 1536
35Vision_LN_8LayerNorm576 × 1536
36Vision_Attn_8Multi-Head Attention576 × 1536
37Vision_Add_8Add576 × 1536
38Vision_FFN_8Feed Forward576 × 1536
39Vision_LN_9LayerNorm576 × 1536
40Vision_Attn_9Multi-Head Attention576 × 1536
41Vision_Add_9Add576 × 1536
42Vision_FFN_9Feed Forward576 × 1536
43Vision_LN_10LayerNorm576 × 1536
44Vision_Attn_10Multi-Head Attention576 × 1536
45Vision_Add_10Add576 × 1536
46Vision_FFN_10Feed Forward576 × 1536
47Vision_LN_11LayerNorm576 × 1536
48Vision_Attn_11Multi-Head Attention576 × 1536
49Vision_Add_11Add576 × 1536
50Vision_FFN_11Feed Forward576 × 1536
51Vision_LN_12LayerNorm576 × 1536
52Vision_Attn_12Multi-Head Attention576 × 1536
53Vision_Add_12Add576 × 1536
54Vision_FFN_12Feed Forward576 × 1536
55Vision_LN_13LayerNorm576 × 1536
56Vision_Attn_13Multi-Head Attention576 × 1536
57Vision_Add_13Add576 × 1536
58Vision_FFN_13Feed Forward576 × 1536
59Vision_LN_14LayerNorm576 × 1536
60Vision_Attn_14Multi-Head Attention576 × 1536
61Vision_Add_14Add576 × 1536
62Vision_FFN_14Feed Forward576 × 1536
63Vision_LN_15LayerNorm576 × 1536
64Vision_Attn_15Multi-Head Attention576 × 1536
65Vision_Add_15Add576 × 1536
66Vision_FFN_15Feed Forward576 × 1536
67Vision_LN_16LayerNorm576 × 1536
68Vision_Attn_16Multi-Head Attention576 × 1536
69Vision_Add_16Add576 × 1536
70Vision_FFN_16Feed Forward576 × 1536
71Vision_LN_17LayerNorm576 × 1536
72Vision_Attn_17Multi-Head Attention576 × 1536
73Vision_Add_17Add576 × 1536
74Vision_FFN_17Feed Forward576 × 1536
75Vision_LN_18LayerNorm576 × 1536
76Vision_Attn_18Multi-Head Attention576 × 1536
77Vision_Add_18Add576 × 1536
78Vision_FFN_18Feed Forward576 × 1536
79Vision_LN_19LayerNorm576 × 1536
80Vision_Attn_19Multi-Head Attention576 × 1536
81Vision_Add_19Add576 × 1536
82Vision_FFN_19Feed Forward576 × 1536
83Vision_LN_20LayerNorm576 × 1536
84Vision_Attn_20Multi-Head Attention576 × 1536
85Vision_Add_20Add576 × 1536
86Vision_FFN_20Feed Forward576 × 1536
87Vision_LN_21LayerNorm576 × 1536
88Vision_Attn_21Multi-Head Attention576 × 1536
89Vision_Add_21Add576 × 1536
90Vision_FFN_21Feed Forward576 × 1536
91Vision_LN_22LayerNorm576 × 1536
92Vision_Attn_22Multi-Head Attention576 × 1536
93Vision_Add_22Add576 × 1536
94Vision_FFN_22Feed Forward576 × 1536
95Vision_LN_23LayerNorm576 × 1536
96Vision_Attn_23Multi-Head Attention576 × 1536
97Vision_Add_23Add576 × 1536
98Vision_FFN_23Feed Forward576 × 1536
99Vision_LN_24LayerNorm576 × 1536
100Vision_Attn_24Multi-Head Attention576 × 1536
101Vision_Add_24Add576 × 1536
102Vision_FFN_24Feed Forward576 × 1536
103Vision_LN_25LayerNorm576 × 1536
104Vision_Attn_25Multi-Head Attention576 × 1536
105Vision_Add_25Add576 × 1536
106Vision_FFN_25Feed Forward576 × 1536
107Vision_LN_26LayerNorm576 × 1536
108Vision_Attn_26Multi-Head Attention576 × 1536
109Vision_Add_26Add576 × 1536
110Vision_FFN_26Feed Forward576 × 1536
111Vision_LN_27LayerNorm576 × 1536
112Vision_Attn_27Multi-Head Attention576 × 1536
113Vision_Add_27Add576 × 1536
114Vision_FFN_27Feed Forward576 × 1536
115Vision_LN_28LayerNorm576 × 1536
116Vision_Attn_28Multi-Head Attention576 × 1536
117Vision_Add_28Add576 × 1536
118Vision_FFN_28Feed Forward576 × 1536
119Vision_LN_29LayerNorm576 × 1536
120Vision_Attn_29Multi-Head Attention576 × 1536
121Vision_Add_29Add576 × 1536
122Vision_FFN_29Feed Forward576 × 1536
123Vision_LN_30LayerNorm576 × 1536
124Vision_Attn_30Multi-Head Attention576 × 1536
125Vision_Add_30Add576 × 1536
126Vision_FFN_30Feed Forward576 × 1536
127Vision_LN_31LayerNorm576 × 1536
128Vision_Attn_31Multi-Head Attention576 × 1536
129Vision_Add_31Add576 × 1536
130Vision_FFN_31Feed Forward576 × 1536
131Vision_LN_32LayerNorm576 × 1536
132Vision_Attn_32Multi-Head Attention576 × 1536
133Vision_Add_32Add576 × 1536
134Vision_FFN_32Feed Forward576 × 1536
135Vision_LN_33LayerNorm576 × 1536
136Vision_Attn_33Multi-Head Attention576 × 1536
137Vision_Add_33Add576 × 1536
138Vision_FFN_33Feed Forward576 × 1536
139Vision_LN_34LayerNorm576 × 1536
140Vision_Attn_34Multi-Head Attention576 × 1536
141Vision_Add_34Add576 × 1536
142Vision_FFN_34Feed Forward576 × 1536
143Vision_LN_35LayerNorm576 × 1536
144Vision_Attn_35Multi-Head Attention576 × 1536
145Vision_Add_35Add576 × 1536
146Vision_FFN_35Feed Forward576 × 1536
147Vision_LN_36LayerNorm576 × 1536
148Vision_Attn_36Multi-Head Attention576 × 1536
149Vision_Add_36Add576 × 1536
150Vision_FFN_36Feed Forward576 × 1536
151Vision_LN_37LayerNorm576 × 1536
152Vision_Attn_37Multi-Head Attention576 × 1536
153Vision_Add_37Add576 × 1536
154Vision_FFN_37Feed Forward576 × 1536
155Vision_LN_38LayerNorm576 × 1536
156Vision_Attn_38Multi-Head Attention576 × 1536
157Vision_Add_38Add576 × 1536
158Vision_FFN_38Feed Forward576 × 1536
159Vision_LN_39LayerNorm576 × 1536
160Vision_Attn_39Multi-Head Attention576 × 1536
161Vision_Add_39Add576 × 1536
162Vision_FFN_39Feed Forward576 × 1536
163Vision_LN_40LayerNorm576 × 1536
164Vision_Attn_40Multi-Head Attention576 × 1536
165Vision_Add_40Add576 × 1536
166Vision_FFN_40Feed Forward576 × 1536
167Vision projectorProjection576 × 1152
168Vision tokensReshape1 × 576 × 1152
169Multimodal fusion (concat tokens)Concatenate1 × 1088 × 1152
170LayerNorm_1_1LayerNorm1 × 1088 × 1152
171Attention_1Multi-Head Attention1 × 1088 × 1152
172Add_1_attnAdd1 × 1088 × 1152
173LayerNorm_1_2LayerNorm1 × 1088 × 1152
174FFN_1Feed Forward1 × 1088 × 1152
175Add_1_ffnAdd1 × 1088 × 1152
176LayerNorm_2_1LayerNorm1 × 1088 × 1152
177Attention_2Multi-Head Attention1 × 1088 × 1152
178Add_2_attnAdd1 × 1088 × 1152
179LayerNorm_2_2LayerNorm1 × 1088 × 1152
180FFN_2Feed Forward1 × 1088 × 1152
181Add_2_ffnAdd1 × 1088 × 1152
182LayerNorm_3_1LayerNorm1 × 1088 × 1152
183Attention_3Multi-Head Attention1 × 1088 × 1152
184Add_3_attnAdd1 × 1088 × 1152
185LayerNorm_3_2LayerNorm1 × 1088 × 1152
186FFN_3Feed Forward1 × 1088 × 1152
187Add_3_ffnAdd1 × 1088 × 1152
188LayerNorm_4_1LayerNorm1 × 1088 × 1152
189Attention_4Multi-Head Attention1 × 1088 × 1152
190Add_4_attnAdd1 × 1088 × 1152
191LayerNorm_4_2LayerNorm1 × 1088 × 1152
192FFN_4Feed Forward1 × 1088 × 1152
193Add_4_ffnAdd1 × 1088 × 1152
194LayerNorm_5_1LayerNorm1 × 1088 × 1152
195Attention_5Multi-Head Attention1 × 1088 × 1152
196Add_5_attnAdd1 × 1088 × 1152
197LayerNorm_5_2LayerNorm1 × 1088 × 1152
198FFN_5Feed Forward1 × 1088 × 1152
199Add_5_ffnAdd1 × 1088 × 1152
200LayerNorm_6_1LayerNorm1 × 1088 × 1152
201Attention_6Multi-Head Attention1 × 1088 × 1152
202Add_6_attnAdd1 × 1088 × 1152
203LayerNorm_6_2LayerNorm1 × 1088 × 1152
204FFN_6Feed Forward1 × 1088 × 1152
205Add_6_ffnAdd1 × 1088 × 1152
206LayerNorm_7_1LayerNorm1 × 1088 × 1152
207Attention_7Multi-Head Attention1 × 1088 × 1152
208Add_7_attnAdd1 × 1088 × 1152
209LayerNorm_7_2LayerNorm1 × 1088 × 1152
210FFN_7Feed Forward1 × 1088 × 1152
211Add_7_ffnAdd1 × 1088 × 1152
212LayerNorm_8_1LayerNorm1 × 1088 × 1152
213Attention_8Multi-Head Attention1 × 1088 × 1152
214Add_8_attnAdd1 × 1088 × 1152
215LayerNorm_8_2LayerNorm1 × 1088 × 1152
216FFN_8Feed Forward1 × 1088 × 1152
217Add_8_ffnAdd1 × 1088 × 1152
218LayerNorm_9_1LayerNorm1 × 1088 × 1152
219Attention_9Multi-Head Attention1 × 1088 × 1152
220Add_9_attnAdd1 × 1088 × 1152
221LayerNorm_9_2LayerNorm1 × 1088 × 1152
222FFN_9Feed Forward1 × 1088 × 1152
223Add_9_ffnAdd1 × 1088 × 1152
224LayerNorm_10_1LayerNorm1 × 1088 × 1152
225Attention_10Multi-Head Attention1 × 1088 × 1152
226Add_10_attnAdd1 × 1088 × 1152
227LayerNorm_10_2LayerNorm1 × 1088 × 1152
228FFN_10Feed Forward1 × 1088 × 1152
229Add_10_ffnAdd1 × 1088 × 1152
230LayerNorm_11_1LayerNorm1 × 1088 × 1152
231Attention_11Multi-Head Attention1 × 1088 × 1152
232Add_11_attnAdd1 × 1088 × 1152
233LayerNorm_11_2LayerNorm1 × 1088 × 1152
234FFN_11Feed Forward1 × 1088 × 1152
235Add_11_ffnAdd1 × 1088 × 1152
236LayerNorm_12_1LayerNorm1 × 1088 × 1152
237Attention_12Multi-Head Attention1 × 1088 × 1152
238Add_12_attnAdd1 × 1088 × 1152
239LayerNorm_12_2LayerNorm1 × 1088 × 1152
240FFN_12Feed Forward1 × 1088 × 1152
241Add_12_ffnAdd1 × 1088 × 1152
242LayerNorm_13_1LayerNorm1 × 1088 × 1152
243Attention_13Multi-Head Attention1 × 1088 × 1152
244Add_13_attnAdd1 × 1088 × 1152
245LayerNorm_13_2LayerNorm1 × 1088 × 1152
246FFN_13Feed Forward1 × 1088 × 1152
247Add_13_ffnAdd1 × 1088 × 1152
248LayerNorm_14_1LayerNorm1 × 1088 × 1152
249Attention_14Multi-Head Attention1 × 1088 × 1152
250Add_14_attnAdd1 × 1088 × 1152
251LayerNorm_14_2LayerNorm1 × 1088 × 1152
252FFN_14Feed Forward1 × 1088 × 1152
253Add_14_ffnAdd1 × 1088 × 1152
254LayerNorm_15_1LayerNorm1 × 1088 × 1152
255Attention_15Multi-Head Attention1 × 1088 × 1152
256Add_15_attnAdd1 × 1088 × 1152
257LayerNorm_15_2LayerNorm1 × 1088 × 1152
258FFN_15Feed Forward1 × 1088 × 1152
259Add_15_ffnAdd1 × 1088 × 1152
260LayerNorm_16_1LayerNorm1 × 1088 × 1152
261Attention_16Multi-Head Attention1 × 1088 × 1152
262Add_16_attnAdd1 × 1088 × 1152
263LayerNorm_16_2LayerNorm1 × 1088 × 1152
264FFN_16Feed Forward1 × 1088 × 1152
265Add_16_ffnAdd1 × 1088 × 1152
266LayerNorm_17_1LayerNorm1 × 1088 × 1152
267Attention_17Multi-Head Attention1 × 1088 × 1152
268Add_17_attnAdd1 × 1088 × 1152
269LayerNorm_17_2LayerNorm1 × 1088 × 1152
270FFN_17Feed Forward1 × 1088 × 1152
271Add_17_ffnAdd1 × 1088 × 1152
272LayerNorm_18_1LayerNorm1 × 1088 × 1152
273Attention_18Multi-Head Attention1 × 1088 × 1152
274Add_18_attnAdd1 × 1088 × 1152
275LayerNorm_18_2LayerNorm1 × 1088 × 1152
276FFN_18Feed Forward1 × 1088 × 1152
277Add_18_ffnAdd1 × 1088 × 1152
278LayerNorm_19_1LayerNorm1 × 1088 × 1152
279Attention_19Multi-Head Attention1 × 1088 × 1152
280Add_19_attnAdd1 × 1088 × 1152
281LayerNorm_19_2LayerNorm1 × 1088 × 1152
282FFN_19Feed Forward1 × 1088 × 1152
283Add_19_ffnAdd1 × 1088 × 1152
284LayerNorm_20_1LayerNorm1 × 1088 × 1152
285Attention_20Multi-Head Attention1 × 1088 × 1152
286Add_20_attnAdd1 × 1088 × 1152
287LayerNorm_20_2LayerNorm1 × 1088 × 1152
288FFN_20Feed Forward1 × 1088 × 1152
289Add_20_ffnAdd1 × 1088 × 1152
290LayerNorm_21_1LayerNorm1 × 1088 × 1152
291Attention_21Multi-Head Attention1 × 1088 × 1152
292Add_21_attnAdd1 × 1088 × 1152
293LayerNorm_21_2LayerNorm1 × 1088 × 1152
294FFN_21Feed Forward1 × 1088 × 1152
295Add_21_ffnAdd1 × 1088 × 1152
296LayerNorm_22_1LayerNorm1 × 1088 × 1152
297Attention_22Multi-Head Attention1 × 1088 × 1152
298Add_22_attnAdd1 × 1088 × 1152
299LayerNorm_22_2LayerNorm1 × 1088 × 1152
300FFN_22Feed Forward1 × 1088 × 1152
301Add_22_ffnAdd1 × 1088 × 1152
302LayerNorm_23_1LayerNorm1 × 1088 × 1152
303Attention_23Multi-Head Attention1 × 1088 × 1152
304Add_23_attnAdd1 × 1088 × 1152
305LayerNorm_23_2LayerNorm1 × 1088 × 1152
306FFN_23Feed Forward1 × 1088 × 1152
307Add_23_ffnAdd1 × 1088 × 1152
308LayerNorm_24_1LayerNorm1 × 1088 × 1152
309Attention_24Multi-Head Attention1 × 1088 × 1152
310Add_24_attnAdd1 × 1088 × 1152
311LayerNorm_24_2LayerNorm1 × 1088 × 1152
312FFN_24Feed Forward1 × 1088 × 1152
313Add_24_ffnAdd1 × 1088 × 1152
314LayerNorm_25_1LayerNorm1 × 1088 × 1152
315Attention_25Multi-Head Attention1 × 1088 × 1152
316Add_25_attnAdd1 × 1088 × 1152
317LayerNorm_25_2LayerNorm1 × 1088 × 1152
318FFN_25Feed Forward1 × 1088 × 1152
319Add_25_ffnAdd1 × 1088 × 1152
320LayerNorm_26_1LayerNorm1 × 1088 × 1152
321Attention_26Multi-Head Attention1 × 1088 × 1152
322Add_26_attnAdd1 × 1088 × 1152
323LayerNorm_26_2LayerNorm1 × 1088 × 1152
324FFN_26Feed Forward1 × 1088 × 1152
325Add_26_ffnAdd1 × 1088 × 1152
326LayerNorm_27_1LayerNorm1 × 1088 × 1152
327Attention_27Multi-Head Attention1 × 1088 × 1152
328Add_27_attnAdd1 × 1088 × 1152
329LayerNorm_27_2LayerNorm1 × 1088 × 1152
330FFN_27Feed Forward1 × 1088 × 1152
331Add_27_ffnAdd1 × 1088 × 1152
332OutputOutput1 × 1088 × 1152

What the verifier says

warn"Positional_Embedding" receives input but its output is not connected. This layer will be unreachable in the forward pass. Fix: Connect the output forward, or add an Output node if this is the final layer.
dead-end
infoAt 67 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace google/siglip2-giant-opt-patch16-384 --plan --share