N Neurarch Architectures Models Checks Data Docs Open the app

Models / qwen2_vl

jina-reranker-m0

Reconstructed from its own config.json with no weights read. 383K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
1.55B
1,547,500,032 parameters
In the published checkpoint
2.44B
2,444,721,665 scalars · safetensors.total, read 2026-04-09
Delta
-36.7%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `modeling.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
175
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$3366.11
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

178 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 32768
2EmbeddingEmbedding1 × 32768 × 1536
3RoPERoPE1 × 32768 × 1536
4Vision inputInput3 × 224 × 224
5PatchEmbedPatch Embed196 × 1536
6Patch_Position_EmbeddingLearned Pos Embed196 × 1536
7Vision encoder (internals not in config)Projection196 × 1536
8Vision tokensReshape1 × 196 × 1536
9Multimodal fusion (concat tokens)Concatenate1 × 32964 × 1536
10RMSNorm_1_1RMSNorm1 × 32964 × 1536
11Attention_1Grouped Query Attn1 × 32964 × 1536
12Add_1_attnAdd1 × 32964 × 1536
13RMSNorm_1_2RMSNorm1 × 32964 × 1536
14FFN_1SwiGLU1 × 32964 × 1536
15Add_1_ffnAdd1 × 32964 × 1536
16RMSNorm_2_1RMSNorm1 × 32964 × 1536
17Attention_2Grouped Query Attn1 × 32964 × 1536
18Add_2_attnAdd1 × 32964 × 1536
19RMSNorm_2_2RMSNorm1 × 32964 × 1536
20FFN_2SwiGLU1 × 32964 × 1536
21Add_2_ffnAdd1 × 32964 × 1536
22RMSNorm_3_1RMSNorm1 × 32964 × 1536
23Attention_3Grouped Query Attn1 × 32964 × 1536
24Add_3_attnAdd1 × 32964 × 1536
25RMSNorm_3_2RMSNorm1 × 32964 × 1536
26FFN_3SwiGLU1 × 32964 × 1536
27Add_3_ffnAdd1 × 32964 × 1536
28RMSNorm_4_1RMSNorm1 × 32964 × 1536
29Attention_4Grouped Query Attn1 × 32964 × 1536
30Add_4_attnAdd1 × 32964 × 1536
31RMSNorm_4_2RMSNorm1 × 32964 × 1536
32FFN_4SwiGLU1 × 32964 × 1536
33Add_4_ffnAdd1 × 32964 × 1536
34RMSNorm_5_1RMSNorm1 × 32964 × 1536
35Attention_5Grouped Query Attn1 × 32964 × 1536
36Add_5_attnAdd1 × 32964 × 1536
37RMSNorm_5_2RMSNorm1 × 32964 × 1536
38FFN_5SwiGLU1 × 32964 × 1536
39Add_5_ffnAdd1 × 32964 × 1536
40RMSNorm_6_1RMSNorm1 × 32964 × 1536
41Attention_6Grouped Query Attn1 × 32964 × 1536
42Add_6_attnAdd1 × 32964 × 1536
43RMSNorm_6_2RMSNorm1 × 32964 × 1536
44FFN_6SwiGLU1 × 32964 × 1536
45Add_6_ffnAdd1 × 32964 × 1536
46RMSNorm_7_1RMSNorm1 × 32964 × 1536
47Attention_7Grouped Query Attn1 × 32964 × 1536
48Add_7_attnAdd1 × 32964 × 1536
49RMSNorm_7_2RMSNorm1 × 32964 × 1536
50FFN_7SwiGLU1 × 32964 × 1536
51Add_7_ffnAdd1 × 32964 × 1536
52RMSNorm_8_1RMSNorm1 × 32964 × 1536
53Attention_8Grouped Query Attn1 × 32964 × 1536
54Add_8_attnAdd1 × 32964 × 1536
55RMSNorm_8_2RMSNorm1 × 32964 × 1536
56FFN_8SwiGLU1 × 32964 × 1536
57Add_8_ffnAdd1 × 32964 × 1536
58RMSNorm_9_1RMSNorm1 × 32964 × 1536
59Attention_9Grouped Query Attn1 × 32964 × 1536
60Add_9_attnAdd1 × 32964 × 1536
61RMSNorm_9_2RMSNorm1 × 32964 × 1536
62FFN_9SwiGLU1 × 32964 × 1536
63Add_9_ffnAdd1 × 32964 × 1536
64RMSNorm_10_1RMSNorm1 × 32964 × 1536
65Attention_10Grouped Query Attn1 × 32964 × 1536
66Add_10_attnAdd1 × 32964 × 1536
67RMSNorm_10_2RMSNorm1 × 32964 × 1536
68FFN_10SwiGLU1 × 32964 × 1536
69Add_10_ffnAdd1 × 32964 × 1536
70RMSNorm_11_1RMSNorm1 × 32964 × 1536
71Attention_11Grouped Query Attn1 × 32964 × 1536
72Add_11_attnAdd1 × 32964 × 1536
73RMSNorm_11_2RMSNorm1 × 32964 × 1536
74FFN_11SwiGLU1 × 32964 × 1536
75Add_11_ffnAdd1 × 32964 × 1536
76RMSNorm_12_1RMSNorm1 × 32964 × 1536
77Attention_12Grouped Query Attn1 × 32964 × 1536
78Add_12_attnAdd1 × 32964 × 1536
79RMSNorm_12_2RMSNorm1 × 32964 × 1536
80FFN_12SwiGLU1 × 32964 × 1536
81Add_12_ffnAdd1 × 32964 × 1536
82RMSNorm_13_1RMSNorm1 × 32964 × 1536
83Attention_13Grouped Query Attn1 × 32964 × 1536
84Add_13_attnAdd1 × 32964 × 1536
85RMSNorm_13_2RMSNorm1 × 32964 × 1536
86FFN_13SwiGLU1 × 32964 × 1536
87Add_13_ffnAdd1 × 32964 × 1536
88RMSNorm_14_1RMSNorm1 × 32964 × 1536
89Attention_14Grouped Query Attn1 × 32964 × 1536
90Add_14_attnAdd1 × 32964 × 1536
91RMSNorm_14_2RMSNorm1 × 32964 × 1536
92FFN_14SwiGLU1 × 32964 × 1536
93Add_14_ffnAdd1 × 32964 × 1536
94RMSNorm_15_1RMSNorm1 × 32964 × 1536
95Attention_15Grouped Query Attn1 × 32964 × 1536
96Add_15_attnAdd1 × 32964 × 1536
97RMSNorm_15_2RMSNorm1 × 32964 × 1536
98FFN_15SwiGLU1 × 32964 × 1536
99Add_15_ffnAdd1 × 32964 × 1536
100RMSNorm_16_1RMSNorm1 × 32964 × 1536
101Attention_16Grouped Query Attn1 × 32964 × 1536
102Add_16_attnAdd1 × 32964 × 1536
103RMSNorm_16_2RMSNorm1 × 32964 × 1536
104FFN_16SwiGLU1 × 32964 × 1536
105Add_16_ffnAdd1 × 32964 × 1536
106RMSNorm_17_1RMSNorm1 × 32964 × 1536
107Attention_17Grouped Query Attn1 × 32964 × 1536
108Add_17_attnAdd1 × 32964 × 1536
109RMSNorm_17_2RMSNorm1 × 32964 × 1536
110FFN_17SwiGLU1 × 32964 × 1536
111Add_17_ffnAdd1 × 32964 × 1536
112RMSNorm_18_1RMSNorm1 × 32964 × 1536
113Attention_18Grouped Query Attn1 × 32964 × 1536
114Add_18_attnAdd1 × 32964 × 1536
115RMSNorm_18_2RMSNorm1 × 32964 × 1536
116FFN_18SwiGLU1 × 32964 × 1536
117Add_18_ffnAdd1 × 32964 × 1536
118RMSNorm_19_1RMSNorm1 × 32964 × 1536
119Attention_19Grouped Query Attn1 × 32964 × 1536
120Add_19_attnAdd1 × 32964 × 1536
121RMSNorm_19_2RMSNorm1 × 32964 × 1536
122FFN_19SwiGLU1 × 32964 × 1536
123Add_19_ffnAdd1 × 32964 × 1536
124RMSNorm_20_1RMSNorm1 × 32964 × 1536
125Attention_20Grouped Query Attn1 × 32964 × 1536
126Add_20_attnAdd1 × 32964 × 1536
127RMSNorm_20_2RMSNorm1 × 32964 × 1536
128FFN_20SwiGLU1 × 32964 × 1536
129Add_20_ffnAdd1 × 32964 × 1536
130RMSNorm_21_1RMSNorm1 × 32964 × 1536
131Attention_21Grouped Query Attn1 × 32964 × 1536
132Add_21_attnAdd1 × 32964 × 1536
133RMSNorm_21_2RMSNorm1 × 32964 × 1536
134FFN_21SwiGLU1 × 32964 × 1536
135Add_21_ffnAdd1 × 32964 × 1536
136RMSNorm_22_1RMSNorm1 × 32964 × 1536
137Attention_22Grouped Query Attn1 × 32964 × 1536
138Add_22_attnAdd1 × 32964 × 1536
139RMSNorm_22_2RMSNorm1 × 32964 × 1536
140FFN_22SwiGLU1 × 32964 × 1536
141Add_22_ffnAdd1 × 32964 × 1536
142RMSNorm_23_1RMSNorm1 × 32964 × 1536
143Attention_23Grouped Query Attn1 × 32964 × 1536
144Add_23_attnAdd1 × 32964 × 1536
145RMSNorm_23_2RMSNorm1 × 32964 × 1536
146FFN_23SwiGLU1 × 32964 × 1536
147Add_23_ffnAdd1 × 32964 × 1536
148RMSNorm_24_1RMSNorm1 × 32964 × 1536
149Attention_24Grouped Query Attn1 × 32964 × 1536
150Add_24_attnAdd1 × 32964 × 1536
151RMSNorm_24_2RMSNorm1 × 32964 × 1536
152FFN_24SwiGLU1 × 32964 × 1536
153Add_24_ffnAdd1 × 32964 × 1536
154RMSNorm_25_1RMSNorm1 × 32964 × 1536
155Attention_25Grouped Query Attn1 × 32964 × 1536
156Add_25_attnAdd1 × 32964 × 1536
157RMSNorm_25_2RMSNorm1 × 32964 × 1536
158FFN_25SwiGLU1 × 32964 × 1536
159Add_25_ffnAdd1 × 32964 × 1536
160RMSNorm_26_1RMSNorm1 × 32964 × 1536
161Attention_26Grouped Query Attn1 × 32964 × 1536
162Add_26_attnAdd1 × 32964 × 1536
163RMSNorm_26_2RMSNorm1 × 32964 × 1536
164FFN_26SwiGLU1 × 32964 × 1536
165Add_26_ffnAdd1 × 32964 × 1536
166RMSNorm_27_1RMSNorm1 × 32964 × 1536
167Attention_27Grouped Query Attn1 × 32964 × 1536
168Add_27_attnAdd1 × 32964 × 1536
169RMSNorm_27_2RMSNorm1 × 32964 × 1536
170FFN_27SwiGLU1 × 32964 × 1536
171Add_27_ffnAdd1 × 32964 × 1536
172RMSNorm_28_1RMSNorm1 × 32964 × 1536
173Attention_28Grouped Query Attn1 × 32964 × 1536
174Add_28_attnAdd1 × 32964 × 1536
175RMSNorm_28_2RMSNorm1 × 32964 × 1536
176FFN_28SwiGLU1 × 32964 × 1536
177Add_28_ffnAdd1 × 32964 × 1536
178OutputOutput1 × 32964 × 1536

What the verifier says

warn"RoPE" receives input but its output is not connected. This layer will be unreachable in the forward pass. Fix: Connect the output forward, or add an Output node if this is the final layer.
dead-end
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 8960 (5.83× embedDim). Expected: ~4096. Fix: Set intermediateSize to 4096 for embedDim=1536.
swiglu-dim-convention
infoAt 28 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace jinaai/jina-reranker-m0 --plan --share