N Neurarch Architectures Models Checks Data Docs Open the app

Models / qwen2

Qwen2.5-3B-Instruct

Reconstructed from its own config.json with no weights read. 6.0M downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
3.09B
3,085,844,480 parameters
In the published checkpoint
3.09B
3,085,938,688 scalars · safetensors.total, read 2024-09-25
Delta
-0.00%

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
218
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$6366.49
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsfits

Structure

220 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 32768
2EmbeddingEmbedding1 × 32768 × 2048
3RoPERoPE1 × 32768 × 2048
4RMSNorm_1_1RMSNorm1 × 32768 × 2048
5Attention_1Grouped Query Attn1 × 32768 × 2048
6Add_1_attnAdd1 × 32768 × 2048
7RMSNorm_1_2RMSNorm1 × 32768 × 2048
8FFN_1SwiGLU1 × 32768 × 2048
9Add_1_ffnAdd1 × 32768 × 2048
10RMSNorm_2_1RMSNorm1 × 32768 × 2048
11Attention_2Grouped Query Attn1 × 32768 × 2048
12Add_2_attnAdd1 × 32768 × 2048
13RMSNorm_2_2RMSNorm1 × 32768 × 2048
14FFN_2SwiGLU1 × 32768 × 2048
15Add_2_ffnAdd1 × 32768 × 2048
16RMSNorm_3_1RMSNorm1 × 32768 × 2048
17Attention_3Grouped Query Attn1 × 32768 × 2048
18Add_3_attnAdd1 × 32768 × 2048
19RMSNorm_3_2RMSNorm1 × 32768 × 2048
20FFN_3SwiGLU1 × 32768 × 2048
21Add_3_ffnAdd1 × 32768 × 2048
22RMSNorm_4_1RMSNorm1 × 32768 × 2048
23Attention_4Grouped Query Attn1 × 32768 × 2048
24Add_4_attnAdd1 × 32768 × 2048
25RMSNorm_4_2RMSNorm1 × 32768 × 2048
26FFN_4SwiGLU1 × 32768 × 2048
27Add_4_ffnAdd1 × 32768 × 2048
28RMSNorm_5_1RMSNorm1 × 32768 × 2048
29Attention_5Grouped Query Attn1 × 32768 × 2048
30Add_5_attnAdd1 × 32768 × 2048
31RMSNorm_5_2RMSNorm1 × 32768 × 2048
32FFN_5SwiGLU1 × 32768 × 2048
33Add_5_ffnAdd1 × 32768 × 2048
34RMSNorm_6_1RMSNorm1 × 32768 × 2048
35Attention_6Grouped Query Attn1 × 32768 × 2048
36Add_6_attnAdd1 × 32768 × 2048
37RMSNorm_6_2RMSNorm1 × 32768 × 2048
38FFN_6SwiGLU1 × 32768 × 2048
39Add_6_ffnAdd1 × 32768 × 2048
40RMSNorm_7_1RMSNorm1 × 32768 × 2048
41Attention_7Grouped Query Attn1 × 32768 × 2048
42Add_7_attnAdd1 × 32768 × 2048
43RMSNorm_7_2RMSNorm1 × 32768 × 2048
44FFN_7SwiGLU1 × 32768 × 2048
45Add_7_ffnAdd1 × 32768 × 2048
46RMSNorm_8_1RMSNorm1 × 32768 × 2048
47Attention_8Grouped Query Attn1 × 32768 × 2048
48Add_8_attnAdd1 × 32768 × 2048
49RMSNorm_8_2RMSNorm1 × 32768 × 2048
50FFN_8SwiGLU1 × 32768 × 2048
51Add_8_ffnAdd1 × 32768 × 2048
52RMSNorm_9_1RMSNorm1 × 32768 × 2048
53Attention_9Grouped Query Attn1 × 32768 × 2048
54Add_9_attnAdd1 × 32768 × 2048
55RMSNorm_9_2RMSNorm1 × 32768 × 2048
56FFN_9SwiGLU1 × 32768 × 2048
57Add_9_ffnAdd1 × 32768 × 2048
58RMSNorm_10_1RMSNorm1 × 32768 × 2048
59Attention_10Grouped Query Attn1 × 32768 × 2048
60Add_10_attnAdd1 × 32768 × 2048
61RMSNorm_10_2RMSNorm1 × 32768 × 2048
62FFN_10SwiGLU1 × 32768 × 2048
63Add_10_ffnAdd1 × 32768 × 2048
64RMSNorm_11_1RMSNorm1 × 32768 × 2048
65Attention_11Grouped Query Attn1 × 32768 × 2048
66Add_11_attnAdd1 × 32768 × 2048
67RMSNorm_11_2RMSNorm1 × 32768 × 2048
68FFN_11SwiGLU1 × 32768 × 2048
69Add_11_ffnAdd1 × 32768 × 2048
70RMSNorm_12_1RMSNorm1 × 32768 × 2048
71Attention_12Grouped Query Attn1 × 32768 × 2048
72Add_12_attnAdd1 × 32768 × 2048
73RMSNorm_12_2RMSNorm1 × 32768 × 2048
74FFN_12SwiGLU1 × 32768 × 2048
75Add_12_ffnAdd1 × 32768 × 2048
76RMSNorm_13_1RMSNorm1 × 32768 × 2048
77Attention_13Grouped Query Attn1 × 32768 × 2048
78Add_13_attnAdd1 × 32768 × 2048
79RMSNorm_13_2RMSNorm1 × 32768 × 2048
80FFN_13SwiGLU1 × 32768 × 2048
81Add_13_ffnAdd1 × 32768 × 2048
82RMSNorm_14_1RMSNorm1 × 32768 × 2048
83Attention_14Grouped Query Attn1 × 32768 × 2048
84Add_14_attnAdd1 × 32768 × 2048
85RMSNorm_14_2RMSNorm1 × 32768 × 2048
86FFN_14SwiGLU1 × 32768 × 2048
87Add_14_ffnAdd1 × 32768 × 2048
88RMSNorm_15_1RMSNorm1 × 32768 × 2048
89Attention_15Grouped Query Attn1 × 32768 × 2048
90Add_15_attnAdd1 × 32768 × 2048
91RMSNorm_15_2RMSNorm1 × 32768 × 2048
92FFN_15SwiGLU1 × 32768 × 2048
93Add_15_ffnAdd1 × 32768 × 2048
94RMSNorm_16_1RMSNorm1 × 32768 × 2048
95Attention_16Grouped Query Attn1 × 32768 × 2048
96Add_16_attnAdd1 × 32768 × 2048
97RMSNorm_16_2RMSNorm1 × 32768 × 2048
98FFN_16SwiGLU1 × 32768 × 2048
99Add_16_ffnAdd1 × 32768 × 2048
100RMSNorm_17_1RMSNorm1 × 32768 × 2048
101Attention_17Grouped Query Attn1 × 32768 × 2048
102Add_17_attnAdd1 × 32768 × 2048
103RMSNorm_17_2RMSNorm1 × 32768 × 2048
104FFN_17SwiGLU1 × 32768 × 2048
105Add_17_ffnAdd1 × 32768 × 2048
106RMSNorm_18_1RMSNorm1 × 32768 × 2048
107Attention_18Grouped Query Attn1 × 32768 × 2048
108Add_18_attnAdd1 × 32768 × 2048
109RMSNorm_18_2RMSNorm1 × 32768 × 2048
110FFN_18SwiGLU1 × 32768 × 2048
111Add_18_ffnAdd1 × 32768 × 2048
112RMSNorm_19_1RMSNorm1 × 32768 × 2048
113Attention_19Grouped Query Attn1 × 32768 × 2048
114Add_19_attnAdd1 × 32768 × 2048
115RMSNorm_19_2RMSNorm1 × 32768 × 2048
116FFN_19SwiGLU1 × 32768 × 2048
117Add_19_ffnAdd1 × 32768 × 2048
118RMSNorm_20_1RMSNorm1 × 32768 × 2048
119Attention_20Grouped Query Attn1 × 32768 × 2048
120Add_20_attnAdd1 × 32768 × 2048
121RMSNorm_20_2RMSNorm1 × 32768 × 2048
122FFN_20SwiGLU1 × 32768 × 2048
123Add_20_ffnAdd1 × 32768 × 2048
124RMSNorm_21_1RMSNorm1 × 32768 × 2048
125Attention_21Grouped Query Attn1 × 32768 × 2048
126Add_21_attnAdd1 × 32768 × 2048
127RMSNorm_21_2RMSNorm1 × 32768 × 2048
128FFN_21SwiGLU1 × 32768 × 2048
129Add_21_ffnAdd1 × 32768 × 2048
130RMSNorm_22_1RMSNorm1 × 32768 × 2048
131Attention_22Grouped Query Attn1 × 32768 × 2048
132Add_22_attnAdd1 × 32768 × 2048
133RMSNorm_22_2RMSNorm1 × 32768 × 2048
134FFN_22SwiGLU1 × 32768 × 2048
135Add_22_ffnAdd1 × 32768 × 2048
136RMSNorm_23_1RMSNorm1 × 32768 × 2048
137Attention_23Grouped Query Attn1 × 32768 × 2048
138Add_23_attnAdd1 × 32768 × 2048
139RMSNorm_23_2RMSNorm1 × 32768 × 2048
140FFN_23SwiGLU1 × 32768 × 2048
141Add_23_ffnAdd1 × 32768 × 2048
142RMSNorm_24_1RMSNorm1 × 32768 × 2048
143Attention_24Grouped Query Attn1 × 32768 × 2048
144Add_24_attnAdd1 × 32768 × 2048
145RMSNorm_24_2RMSNorm1 × 32768 × 2048
146FFN_24SwiGLU1 × 32768 × 2048
147Add_24_ffnAdd1 × 32768 × 2048
148RMSNorm_25_1RMSNorm1 × 32768 × 2048
149Attention_25Grouped Query Attn1 × 32768 × 2048
150Add_25_attnAdd1 × 32768 × 2048
151RMSNorm_25_2RMSNorm1 × 32768 × 2048
152FFN_25SwiGLU1 × 32768 × 2048
153Add_25_ffnAdd1 × 32768 × 2048
154RMSNorm_26_1RMSNorm1 × 32768 × 2048
155Attention_26Grouped Query Attn1 × 32768 × 2048
156Add_26_attnAdd1 × 32768 × 2048
157RMSNorm_26_2RMSNorm1 × 32768 × 2048
158FFN_26SwiGLU1 × 32768 × 2048
159Add_26_ffnAdd1 × 32768 × 2048
160RMSNorm_27_1RMSNorm1 × 32768 × 2048
161Attention_27Grouped Query Attn1 × 32768 × 2048
162Add_27_attnAdd1 × 32768 × 2048
163RMSNorm_27_2RMSNorm1 × 32768 × 2048
164FFN_27SwiGLU1 × 32768 × 2048
165Add_27_ffnAdd1 × 32768 × 2048
166RMSNorm_28_1RMSNorm1 × 32768 × 2048
167Attention_28Grouped Query Attn1 × 32768 × 2048
168Add_28_attnAdd1 × 32768 × 2048
169RMSNorm_28_2RMSNorm1 × 32768 × 2048
170FFN_28SwiGLU1 × 32768 × 2048
171Add_28_ffnAdd1 × 32768 × 2048
172RMSNorm_29_1RMSNorm1 × 32768 × 2048
173Attention_29Grouped Query Attn1 × 32768 × 2048
174Add_29_attnAdd1 × 32768 × 2048
175RMSNorm_29_2RMSNorm1 × 32768 × 2048
176FFN_29SwiGLU1 × 32768 × 2048
177Add_29_ffnAdd1 × 32768 × 2048
178RMSNorm_30_1RMSNorm1 × 32768 × 2048
179Attention_30Grouped Query Attn1 × 32768 × 2048
180Add_30_attnAdd1 × 32768 × 2048
181RMSNorm_30_2RMSNorm1 × 32768 × 2048
182FFN_30SwiGLU1 × 32768 × 2048
183Add_30_ffnAdd1 × 32768 × 2048
184RMSNorm_31_1RMSNorm1 × 32768 × 2048
185Attention_31Grouped Query Attn1 × 32768 × 2048
186Add_31_attnAdd1 × 32768 × 2048
187RMSNorm_31_2RMSNorm1 × 32768 × 2048
188FFN_31SwiGLU1 × 32768 × 2048
189Add_31_ffnAdd1 × 32768 × 2048
190RMSNorm_32_1RMSNorm1 × 32768 × 2048
191Attention_32Grouped Query Attn1 × 32768 × 2048
192Add_32_attnAdd1 × 32768 × 2048
193RMSNorm_32_2RMSNorm1 × 32768 × 2048
194FFN_32SwiGLU1 × 32768 × 2048
195Add_32_ffnAdd1 × 32768 × 2048
196RMSNorm_33_1RMSNorm1 × 32768 × 2048
197Attention_33Grouped Query Attn1 × 32768 × 2048
198Add_33_attnAdd1 × 32768 × 2048
199RMSNorm_33_2RMSNorm1 × 32768 × 2048
200FFN_33SwiGLU1 × 32768 × 2048
201Add_33_ffnAdd1 × 32768 × 2048
202RMSNorm_34_1RMSNorm1 × 32768 × 2048
203Attention_34Grouped Query Attn1 × 32768 × 2048
204Add_34_attnAdd1 × 32768 × 2048
205RMSNorm_34_2RMSNorm1 × 32768 × 2048
206FFN_34SwiGLU1 × 32768 × 2048
207Add_34_ffnAdd1 × 32768 × 2048
208RMSNorm_35_1RMSNorm1 × 32768 × 2048
209Attention_35Grouped Query Attn1 × 32768 × 2048
210Add_35_attnAdd1 × 32768 × 2048
211RMSNorm_35_2RMSNorm1 × 32768 × 2048
212FFN_35SwiGLU1 × 32768 × 2048
213Add_35_ffnAdd1 × 32768 × 2048
214RMSNorm_36_1RMSNorm1 × 32768 × 2048
215Attention_36Grouped Query Attn1 × 32768 × 2048
216Add_36_attnAdd1 × 32768 × 2048
217RMSNorm_36_2RMSNorm1 × 32768 × 2048
218FFN_36SwiGLU1 × 32768 × 2048
219Add_36_ffnAdd1 × 32768 × 2048
220OutputOutput1 × 32768 × 2048

What the verifier says

infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoLLaMA uses intermediateSize ≈ ⌊(8/3 × D) / 256⌋ × 256. Current: 11008 (5.38× embedDim). Expected: ~5376. Fix: Set intermediateSize to 5376 for embedDim=2048.
swiglu-dim-convention
infoAt 36 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace Qwen/Qwen2.5-3B-Instruct --plan --share

Other qwen2 checkpoints

gte-Qwen2-1.5B-instruct
1.78B derived · +0.01% against the checkpoint
gte-Qwen2-7B-instruct
7.61B derived · +0.00% against the checkpoint
Qwen2.5-0.5B-Instruct
494M derived · -0.01% against the checkpoint
Qwen2.5-1.5B-Instruct
1.54B derived · -0.00% against the checkpoint