N Neurarch Architectures Models Checks Data Docs Open the app

Models / qwen3

voyage-4-nano

Reconstructed from its own config.json with no weights read. 292K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
344M
344,350,720 parameters
In the published checkpoint
346M
346,451,968 scalars · safetensors.total, read 2026-03-02
Delta
-0.61%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `modeling_qwen3_bidirectional.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
74
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$1832.70
CardMemory
T4 (16GB)weights + activationsfits
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

76 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 40960
2EmbeddingEmbedding1 × 40960 × 1024
3RoPERoPE1 × 40960 × 1024
4RMSNorm_1_1RMSNorm1 × 40960 × 1024
5Attention_1Grouped Query Attn1 × 40960 × 1024
6Add_1_attnAdd1 × 40960 × 1024
7RMSNorm_1_2RMSNorm1 × 40960 × 1024
8FFN_1SwiGLU1 × 40960 × 1024
9Add_1_ffnAdd1 × 40960 × 1024
10RMSNorm_2_1RMSNorm1 × 40960 × 1024
11Attention_2Grouped Query Attn1 × 40960 × 1024
12Add_2_attnAdd1 × 40960 × 1024
13RMSNorm_2_2RMSNorm1 × 40960 × 1024
14FFN_2SwiGLU1 × 40960 × 1024
15Add_2_ffnAdd1 × 40960 × 1024
16RMSNorm_3_1RMSNorm1 × 40960 × 1024
17Attention_3Grouped Query Attn1 × 40960 × 1024
18Add_3_attnAdd1 × 40960 × 1024
19RMSNorm_3_2RMSNorm1 × 40960 × 1024
20FFN_3SwiGLU1 × 40960 × 1024
21Add_3_ffnAdd1 × 40960 × 1024
22RMSNorm_4_1RMSNorm1 × 40960 × 1024
23Attention_4Grouped Query Attn1 × 40960 × 1024
24Add_4_attnAdd1 × 40960 × 1024
25RMSNorm_4_2RMSNorm1 × 40960 × 1024
26FFN_4SwiGLU1 × 40960 × 1024
27Add_4_ffnAdd1 × 40960 × 1024
28RMSNorm_5_1RMSNorm1 × 40960 × 1024
29Attention_5Grouped Query Attn1 × 40960 × 1024
30Add_5_attnAdd1 × 40960 × 1024
31RMSNorm_5_2RMSNorm1 × 40960 × 1024
32FFN_5SwiGLU1 × 40960 × 1024
33Add_5_ffnAdd1 × 40960 × 1024
34RMSNorm_6_1RMSNorm1 × 40960 × 1024
35Attention_6Grouped Query Attn1 × 40960 × 1024
36Add_6_attnAdd1 × 40960 × 1024
37RMSNorm_6_2RMSNorm1 × 40960 × 1024
38FFN_6SwiGLU1 × 40960 × 1024
39Add_6_ffnAdd1 × 40960 × 1024
40RMSNorm_7_1RMSNorm1 × 40960 × 1024
41Attention_7Grouped Query Attn1 × 40960 × 1024
42Add_7_attnAdd1 × 40960 × 1024
43RMSNorm_7_2RMSNorm1 × 40960 × 1024
44FFN_7SwiGLU1 × 40960 × 1024
45Add_7_ffnAdd1 × 40960 × 1024
46RMSNorm_8_1RMSNorm1 × 40960 × 1024
47Attention_8Grouped Query Attn1 × 40960 × 1024
48Add_8_attnAdd1 × 40960 × 1024
49RMSNorm_8_2RMSNorm1 × 40960 × 1024
50FFN_8SwiGLU1 × 40960 × 1024
51Add_8_ffnAdd1 × 40960 × 1024
52RMSNorm_9_1RMSNorm1 × 40960 × 1024
53Attention_9Grouped Query Attn1 × 40960 × 1024
54Add_9_attnAdd1 × 40960 × 1024
55RMSNorm_9_2RMSNorm1 × 40960 × 1024
56FFN_9SwiGLU1 × 40960 × 1024
57Add_9_ffnAdd1 × 40960 × 1024
58RMSNorm_10_1RMSNorm1 × 40960 × 1024
59Attention_10Grouped Query Attn1 × 40960 × 1024
60Add_10_attnAdd1 × 40960 × 1024
61RMSNorm_10_2RMSNorm1 × 40960 × 1024
62FFN_10SwiGLU1 × 40960 × 1024
63Add_10_ffnAdd1 × 40960 × 1024
64RMSNorm_11_1RMSNorm1 × 40960 × 1024
65Attention_11Grouped Query Attn1 × 40960 × 1024
66Add_11_attnAdd1 × 40960 × 1024
67RMSNorm_11_2RMSNorm1 × 40960 × 1024
68FFN_11SwiGLU1 × 40960 × 1024
69Add_11_ffnAdd1 × 40960 × 1024
70RMSNorm_12_1RMSNorm1 × 40960 × 1024
71Attention_12Grouped Query Attn1 × 40960 × 1024
72Add_12_attnAdd1 × 40960 × 1024
73RMSNorm_12_2RMSNorm1 × 40960 × 1024
74FFN_12SwiGLU1 × 40960 × 1024
75Add_12_ffnAdd1 × 40960 × 1024
76OutputOutput1 × 40960 × 1024

What the verifier says

infoAt 12 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace voyageai/voyage-4-nano --plan --share

Other qwen3 checkpoints

jina-reranker-v3
596M derived · -0.13% against the checkpoint
Kimi-K3-DSpark
4.26B derived · +89.3% against the checkpoint
Qwen3-0.6B
596M derived · -20.7% against the checkpoint
Qwen3-1.7B
1.72B derived · -15.3% against the checkpoint