N Neurarch Architectures Models Checks Data Docs Open the app

Models / nemotron_h_puzzle

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4

Reconstructed from its own config.json with no weights read. 156K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
1629B
1,629,269,336,064 parameters
In the published checkpoint
44.54B
44,539,891,200 scalars · safetensors.total, read 2026-07-07
Delta
+3558%

quantized This checkpoint is stored quantized (modelopt). The tensor count in the file counts stored elements under a packing scheme, not logical parameters, so the two numbers below are not measuring the same thing in either direction.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
76
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$353921.42
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsdoes not fit

Structure

78 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 262144
2EmbeddingEmbedding1 × 262144 × 4096
3Positional_EmbeddingLearned Pos Embed1 × 262144 × 4096
4LayerNorm_1_1LayerNorm1 × 262144 × 4096
5Attention_1Grouped Query Attn1 × 262144 × 4096
6Add_1_attnAdd1 × 262144 × 4096
7LayerNorm_1_2LayerNorm1 × 262144 × 4096
8MoE_1Shared-Expert MoE1 × 262144 × 4096
9Add_1_ffnAdd1 × 262144 × 4096
10LayerNorm_2_1LayerNorm1 × 262144 × 4096
11Attention_2Grouped Query Attn1 × 262144 × 4096
12Add_2_attnAdd1 × 262144 × 4096
13LayerNorm_2_2LayerNorm1 × 262144 × 4096
14MoE_2Shared-Expert MoE1 × 262144 × 4096
15Add_2_ffnAdd1 × 262144 × 4096
16LayerNorm_3_1LayerNorm1 × 262144 × 4096
17Attention_3Grouped Query Attn1 × 262144 × 4096
18Add_3_attnAdd1 × 262144 × 4096
19LayerNorm_3_2LayerNorm1 × 262144 × 4096
20MoE_3Shared-Expert MoE1 × 262144 × 4096
21Add_3_ffnAdd1 × 262144 × 4096
22LayerNorm_4_1LayerNorm1 × 262144 × 4096
23Attention_4Grouped Query Attn1 × 262144 × 4096
24Add_4_attnAdd1 × 262144 × 4096
25LayerNorm_4_2LayerNorm1 × 262144 × 4096
26MoE_4Shared-Expert MoE1 × 262144 × 4096
27Add_4_ffnAdd1 × 262144 × 4096
28LayerNorm_5_1LayerNorm1 × 262144 × 4096
29Attention_5Grouped Query Attn1 × 262144 × 4096
30Add_5_attnAdd1 × 262144 × 4096
31LayerNorm_5_2LayerNorm1 × 262144 × 4096
32MoE_5Shared-Expert MoE1 × 262144 × 4096
33Add_5_ffnAdd1 × 262144 × 4096
34LayerNorm_6_1LayerNorm1 × 262144 × 4096
35Attention_6Grouped Query Attn1 × 262144 × 4096
36Add_6_attnAdd1 × 262144 × 4096
37LayerNorm_6_2LayerNorm1 × 262144 × 4096
38MoE_6Shared-Expert MoE1 × 262144 × 4096
39Add_6_ffnAdd1 × 262144 × 4096
40LayerNorm_7_1LayerNorm1 × 262144 × 4096
41Attention_7Grouped Query Attn1 × 262144 × 4096
42Add_7_attnAdd1 × 262144 × 4096
43LayerNorm_7_2LayerNorm1 × 262144 × 4096
44MoE_7Shared-Expert MoE1 × 262144 × 4096
45Add_7_ffnAdd1 × 262144 × 4096
46LayerNorm_8_1LayerNorm1 × 262144 × 4096
47Attention_8Grouped Query Attn1 × 262144 × 4096
48Add_8_attnAdd1 × 262144 × 4096
49LayerNorm_8_2LayerNorm1 × 262144 × 4096
50MoE_8Shared-Expert MoE1 × 262144 × 4096
51Add_8_ffnAdd1 × 262144 × 4096
52LayerNorm_9_1LayerNorm1 × 262144 × 4096
53Attention_9Grouped Query Attn1 × 262144 × 4096
54Add_9_attnAdd1 × 262144 × 4096
55LayerNorm_9_2LayerNorm1 × 262144 × 4096
56MoE_9Shared-Expert MoE1 × 262144 × 4096
57Add_9_ffnAdd1 × 262144 × 4096
58LayerNorm_10_1LayerNorm1 × 262144 × 4096
59Attention_10Grouped Query Attn1 × 262144 × 4096
60Add_10_attnAdd1 × 262144 × 4096
61LayerNorm_10_2LayerNorm1 × 262144 × 4096
62MoE_10Shared-Expert MoE1 × 262144 × 4096
63Add_10_ffnAdd1 × 262144 × 4096
64LayerNorm_11_1LayerNorm1 × 262144 × 4096
65Attention_11Grouped Query Attn1 × 262144 × 4096
66Add_11_attnAdd1 × 262144 × 4096
67LayerNorm_11_2LayerNorm1 × 262144 × 4096
68MoE_11Shared-Expert MoE1 × 262144 × 4096
69Add_11_ffnAdd1 × 262144 × 4096
70LayerNorm_12_1LayerNorm1 × 262144 × 4096
71Attention_12Grouped Query Attn1 × 262144 × 4096
72Add_12_attnAdd1 × 262144 × 4096
73LayerNorm_12_2LayerNorm1 × 262144 × 4096
74MoE_12Shared-Expert MoE1 × 262144 × 4096
75Add_12_ffnAdd1 × 262144 × 4096
76Final_LayerNormLayerNorm1 × 262144 × 4096
77LM_HeadLinear1 × 262144 × 131072
78OutputOutput1 × 262144 × 131072

What the verifier says

infoAt 12 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4 --plan --share