N Neurarch Architectures Models Checks Data Docs Open the app

Models / nemotron_labs_diffusion

Nemotron-Labs-Diffusion-8B

Reconstructed from its own config.json with no weights read. 107K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
7.89B
7,886,753,792 parameters
In the published checkpoint
8.49B
8,489,553,920 scalars · safetensors.total, read 2026-06-03
Delta
-7.10%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `configuration_nemotron_labs_diffusion.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
138
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$419746.59
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsdoes not fit

Structure

140 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 262144
2EmbeddingEmbedding1 × 262144 × 4096
3Positional_EmbeddingLearned Pos Embed1 × 262144 × 4096
4Attention_1Multi-Head Attention1 × 262144 × 4096
5Add_1Add1 × 262144 × 4096
6LayerNorm_1_1LayerNorm1 × 262144 × 4096
7FFN_1Feed Forward1 × 262144 × 4096
8Attention_2Multi-Head Attention1 × 262144 × 4096
9Add_2Add1 × 262144 × 4096
10LayerNorm_2_1LayerNorm1 × 262144 × 4096
11FFN_2Feed Forward1 × 262144 × 4096
12Attention_3Multi-Head Attention1 × 262144 × 4096
13Add_3Add1 × 262144 × 4096
14LayerNorm_3_1LayerNorm1 × 262144 × 4096
15FFN_3Feed Forward1 × 262144 × 4096
16Attention_4Multi-Head Attention1 × 262144 × 4096
17Add_4Add1 × 262144 × 4096
18LayerNorm_4_1LayerNorm1 × 262144 × 4096
19FFN_4Feed Forward1 × 262144 × 4096
20Attention_5Multi-Head Attention1 × 262144 × 4096
21Add_5Add1 × 262144 × 4096
22LayerNorm_5_1LayerNorm1 × 262144 × 4096
23FFN_5Feed Forward1 × 262144 × 4096
24Attention_6Multi-Head Attention1 × 262144 × 4096
25Add_6Add1 × 262144 × 4096
26LayerNorm_6_1LayerNorm1 × 262144 × 4096
27FFN_6Feed Forward1 × 262144 × 4096
28Attention_7Multi-Head Attention1 × 262144 × 4096
29Add_7Add1 × 262144 × 4096
30LayerNorm_7_1LayerNorm1 × 262144 × 4096
31FFN_7Feed Forward1 × 262144 × 4096
32Attention_8Multi-Head Attention1 × 262144 × 4096
33Add_8Add1 × 262144 × 4096
34LayerNorm_8_1LayerNorm1 × 262144 × 4096
35FFN_8Feed Forward1 × 262144 × 4096
36Attention_9Multi-Head Attention1 × 262144 × 4096
37Add_9Add1 × 262144 × 4096
38LayerNorm_9_1LayerNorm1 × 262144 × 4096
39FFN_9Feed Forward1 × 262144 × 4096
40Attention_10Multi-Head Attention1 × 262144 × 4096
41Add_10Add1 × 262144 × 4096
42LayerNorm_10_1LayerNorm1 × 262144 × 4096
43FFN_10Feed Forward1 × 262144 × 4096
44Attention_11Multi-Head Attention1 × 262144 × 4096
45Add_11Add1 × 262144 × 4096
46LayerNorm_11_1LayerNorm1 × 262144 × 4096
47FFN_11Feed Forward1 × 262144 × 4096
48Attention_12Multi-Head Attention1 × 262144 × 4096
49Add_12Add1 × 262144 × 4096
50LayerNorm_12_1LayerNorm1 × 262144 × 4096
51FFN_12Feed Forward1 × 262144 × 4096
52Attention_13Multi-Head Attention1 × 262144 × 4096
53Add_13Add1 × 262144 × 4096
54LayerNorm_13_1LayerNorm1 × 262144 × 4096
55FFN_13Feed Forward1 × 262144 × 4096
56Attention_14Multi-Head Attention1 × 262144 × 4096
57Add_14Add1 × 262144 × 4096
58LayerNorm_14_1LayerNorm1 × 262144 × 4096
59FFN_14Feed Forward1 × 262144 × 4096
60Attention_15Multi-Head Attention1 × 262144 × 4096
61Add_15Add1 × 262144 × 4096
62LayerNorm_15_1LayerNorm1 × 262144 × 4096
63FFN_15Feed Forward1 × 262144 × 4096
64Attention_16Multi-Head Attention1 × 262144 × 4096
65Add_16Add1 × 262144 × 4096
66LayerNorm_16_1LayerNorm1 × 262144 × 4096
67FFN_16Feed Forward1 × 262144 × 4096
68Attention_17Multi-Head Attention1 × 262144 × 4096
69Add_17Add1 × 262144 × 4096
70LayerNorm_17_1LayerNorm1 × 262144 × 4096
71FFN_17Feed Forward1 × 262144 × 4096
72Attention_18Multi-Head Attention1 × 262144 × 4096
73Add_18Add1 × 262144 × 4096
74LayerNorm_18_1LayerNorm1 × 262144 × 4096
75FFN_18Feed Forward1 × 262144 × 4096
76Attention_19Multi-Head Attention1 × 262144 × 4096
77Add_19Add1 × 262144 × 4096
78LayerNorm_19_1LayerNorm1 × 262144 × 4096
79FFN_19Feed Forward1 × 262144 × 4096
80Attention_20Multi-Head Attention1 × 262144 × 4096
81Add_20Add1 × 262144 × 4096
82LayerNorm_20_1LayerNorm1 × 262144 × 4096
83FFN_20Feed Forward1 × 262144 × 4096
84Attention_21Multi-Head Attention1 × 262144 × 4096
85Add_21Add1 × 262144 × 4096
86LayerNorm_21_1LayerNorm1 × 262144 × 4096
87FFN_21Feed Forward1 × 262144 × 4096
88Attention_22Multi-Head Attention1 × 262144 × 4096
89Add_22Add1 × 262144 × 4096
90LayerNorm_22_1LayerNorm1 × 262144 × 4096
91FFN_22Feed Forward1 × 262144 × 4096
92Attention_23Multi-Head Attention1 × 262144 × 4096
93Add_23Add1 × 262144 × 4096
94LayerNorm_23_1LayerNorm1 × 262144 × 4096
95FFN_23Feed Forward1 × 262144 × 4096
96Attention_24Multi-Head Attention1 × 262144 × 4096
97Add_24Add1 × 262144 × 4096
98LayerNorm_24_1LayerNorm1 × 262144 × 4096
99FFN_24Feed Forward1 × 262144 × 4096
100Attention_25Multi-Head Attention1 × 262144 × 4096
101Add_25Add1 × 262144 × 4096
102LayerNorm_25_1LayerNorm1 × 262144 × 4096
103FFN_25Feed Forward1 × 262144 × 4096
104Attention_26Multi-Head Attention1 × 262144 × 4096
105Add_26Add1 × 262144 × 4096
106LayerNorm_26_1LayerNorm1 × 262144 × 4096
107FFN_26Feed Forward1 × 262144 × 4096
108Attention_27Multi-Head Attention1 × 262144 × 4096
109Add_27Add1 × 262144 × 4096
110LayerNorm_27_1LayerNorm1 × 262144 × 4096
111FFN_27Feed Forward1 × 262144 × 4096
112Attention_28Multi-Head Attention1 × 262144 × 4096
113Add_28Add1 × 262144 × 4096
114LayerNorm_28_1LayerNorm1 × 262144 × 4096
115FFN_28Feed Forward1 × 262144 × 4096
116Attention_29Multi-Head Attention1 × 262144 × 4096
117Add_29Add1 × 262144 × 4096
118LayerNorm_29_1LayerNorm1 × 262144 × 4096
119FFN_29Feed Forward1 × 262144 × 4096
120Attention_30Multi-Head Attention1 × 262144 × 4096
121Add_30Add1 × 262144 × 4096
122LayerNorm_30_1LayerNorm1 × 262144 × 4096
123FFN_30Feed Forward1 × 262144 × 4096
124Attention_31Multi-Head Attention1 × 262144 × 4096
125Add_31Add1 × 262144 × 4096
126LayerNorm_31_1LayerNorm1 × 262144 × 4096
127FFN_31Feed Forward1 × 262144 × 4096
128Attention_32Multi-Head Attention1 × 262144 × 4096
129Add_32Add1 × 262144 × 4096
130LayerNorm_32_1LayerNorm1 × 262144 × 4096
131FFN_32Feed Forward1 × 262144 × 4096
132Attention_33Multi-Head Attention1 × 262144 × 4096
133Add_33Add1 × 262144 × 4096
134LayerNorm_33_1LayerNorm1 × 262144 × 4096
135FFN_33Feed Forward1 × 262144 × 4096
136Attention_34Multi-Head Attention1 × 262144 × 4096
137Add_34Add1 × 262144 × 4096
138LayerNorm_34_1LayerNorm1 × 262144 × 4096
139FFN_34Feed Forward1 × 262144 × 4096
140OutputOutput1 × 262144 × 4096

What the verifier says

info34 attention layers at embedDim 4096 cache full per-head K/V: about 544 KB per token at fp16, which dominates memory at long context. Grouped-query attention (e.g. 8:1) would cut this ~8×; multi-head latent attention (MLA) shrinks it ~10× or more. This is the move production LLMs make; it does not change the parameter count. Fix: Switch attention to groupedQueryAttention (set numKVHeads below numHeads, e.g. numHeads/4) or mla (a low-rank cached latent).
full-mha-serving-cost
infoAt 34 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init
warnAcross 34 attention layers this design caches 544 KB per token, so a single 8,192-token sequence needs ~4.6 GB of KV cache before weights or activations. That exceeds the 4 GB budget this rule assumes for serving headroom. Fix: Cut KV width: raise the GQA ratio (fewer numKVHeads), switch to MLA, reduce depth or embedDim, or accept a shorter serving context.
kv-cache-context-budget

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace nvidia/Nemotron-Labs-Diffusion-8B --plan --share