N Neurarch Architectures Models Checks Data Docs Open the app

Models / bart

bart-large-mnli

Reconstructed from its own config.json with no weights read. 3.2M downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
405M
405,185,536 parameters
In the published checkpoint
407M
407,344,133 scalars · safetensors.total, read 2026-09-06
Delta
-0.53%

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
136
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$14.56
CardMemory
T4 (16GB)weights + activationsfits
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

138 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 1024
2Encoder_EmbeddingEmbedding1 × 1024 × 1024
3Encoder_Position_EmbeddingLearned Pos Embed1 × 1024 × 1024
4Enc_Attn_1Multi-Head Attention1 × 1024 × 1024
5Enc_Add_1Add1 × 1024 × 1024
6Enc_Norm_1LayerNorm1 × 1024 × 1024
7Enc_FFN_1Feed Forward1 × 1024 × 1024
8Enc_Attn_2Multi-Head Attention1 × 1024 × 1024
9Enc_Add_2Add1 × 1024 × 1024
10Enc_Norm_2LayerNorm1 × 1024 × 1024
11Enc_FFN_2Feed Forward1 × 1024 × 1024
12Enc_Attn_3Multi-Head Attention1 × 1024 × 1024
13Enc_Add_3Add1 × 1024 × 1024
14Enc_Norm_3LayerNorm1 × 1024 × 1024
15Enc_FFN_3Feed Forward1 × 1024 × 1024
16Enc_Attn_4Multi-Head Attention1 × 1024 × 1024
17Enc_Add_4Add1 × 1024 × 1024
18Enc_Norm_4LayerNorm1 × 1024 × 1024
19Enc_FFN_4Feed Forward1 × 1024 × 1024
20Enc_Attn_5Multi-Head Attention1 × 1024 × 1024
21Enc_Add_5Add1 × 1024 × 1024
22Enc_Norm_5LayerNorm1 × 1024 × 1024
23Enc_FFN_5Feed Forward1 × 1024 × 1024
24Enc_Attn_6Multi-Head Attention1 × 1024 × 1024
25Enc_Add_6Add1 × 1024 × 1024
26Enc_Norm_6LayerNorm1 × 1024 × 1024
27Enc_FFN_6Feed Forward1 × 1024 × 1024
28Enc_Attn_7Multi-Head Attention1 × 1024 × 1024
29Enc_Add_7Add1 × 1024 × 1024
30Enc_Norm_7LayerNorm1 × 1024 × 1024
31Enc_FFN_7Feed Forward1 × 1024 × 1024
32Enc_Attn_8Multi-Head Attention1 × 1024 × 1024
33Enc_Add_8Add1 × 1024 × 1024
34Enc_Norm_8LayerNorm1 × 1024 × 1024
35Enc_FFN_8Feed Forward1 × 1024 × 1024
36Enc_Attn_9Multi-Head Attention1 × 1024 × 1024
37Enc_Add_9Add1 × 1024 × 1024
38Enc_Norm_9LayerNorm1 × 1024 × 1024
39Enc_FFN_9Feed Forward1 × 1024 × 1024
40Enc_Attn_10Multi-Head Attention1 × 1024 × 1024
41Enc_Add_10Add1 × 1024 × 1024
42Enc_Norm_10LayerNorm1 × 1024 × 1024
43Enc_FFN_10Feed Forward1 × 1024 × 1024
44Enc_Attn_11Multi-Head Attention1 × 1024 × 1024
45Enc_Add_11Add1 × 1024 × 1024
46Enc_Norm_11LayerNorm1 × 1024 × 1024
47Enc_FFN_11Feed Forward1 × 1024 × 1024
48Enc_Attn_12Multi-Head Attention1 × 1024 × 1024
49Enc_Add_12Add1 × 1024 × 1024
50Enc_Norm_12LayerNorm1 × 1024 × 1024
51Enc_FFN_12Feed Forward1 × 1024 × 1024
52Decoder_EmbeddingEmbedding1 × 1024 × 1024
53Dec_SelfAttn_1Multi-Head Attention1 × 1024 × 1024
54Dec_Add_1Add1 × 1024 × 1024
55Dec_Norm1_1LayerNorm1 × 1024 × 1024
56Dec_CrossAttn_1Cross-Attention1 × 1024 × 1024
57Dec_Add2_1Add1 × 1024 × 1024
58Dec_Norm2_1LayerNorm1 × 1024 × 1024
59Dec_FFN_1Feed Forward1 × 1024 × 1024
60Dec_SelfAttn_2Multi-Head Attention1 × 1024 × 1024
61Dec_Add_2Add1 × 1024 × 1024
62Dec_Norm1_2LayerNorm1 × 1024 × 1024
63Dec_CrossAttn_2Cross-Attention1 × 1024 × 1024
64Dec_Add2_2Add1 × 1024 × 1024
65Dec_Norm2_2LayerNorm1 × 1024 × 1024
66Dec_FFN_2Feed Forward1 × 1024 × 1024
67Dec_SelfAttn_3Multi-Head Attention1 × 1024 × 1024
68Dec_Add_3Add1 × 1024 × 1024
69Dec_Norm1_3LayerNorm1 × 1024 × 1024
70Dec_CrossAttn_3Cross-Attention1 × 1024 × 1024
71Dec_Add2_3Add1 × 1024 × 1024
72Dec_Norm2_3LayerNorm1 × 1024 × 1024
73Dec_FFN_3Feed Forward1 × 1024 × 1024
74Dec_SelfAttn_4Multi-Head Attention1 × 1024 × 1024
75Dec_Add_4Add1 × 1024 × 1024
76Dec_Norm1_4LayerNorm1 × 1024 × 1024
77Dec_CrossAttn_4Cross-Attention1 × 1024 × 1024
78Dec_Add2_4Add1 × 1024 × 1024
79Dec_Norm2_4LayerNorm1 × 1024 × 1024
80Dec_FFN_4Feed Forward1 × 1024 × 1024
81Dec_SelfAttn_5Multi-Head Attention1 × 1024 × 1024
82Dec_Add_5Add1 × 1024 × 1024
83Dec_Norm1_5LayerNorm1 × 1024 × 1024
84Dec_CrossAttn_5Cross-Attention1 × 1024 × 1024
85Dec_Add2_5Add1 × 1024 × 1024
86Dec_Norm2_5LayerNorm1 × 1024 × 1024
87Dec_FFN_5Feed Forward1 × 1024 × 1024
88Dec_SelfAttn_6Multi-Head Attention1 × 1024 × 1024
89Dec_Add_6Add1 × 1024 × 1024
90Dec_Norm1_6LayerNorm1 × 1024 × 1024
91Dec_CrossAttn_6Cross-Attention1 × 1024 × 1024
92Dec_Add2_6Add1 × 1024 × 1024
93Dec_Norm2_6LayerNorm1 × 1024 × 1024
94Dec_FFN_6Feed Forward1 × 1024 × 1024
95Dec_SelfAttn_7Multi-Head Attention1 × 1024 × 1024
96Dec_Add_7Add1 × 1024 × 1024
97Dec_Norm1_7LayerNorm1 × 1024 × 1024
98Dec_CrossAttn_7Cross-Attention1 × 1024 × 1024
99Dec_Add2_7Add1 × 1024 × 1024
100Dec_Norm2_7LayerNorm1 × 1024 × 1024
101Dec_FFN_7Feed Forward1 × 1024 × 1024
102Dec_SelfAttn_8Multi-Head Attention1 × 1024 × 1024
103Dec_Add_8Add1 × 1024 × 1024
104Dec_Norm1_8LayerNorm1 × 1024 × 1024
105Dec_CrossAttn_8Cross-Attention1 × 1024 × 1024
106Dec_Add2_8Add1 × 1024 × 1024
107Dec_Norm2_8LayerNorm1 × 1024 × 1024
108Dec_FFN_8Feed Forward1 × 1024 × 1024
109Dec_SelfAttn_9Multi-Head Attention1 × 1024 × 1024
110Dec_Add_9Add1 × 1024 × 1024
111Dec_Norm1_9LayerNorm1 × 1024 × 1024
112Dec_CrossAttn_9Cross-Attention1 × 1024 × 1024
113Dec_Add2_9Add1 × 1024 × 1024
114Dec_Norm2_9LayerNorm1 × 1024 × 1024
115Dec_FFN_9Feed Forward1 × 1024 × 1024
116Dec_SelfAttn_10Multi-Head Attention1 × 1024 × 1024
117Dec_Add_10Add1 × 1024 × 1024
118Dec_Norm1_10LayerNorm1 × 1024 × 1024
119Dec_CrossAttn_10Cross-Attention1 × 1024 × 1024
120Dec_Add2_10Add1 × 1024 × 1024
121Dec_Norm2_10LayerNorm1 × 1024 × 1024
122Dec_FFN_10Feed Forward1 × 1024 × 1024
123Dec_SelfAttn_11Multi-Head Attention1 × 1024 × 1024
124Dec_Add_11Add1 × 1024 × 1024
125Dec_Norm1_11LayerNorm1 × 1024 × 1024
126Dec_CrossAttn_11Cross-Attention1 × 1024 × 1024
127Dec_Add2_11Add1 × 1024 × 1024
128Dec_Norm2_11LayerNorm1 × 1024 × 1024
129Dec_FFN_11Feed Forward1 × 1024 × 1024
130Dec_SelfAttn_12Multi-Head Attention1 × 1024 × 1024
131Dec_Add_12Add1 × 1024 × 1024
132Dec_Norm1_12LayerNorm1 × 1024 × 1024
133Dec_CrossAttn_12Cross-Attention1 × 1024 × 1024
134Dec_Add2_12Add1 × 1024 × 1024
135Dec_Norm2_12LayerNorm1 × 1024 × 1024
136Dec_FFN_12Feed Forward1 × 1024 × 1024
137LM_HeadLinear1 × 1024 × 50265
138OutputOutput1 × 1024 × 50265

What the verifier says

infoAt 24 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace facebook/bart-large-mnli --plan --share