N Neurarch Architectures Models Checks Data Docs Open the app

Models / t5

t5-base

Reconstructed from its own config.json with no weights read. 4.4M downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
223M
223,113,600 parameters
In the published checkpoint
223M
222,903,936 scalars · safetensors.total, read 2024-02-14
Delta
+0.09%

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
136
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$3.95
CardMemory
T4 (16GB)weights + activationsfits
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

138 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 512
2Encoder_EmbeddingEmbedding1 × 512 × 768
3Relative_Position_BiasRelative Position Bias1 × 512 × 768
4Enc_Attn_1Multi-Head Attention1 × 512 × 768
5Enc_Add_1Add1 × 512 × 768
6Enc_Norm_1LayerNorm1 × 512 × 768
7Enc_FFN_1Feed Forward1 × 512 × 768
8Enc_Attn_2Multi-Head Attention1 × 512 × 768
9Enc_Add_2Add1 × 512 × 768
10Enc_Norm_2LayerNorm1 × 512 × 768
11Enc_FFN_2Feed Forward1 × 512 × 768
12Enc_Attn_3Multi-Head Attention1 × 512 × 768
13Enc_Add_3Add1 × 512 × 768
14Enc_Norm_3LayerNorm1 × 512 × 768
15Enc_FFN_3Feed Forward1 × 512 × 768
16Enc_Attn_4Multi-Head Attention1 × 512 × 768
17Enc_Add_4Add1 × 512 × 768
18Enc_Norm_4LayerNorm1 × 512 × 768
19Enc_FFN_4Feed Forward1 × 512 × 768
20Enc_Attn_5Multi-Head Attention1 × 512 × 768
21Enc_Add_5Add1 × 512 × 768
22Enc_Norm_5LayerNorm1 × 512 × 768
23Enc_FFN_5Feed Forward1 × 512 × 768
24Enc_Attn_6Multi-Head Attention1 × 512 × 768
25Enc_Add_6Add1 × 512 × 768
26Enc_Norm_6LayerNorm1 × 512 × 768
27Enc_FFN_6Feed Forward1 × 512 × 768
28Enc_Attn_7Multi-Head Attention1 × 512 × 768
29Enc_Add_7Add1 × 512 × 768
30Enc_Norm_7LayerNorm1 × 512 × 768
31Enc_FFN_7Feed Forward1 × 512 × 768
32Enc_Attn_8Multi-Head Attention1 × 512 × 768
33Enc_Add_8Add1 × 512 × 768
34Enc_Norm_8LayerNorm1 × 512 × 768
35Enc_FFN_8Feed Forward1 × 512 × 768
36Enc_Attn_9Multi-Head Attention1 × 512 × 768
37Enc_Add_9Add1 × 512 × 768
38Enc_Norm_9LayerNorm1 × 512 × 768
39Enc_FFN_9Feed Forward1 × 512 × 768
40Enc_Attn_10Multi-Head Attention1 × 512 × 768
41Enc_Add_10Add1 × 512 × 768
42Enc_Norm_10LayerNorm1 × 512 × 768
43Enc_FFN_10Feed Forward1 × 512 × 768
44Enc_Attn_11Multi-Head Attention1 × 512 × 768
45Enc_Add_11Add1 × 512 × 768
46Enc_Norm_11LayerNorm1 × 512 × 768
47Enc_FFN_11Feed Forward1 × 512 × 768
48Enc_Attn_12Multi-Head Attention1 × 512 × 768
49Enc_Add_12Add1 × 512 × 768
50Enc_Norm_12LayerNorm1 × 512 × 768
51Enc_FFN_12Feed Forward1 × 512 × 768
52Decoder_EmbeddingEmbedding1 × 512 × 768
53Dec_SelfAttn_1Multi-Head Attention1 × 512 × 768
54Dec_Add_1Add1 × 512 × 768
55Dec_Norm1_1LayerNorm1 × 512 × 768
56Dec_CrossAttn_1Cross-Attention1 × 512 × 768
57Dec_Add2_1Add1 × 512 × 768
58Dec_Norm2_1LayerNorm1 × 512 × 768
59Dec_FFN_1Feed Forward1 × 512 × 768
60Dec_SelfAttn_2Multi-Head Attention1 × 512 × 768
61Dec_Add_2Add1 × 512 × 768
62Dec_Norm1_2LayerNorm1 × 512 × 768
63Dec_CrossAttn_2Cross-Attention1 × 512 × 768
64Dec_Add2_2Add1 × 512 × 768
65Dec_Norm2_2LayerNorm1 × 512 × 768
66Dec_FFN_2Feed Forward1 × 512 × 768
67Dec_SelfAttn_3Multi-Head Attention1 × 512 × 768
68Dec_Add_3Add1 × 512 × 768
69Dec_Norm1_3LayerNorm1 × 512 × 768
70Dec_CrossAttn_3Cross-Attention1 × 512 × 768
71Dec_Add2_3Add1 × 512 × 768
72Dec_Norm2_3LayerNorm1 × 512 × 768
73Dec_FFN_3Feed Forward1 × 512 × 768
74Dec_SelfAttn_4Multi-Head Attention1 × 512 × 768
75Dec_Add_4Add1 × 512 × 768
76Dec_Norm1_4LayerNorm1 × 512 × 768
77Dec_CrossAttn_4Cross-Attention1 × 512 × 768
78Dec_Add2_4Add1 × 512 × 768
79Dec_Norm2_4LayerNorm1 × 512 × 768
80Dec_FFN_4Feed Forward1 × 512 × 768
81Dec_SelfAttn_5Multi-Head Attention1 × 512 × 768
82Dec_Add_5Add1 × 512 × 768
83Dec_Norm1_5LayerNorm1 × 512 × 768
84Dec_CrossAttn_5Cross-Attention1 × 512 × 768
85Dec_Add2_5Add1 × 512 × 768
86Dec_Norm2_5LayerNorm1 × 512 × 768
87Dec_FFN_5Feed Forward1 × 512 × 768
88Dec_SelfAttn_6Multi-Head Attention1 × 512 × 768
89Dec_Add_6Add1 × 512 × 768
90Dec_Norm1_6LayerNorm1 × 512 × 768
91Dec_CrossAttn_6Cross-Attention1 × 512 × 768
92Dec_Add2_6Add1 × 512 × 768
93Dec_Norm2_6LayerNorm1 × 512 × 768
94Dec_FFN_6Feed Forward1 × 512 × 768
95Dec_SelfAttn_7Multi-Head Attention1 × 512 × 768
96Dec_Add_7Add1 × 512 × 768
97Dec_Norm1_7LayerNorm1 × 512 × 768
98Dec_CrossAttn_7Cross-Attention1 × 512 × 768
99Dec_Add2_7Add1 × 512 × 768
100Dec_Norm2_7LayerNorm1 × 512 × 768
101Dec_FFN_7Feed Forward1 × 512 × 768
102Dec_SelfAttn_8Multi-Head Attention1 × 512 × 768
103Dec_Add_8Add1 × 512 × 768
104Dec_Norm1_8LayerNorm1 × 512 × 768
105Dec_CrossAttn_8Cross-Attention1 × 512 × 768
106Dec_Add2_8Add1 × 512 × 768
107Dec_Norm2_8LayerNorm1 × 512 × 768
108Dec_FFN_8Feed Forward1 × 512 × 768
109Dec_SelfAttn_9Multi-Head Attention1 × 512 × 768
110Dec_Add_9Add1 × 512 × 768
111Dec_Norm1_9LayerNorm1 × 512 × 768
112Dec_CrossAttn_9Cross-Attention1 × 512 × 768
113Dec_Add2_9Add1 × 512 × 768
114Dec_Norm2_9LayerNorm1 × 512 × 768
115Dec_FFN_9Feed Forward1 × 512 × 768
116Dec_SelfAttn_10Multi-Head Attention1 × 512 × 768
117Dec_Add_10Add1 × 512 × 768
118Dec_Norm1_10LayerNorm1 × 512 × 768
119Dec_CrossAttn_10Cross-Attention1 × 512 × 768
120Dec_Add2_10Add1 × 512 × 768
121Dec_Norm2_10LayerNorm1 × 512 × 768
122Dec_FFN_10Feed Forward1 × 512 × 768
123Dec_SelfAttn_11Multi-Head Attention1 × 512 × 768
124Dec_Add_11Add1 × 512 × 768
125Dec_Norm1_11LayerNorm1 × 512 × 768
126Dec_CrossAttn_11Cross-Attention1 × 512 × 768
127Dec_Add2_11Add1 × 512 × 768
128Dec_Norm2_11LayerNorm1 × 512 × 768
129Dec_FFN_11Feed Forward1 × 512 × 768
130Dec_SelfAttn_12Multi-Head Attention1 × 512 × 768
131Dec_Add_12Add1 × 512 × 768
132Dec_Norm1_12LayerNorm1 × 512 × 768
133Dec_CrossAttn_12Cross-Attention1 × 512 × 768
134Dec_Add2_12Add1 × 512 × 768
135Dec_Norm2_12LayerNorm1 × 512 × 768
136Dec_FFN_12Feed Forward1 × 512 × 768
137LM_HeadLinear1 × 512 × 32128
138OutputOutput1 × 512 × 32128

What the verifier says

infoAt 24 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace google-t5/t5-base --plan --share

Other t5 checkpoints

t5-small
60.6M derived · +0.11% against the checkpoint