N Neurarch Architectures Models Checks Data Docs Open the app

Models / vit

vit-base-patch16-224

Reconstructed from its own config.json with no weights read. 5.1M downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
85.4M
85,357,056 parameters
In the published checkpoint
86.6M
86,567,656 scalars · safetensors.total, read 2026-09-06
Delta
-1.40%

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
74
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$0.98
CardMemory
T4 (16GB)weights + activationsfits
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

76 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 3 × 224 × 224
2PatchEmbeddingEmbedding1 × 3 × 224 × 224 × 768
3Patch_Position_EmbeddingLearned Pos Embed1 × 3 × 224 × 224 × 768
4Attention_1Multi-Head Attention1 × 3 × 224 × 224 × 768
5Add_1Add1 × 3 × 224 × 224 × 768
6LayerNorm_1_1LayerNorm1 × 3 × 224 × 224 × 768
7FFN_1Feed Forward1 × 3 × 224 × 224 × 768
8Add_1_2Add1 × 3 × 224 × 224 × 768
9LayerNorm_1_2LayerNorm1 × 3 × 224 × 224 × 768
10Attention_2Multi-Head Attention1 × 3 × 224 × 224 × 768
11Add_2Add1 × 3 × 224 × 224 × 768
12LayerNorm_2_1LayerNorm1 × 3 × 224 × 224 × 768
13FFN_2Feed Forward1 × 3 × 224 × 224 × 768
14Add_2_2Add1 × 3 × 224 × 224 × 768
15LayerNorm_2_2LayerNorm1 × 3 × 224 × 224 × 768
16Attention_3Multi-Head Attention1 × 3 × 224 × 224 × 768
17Add_3Add1 × 3 × 224 × 224 × 768
18LayerNorm_3_1LayerNorm1 × 3 × 224 × 224 × 768
19FFN_3Feed Forward1 × 3 × 224 × 224 × 768
20Add_3_2Add1 × 3 × 224 × 224 × 768
21LayerNorm_3_2LayerNorm1 × 3 × 224 × 224 × 768
22Attention_4Multi-Head Attention1 × 3 × 224 × 224 × 768
23Add_4Add1 × 3 × 224 × 224 × 768
24LayerNorm_4_1LayerNorm1 × 3 × 224 × 224 × 768
25FFN_4Feed Forward1 × 3 × 224 × 224 × 768
26Add_4_2Add1 × 3 × 224 × 224 × 768
27LayerNorm_4_2LayerNorm1 × 3 × 224 × 224 × 768
28Attention_5Multi-Head Attention1 × 3 × 224 × 224 × 768
29Add_5Add1 × 3 × 224 × 224 × 768
30LayerNorm_5_1LayerNorm1 × 3 × 224 × 224 × 768
31FFN_5Feed Forward1 × 3 × 224 × 224 × 768
32Add_5_2Add1 × 3 × 224 × 224 × 768
33LayerNorm_5_2LayerNorm1 × 3 × 224 × 224 × 768
34Attention_6Multi-Head Attention1 × 3 × 224 × 224 × 768
35Add_6Add1 × 3 × 224 × 224 × 768
36LayerNorm_6_1LayerNorm1 × 3 × 224 × 224 × 768
37FFN_6Feed Forward1 × 3 × 224 × 224 × 768
38Add_6_2Add1 × 3 × 224 × 224 × 768
39LayerNorm_6_2LayerNorm1 × 3 × 224 × 224 × 768
40Attention_7Multi-Head Attention1 × 3 × 224 × 224 × 768
41Add_7Add1 × 3 × 224 × 224 × 768
42LayerNorm_7_1LayerNorm1 × 3 × 224 × 224 × 768
43FFN_7Feed Forward1 × 3 × 224 × 224 × 768
44Add_7_2Add1 × 3 × 224 × 224 × 768
45LayerNorm_7_2LayerNorm1 × 3 × 224 × 224 × 768
46Attention_8Multi-Head Attention1 × 3 × 224 × 224 × 768
47Add_8Add1 × 3 × 224 × 224 × 768
48LayerNorm_8_1LayerNorm1 × 3 × 224 × 224 × 768
49FFN_8Feed Forward1 × 3 × 224 × 224 × 768
50Add_8_2Add1 × 3 × 224 × 224 × 768
51LayerNorm_8_2LayerNorm1 × 3 × 224 × 224 × 768
52Attention_9Multi-Head Attention1 × 3 × 224 × 224 × 768
53Add_9Add1 × 3 × 224 × 224 × 768
54LayerNorm_9_1LayerNorm1 × 3 × 224 × 224 × 768
55FFN_9Feed Forward1 × 3 × 224 × 224 × 768
56Add_9_2Add1 × 3 × 224 × 224 × 768
57LayerNorm_9_2LayerNorm1 × 3 × 224 × 224 × 768
58Attention_10Multi-Head Attention1 × 3 × 224 × 224 × 768
59Add_10Add1 × 3 × 224 × 224 × 768
60LayerNorm_10_1LayerNorm1 × 3 × 224 × 224 × 768
61FFN_10Feed Forward1 × 3 × 224 × 224 × 768
62Add_10_2Add1 × 3 × 224 × 224 × 768
63LayerNorm_10_2LayerNorm1 × 3 × 224 × 224 × 768
64Attention_11Multi-Head Attention1 × 3 × 224 × 224 × 768
65Add_11Add1 × 3 × 224 × 224 × 768
66LayerNorm_11_1LayerNorm1 × 3 × 224 × 224 × 768
67FFN_11Feed Forward1 × 3 × 224 × 224 × 768
68Add_11_2Add1 × 3 × 224 × 224 × 768
69LayerNorm_11_2LayerNorm1 × 3 × 224 × 224 × 768
70Attention_12Multi-Head Attention1 × 3 × 224 × 224 × 768
71Add_12Add1 × 3 × 224 × 224 × 768
72LayerNorm_12_1LayerNorm1 × 3 × 224 × 224 × 768
73FFN_12Feed Forward1 × 3 × 224 × 224 × 768
74Add_12_2Add1 × 3 × 224 × 224 × 768
75LayerNorm_12_2LayerNorm1 × 3 × 224 × 224 × 768
76OutputOutput1 × 3 × 224 × 224 × 768

What the verifier says

warn"LayerNorm_12_2" (layerNorm) is the last layer before Output. Normalizing the raw logits constrains the output range and breaks standard loss functions. Fix: Move normalization before the final Linear/Conv layer.
bn-at-output
infoAt 12 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace google/vit-base-patch16-224 --plan --share