N Neurarch Architectures Models Checks Data Docs Open the app

Models / granite4_vision

granite-vision-4.1-4b

Reconstructed from its own config.json with no weights read. 115K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
3.32B
3,319,787,520 parameters
In the published checkpoint
4.00B
3,997,203,904 scalars · safetensors.total, read 2026-07-13
Delta
-16.9%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `configuration.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
162
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$79217.40
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsfits

Structure

164 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 131072
2EmbeddingEmbedding1 × 131072 × 2560
3Positional_EmbeddingLearned Pos Embed1 × 131072 × 2560
4Attention_1Multi-Head Attention1 × 131072 × 2560
5Add_1Add1 × 131072 × 2560
6LayerNorm_1_1LayerNorm1 × 131072 × 2560
7FFN_1Feed Forward1 × 131072 × 2560
8Attention_2Multi-Head Attention1 × 131072 × 2560
9Add_2Add1 × 131072 × 2560
10LayerNorm_2_1LayerNorm1 × 131072 × 2560
11FFN_2Feed Forward1 × 131072 × 2560
12Attention_3Multi-Head Attention1 × 131072 × 2560
13Add_3Add1 × 131072 × 2560
14LayerNorm_3_1LayerNorm1 × 131072 × 2560
15FFN_3Feed Forward1 × 131072 × 2560
16Attention_4Multi-Head Attention1 × 131072 × 2560
17Add_4Add1 × 131072 × 2560
18LayerNorm_4_1LayerNorm1 × 131072 × 2560
19FFN_4Feed Forward1 × 131072 × 2560
20Attention_5Multi-Head Attention1 × 131072 × 2560
21Add_5Add1 × 131072 × 2560
22LayerNorm_5_1LayerNorm1 × 131072 × 2560
23FFN_5Feed Forward1 × 131072 × 2560
24Attention_6Multi-Head Attention1 × 131072 × 2560
25Add_6Add1 × 131072 × 2560
26LayerNorm_6_1LayerNorm1 × 131072 × 2560
27FFN_6Feed Forward1 × 131072 × 2560
28Attention_7Multi-Head Attention1 × 131072 × 2560
29Add_7Add1 × 131072 × 2560
30LayerNorm_7_1LayerNorm1 × 131072 × 2560
31FFN_7Feed Forward1 × 131072 × 2560
32Attention_8Multi-Head Attention1 × 131072 × 2560
33Add_8Add1 × 131072 × 2560
34LayerNorm_8_1LayerNorm1 × 131072 × 2560
35FFN_8Feed Forward1 × 131072 × 2560
36Attention_9Multi-Head Attention1 × 131072 × 2560
37Add_9Add1 × 131072 × 2560
38LayerNorm_9_1LayerNorm1 × 131072 × 2560
39FFN_9Feed Forward1 × 131072 × 2560
40Attention_10Multi-Head Attention1 × 131072 × 2560
41Add_10Add1 × 131072 × 2560
42LayerNorm_10_1LayerNorm1 × 131072 × 2560
43FFN_10Feed Forward1 × 131072 × 2560
44Attention_11Multi-Head Attention1 × 131072 × 2560
45Add_11Add1 × 131072 × 2560
46LayerNorm_11_1LayerNorm1 × 131072 × 2560
47FFN_11Feed Forward1 × 131072 × 2560
48Attention_12Multi-Head Attention1 × 131072 × 2560
49Add_12Add1 × 131072 × 2560
50LayerNorm_12_1LayerNorm1 × 131072 × 2560
51FFN_12Feed Forward1 × 131072 × 2560
52Attention_13Multi-Head Attention1 × 131072 × 2560
53Add_13Add1 × 131072 × 2560
54LayerNorm_13_1LayerNorm1 × 131072 × 2560
55FFN_13Feed Forward1 × 131072 × 2560
56Attention_14Multi-Head Attention1 × 131072 × 2560
57Add_14Add1 × 131072 × 2560
58LayerNorm_14_1LayerNorm1 × 131072 × 2560
59FFN_14Feed Forward1 × 131072 × 2560
60Attention_15Multi-Head Attention1 × 131072 × 2560
61Add_15Add1 × 131072 × 2560
62LayerNorm_15_1LayerNorm1 × 131072 × 2560
63FFN_15Feed Forward1 × 131072 × 2560
64Attention_16Multi-Head Attention1 × 131072 × 2560
65Add_16Add1 × 131072 × 2560
66LayerNorm_16_1LayerNorm1 × 131072 × 2560
67FFN_16Feed Forward1 × 131072 × 2560
68Attention_17Multi-Head Attention1 × 131072 × 2560
69Add_17Add1 × 131072 × 2560
70LayerNorm_17_1LayerNorm1 × 131072 × 2560
71FFN_17Feed Forward1 × 131072 × 2560
72Attention_18Multi-Head Attention1 × 131072 × 2560
73Add_18Add1 × 131072 × 2560
74LayerNorm_18_1LayerNorm1 × 131072 × 2560
75FFN_18Feed Forward1 × 131072 × 2560
76Attention_19Multi-Head Attention1 × 131072 × 2560
77Add_19Add1 × 131072 × 2560
78LayerNorm_19_1LayerNorm1 × 131072 × 2560
79FFN_19Feed Forward1 × 131072 × 2560
80Attention_20Multi-Head Attention1 × 131072 × 2560
81Add_20Add1 × 131072 × 2560
82LayerNorm_20_1LayerNorm1 × 131072 × 2560
83FFN_20Feed Forward1 × 131072 × 2560
84Attention_21Multi-Head Attention1 × 131072 × 2560
85Add_21Add1 × 131072 × 2560
86LayerNorm_21_1LayerNorm1 × 131072 × 2560
87FFN_21Feed Forward1 × 131072 × 2560
88Attention_22Multi-Head Attention1 × 131072 × 2560
89Add_22Add1 × 131072 × 2560
90LayerNorm_22_1LayerNorm1 × 131072 × 2560
91FFN_22Feed Forward1 × 131072 × 2560
92Attention_23Multi-Head Attention1 × 131072 × 2560
93Add_23Add1 × 131072 × 2560
94LayerNorm_23_1LayerNorm1 × 131072 × 2560
95FFN_23Feed Forward1 × 131072 × 2560
96Attention_24Multi-Head Attention1 × 131072 × 2560
97Add_24Add1 × 131072 × 2560
98LayerNorm_24_1LayerNorm1 × 131072 × 2560
99FFN_24Feed Forward1 × 131072 × 2560
100Attention_25Multi-Head Attention1 × 131072 × 2560
101Add_25Add1 × 131072 × 2560
102LayerNorm_25_1LayerNorm1 × 131072 × 2560
103FFN_25Feed Forward1 × 131072 × 2560
104Attention_26Multi-Head Attention1 × 131072 × 2560
105Add_26Add1 × 131072 × 2560
106LayerNorm_26_1LayerNorm1 × 131072 × 2560
107FFN_26Feed Forward1 × 131072 × 2560
108Attention_27Multi-Head Attention1 × 131072 × 2560
109Add_27Add1 × 131072 × 2560
110LayerNorm_27_1LayerNorm1 × 131072 × 2560
111FFN_27Feed Forward1 × 131072 × 2560
112Attention_28Multi-Head Attention1 × 131072 × 2560
113Add_28Add1 × 131072 × 2560
114LayerNorm_28_1LayerNorm1 × 131072 × 2560
115FFN_28Feed Forward1 × 131072 × 2560
116Attention_29Multi-Head Attention1 × 131072 × 2560
117Add_29Add1 × 131072 × 2560
118LayerNorm_29_1LayerNorm1 × 131072 × 2560
119FFN_29Feed Forward1 × 131072 × 2560
120Attention_30Multi-Head Attention1 × 131072 × 2560
121Add_30Add1 × 131072 × 2560
122LayerNorm_30_1LayerNorm1 × 131072 × 2560
123FFN_30Feed Forward1 × 131072 × 2560
124Attention_31Multi-Head Attention1 × 131072 × 2560
125Add_31Add1 × 131072 × 2560
126LayerNorm_31_1LayerNorm1 × 131072 × 2560
127FFN_31Feed Forward1 × 131072 × 2560
128Attention_32Multi-Head Attention1 × 131072 × 2560
129Add_32Add1 × 131072 × 2560
130LayerNorm_32_1LayerNorm1 × 131072 × 2560
131FFN_32Feed Forward1 × 131072 × 2560
132Attention_33Multi-Head Attention1 × 131072 × 2560
133Add_33Add1 × 131072 × 2560
134LayerNorm_33_1LayerNorm1 × 131072 × 2560
135FFN_33Feed Forward1 × 131072 × 2560
136Attention_34Multi-Head Attention1 × 131072 × 2560
137Add_34Add1 × 131072 × 2560
138LayerNorm_34_1LayerNorm1 × 131072 × 2560
139FFN_34Feed Forward1 × 131072 × 2560
140Attention_35Multi-Head Attention1 × 131072 × 2560
141Add_35Add1 × 131072 × 2560
142LayerNorm_35_1LayerNorm1 × 131072 × 2560
143FFN_35Feed Forward1 × 131072 × 2560
144Attention_36Multi-Head Attention1 × 131072 × 2560
145Add_36Add1 × 131072 × 2560
146LayerNorm_36_1LayerNorm1 × 131072 × 2560
147FFN_36Feed Forward1 × 131072 × 2560
148Attention_37Multi-Head Attention1 × 131072 × 2560
149Add_37Add1 × 131072 × 2560
150LayerNorm_37_1LayerNorm1 × 131072 × 2560
151FFN_37Feed Forward1 × 131072 × 2560
152Attention_38Multi-Head Attention1 × 131072 × 2560
153Add_38Add1 × 131072 × 2560
154LayerNorm_38_1LayerNorm1 × 131072 × 2560
155FFN_38Feed Forward1 × 131072 × 2560
156Attention_39Multi-Head Attention1 × 131072 × 2560
157Add_39Add1 × 131072 × 2560
158LayerNorm_39_1LayerNorm1 × 131072 × 2560
159FFN_39Feed Forward1 × 131072 × 2560
160Attention_40Multi-Head Attention1 × 131072 × 2560
161Add_40Add1 × 131072 × 2560
162LayerNorm_40_1LayerNorm1 × 131072 × 2560
163FFN_40Feed Forward1 × 131072 × 2560
164OutputOutput1 × 131072 × 2560

What the verifier says

info40 attention layers at embedDim 2560 cache full per-head K/V: about 400 KB per token at fp16, which dominates memory at long context. Grouped-query attention (e.g. 8:1) would cut this ~8×; multi-head latent attention (MLA) shrinks it ~10× or more. This is the move production LLMs make; it does not change the parameter count. Fix: Switch attention to groupedQueryAttention (set numKVHeads below numHeads, e.g. numHeads/4) or mla (a low-rank cached latent).
full-mha-serving-cost
infoAt 40 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace ibm-granite/granite-vision-4.1-4b --plan --share