N Neurarch Architectures Models Checks Data Docs Open the app

Models / falcon

falcon-7b

Reconstructed from its own config.json with no weights read. 362K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
3.84B
3,837,851,648 parameters
In the published checkpoint
7.22B
7,217,189,760 scalars · safetensors.total, read 2024-10-12
Delta
-46.8%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `configuration_falcon.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
194
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$69.25
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsfits

Structure

196 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 512
2EmbeddingEmbedding1 × 512 × 4544
3Positional_EmbeddingLearned Pos Embed1 × 512 × 4544
4LayerNorm_1_1LayerNorm1 × 512 × 4544
5Attention_1Multi-Head Attention1 × 512 × 4544
6Add_1_attnAdd1 × 512 × 4544
7LayerNorm_1_2LayerNorm1 × 512 × 4544
8FFN_1Feed Forward1 × 512 × 4544
9Add_1_ffnAdd1 × 512 × 4544
10LayerNorm_2_1LayerNorm1 × 512 × 4544
11Attention_2Multi-Head Attention1 × 512 × 4544
12Add_2_attnAdd1 × 512 × 4544
13LayerNorm_2_2LayerNorm1 × 512 × 4544
14FFN_2Feed Forward1 × 512 × 4544
15Add_2_ffnAdd1 × 512 × 4544
16LayerNorm_3_1LayerNorm1 × 512 × 4544
17Attention_3Multi-Head Attention1 × 512 × 4544
18Add_3_attnAdd1 × 512 × 4544
19LayerNorm_3_2LayerNorm1 × 512 × 4544
20FFN_3Feed Forward1 × 512 × 4544
21Add_3_ffnAdd1 × 512 × 4544
22LayerNorm_4_1LayerNorm1 × 512 × 4544
23Attention_4Multi-Head Attention1 × 512 × 4544
24Add_4_attnAdd1 × 512 × 4544
25LayerNorm_4_2LayerNorm1 × 512 × 4544
26FFN_4Feed Forward1 × 512 × 4544
27Add_4_ffnAdd1 × 512 × 4544
28LayerNorm_5_1LayerNorm1 × 512 × 4544
29Attention_5Multi-Head Attention1 × 512 × 4544
30Add_5_attnAdd1 × 512 × 4544
31LayerNorm_5_2LayerNorm1 × 512 × 4544
32FFN_5Feed Forward1 × 512 × 4544
33Add_5_ffnAdd1 × 512 × 4544
34LayerNorm_6_1LayerNorm1 × 512 × 4544
35Attention_6Multi-Head Attention1 × 512 × 4544
36Add_6_attnAdd1 × 512 × 4544
37LayerNorm_6_2LayerNorm1 × 512 × 4544
38FFN_6Feed Forward1 × 512 × 4544
39Add_6_ffnAdd1 × 512 × 4544
40LayerNorm_7_1LayerNorm1 × 512 × 4544
41Attention_7Multi-Head Attention1 × 512 × 4544
42Add_7_attnAdd1 × 512 × 4544
43LayerNorm_7_2LayerNorm1 × 512 × 4544
44FFN_7Feed Forward1 × 512 × 4544
45Add_7_ffnAdd1 × 512 × 4544
46LayerNorm_8_1LayerNorm1 × 512 × 4544
47Attention_8Multi-Head Attention1 × 512 × 4544
48Add_8_attnAdd1 × 512 × 4544
49LayerNorm_8_2LayerNorm1 × 512 × 4544
50FFN_8Feed Forward1 × 512 × 4544
51Add_8_ffnAdd1 × 512 × 4544
52LayerNorm_9_1LayerNorm1 × 512 × 4544
53Attention_9Multi-Head Attention1 × 512 × 4544
54Add_9_attnAdd1 × 512 × 4544
55LayerNorm_9_2LayerNorm1 × 512 × 4544
56FFN_9Feed Forward1 × 512 × 4544
57Add_9_ffnAdd1 × 512 × 4544
58LayerNorm_10_1LayerNorm1 × 512 × 4544
59Attention_10Multi-Head Attention1 × 512 × 4544
60Add_10_attnAdd1 × 512 × 4544
61LayerNorm_10_2LayerNorm1 × 512 × 4544
62FFN_10Feed Forward1 × 512 × 4544
63Add_10_ffnAdd1 × 512 × 4544
64LayerNorm_11_1LayerNorm1 × 512 × 4544
65Attention_11Multi-Head Attention1 × 512 × 4544
66Add_11_attnAdd1 × 512 × 4544
67LayerNorm_11_2LayerNorm1 × 512 × 4544
68FFN_11Feed Forward1 × 512 × 4544
69Add_11_ffnAdd1 × 512 × 4544
70LayerNorm_12_1LayerNorm1 × 512 × 4544
71Attention_12Multi-Head Attention1 × 512 × 4544
72Add_12_attnAdd1 × 512 × 4544
73LayerNorm_12_2LayerNorm1 × 512 × 4544
74FFN_12Feed Forward1 × 512 × 4544
75Add_12_ffnAdd1 × 512 × 4544
76LayerNorm_13_1LayerNorm1 × 512 × 4544
77Attention_13Multi-Head Attention1 × 512 × 4544
78Add_13_attnAdd1 × 512 × 4544
79LayerNorm_13_2LayerNorm1 × 512 × 4544
80FFN_13Feed Forward1 × 512 × 4544
81Add_13_ffnAdd1 × 512 × 4544
82LayerNorm_14_1LayerNorm1 × 512 × 4544
83Attention_14Multi-Head Attention1 × 512 × 4544
84Add_14_attnAdd1 × 512 × 4544
85LayerNorm_14_2LayerNorm1 × 512 × 4544
86FFN_14Feed Forward1 × 512 × 4544
87Add_14_ffnAdd1 × 512 × 4544
88LayerNorm_15_1LayerNorm1 × 512 × 4544
89Attention_15Multi-Head Attention1 × 512 × 4544
90Add_15_attnAdd1 × 512 × 4544
91LayerNorm_15_2LayerNorm1 × 512 × 4544
92FFN_15Feed Forward1 × 512 × 4544
93Add_15_ffnAdd1 × 512 × 4544
94LayerNorm_16_1LayerNorm1 × 512 × 4544
95Attention_16Multi-Head Attention1 × 512 × 4544
96Add_16_attnAdd1 × 512 × 4544
97LayerNorm_16_2LayerNorm1 × 512 × 4544
98FFN_16Feed Forward1 × 512 × 4544
99Add_16_ffnAdd1 × 512 × 4544
100LayerNorm_17_1LayerNorm1 × 512 × 4544
101Attention_17Multi-Head Attention1 × 512 × 4544
102Add_17_attnAdd1 × 512 × 4544
103LayerNorm_17_2LayerNorm1 × 512 × 4544
104FFN_17Feed Forward1 × 512 × 4544
105Add_17_ffnAdd1 × 512 × 4544
106LayerNorm_18_1LayerNorm1 × 512 × 4544
107Attention_18Multi-Head Attention1 × 512 × 4544
108Add_18_attnAdd1 × 512 × 4544
109LayerNorm_18_2LayerNorm1 × 512 × 4544
110FFN_18Feed Forward1 × 512 × 4544
111Add_18_ffnAdd1 × 512 × 4544
112LayerNorm_19_1LayerNorm1 × 512 × 4544
113Attention_19Multi-Head Attention1 × 512 × 4544
114Add_19_attnAdd1 × 512 × 4544
115LayerNorm_19_2LayerNorm1 × 512 × 4544
116FFN_19Feed Forward1 × 512 × 4544
117Add_19_ffnAdd1 × 512 × 4544
118LayerNorm_20_1LayerNorm1 × 512 × 4544
119Attention_20Multi-Head Attention1 × 512 × 4544
120Add_20_attnAdd1 × 512 × 4544
121LayerNorm_20_2LayerNorm1 × 512 × 4544
122FFN_20Feed Forward1 × 512 × 4544
123Add_20_ffnAdd1 × 512 × 4544
124LayerNorm_21_1LayerNorm1 × 512 × 4544
125Attention_21Multi-Head Attention1 × 512 × 4544
126Add_21_attnAdd1 × 512 × 4544
127LayerNorm_21_2LayerNorm1 × 512 × 4544
128FFN_21Feed Forward1 × 512 × 4544
129Add_21_ffnAdd1 × 512 × 4544
130LayerNorm_22_1LayerNorm1 × 512 × 4544
131Attention_22Multi-Head Attention1 × 512 × 4544
132Add_22_attnAdd1 × 512 × 4544
133LayerNorm_22_2LayerNorm1 × 512 × 4544
134FFN_22Feed Forward1 × 512 × 4544
135Add_22_ffnAdd1 × 512 × 4544
136LayerNorm_23_1LayerNorm1 × 512 × 4544
137Attention_23Multi-Head Attention1 × 512 × 4544
138Add_23_attnAdd1 × 512 × 4544
139LayerNorm_23_2LayerNorm1 × 512 × 4544
140FFN_23Feed Forward1 × 512 × 4544
141Add_23_ffnAdd1 × 512 × 4544
142LayerNorm_24_1LayerNorm1 × 512 × 4544
143Attention_24Multi-Head Attention1 × 512 × 4544
144Add_24_attnAdd1 × 512 × 4544
145LayerNorm_24_2LayerNorm1 × 512 × 4544
146FFN_24Feed Forward1 × 512 × 4544
147Add_24_ffnAdd1 × 512 × 4544
148LayerNorm_25_1LayerNorm1 × 512 × 4544
149Attention_25Multi-Head Attention1 × 512 × 4544
150Add_25_attnAdd1 × 512 × 4544
151LayerNorm_25_2LayerNorm1 × 512 × 4544
152FFN_25Feed Forward1 × 512 × 4544
153Add_25_ffnAdd1 × 512 × 4544
154LayerNorm_26_1LayerNorm1 × 512 × 4544
155Attention_26Multi-Head Attention1 × 512 × 4544
156Add_26_attnAdd1 × 512 × 4544
157LayerNorm_26_2LayerNorm1 × 512 × 4544
158FFN_26Feed Forward1 × 512 × 4544
159Add_26_ffnAdd1 × 512 × 4544
160LayerNorm_27_1LayerNorm1 × 512 × 4544
161Attention_27Multi-Head Attention1 × 512 × 4544
162Add_27_attnAdd1 × 512 × 4544
163LayerNorm_27_2LayerNorm1 × 512 × 4544
164FFN_27Feed Forward1 × 512 × 4544
165Add_27_ffnAdd1 × 512 × 4544
166LayerNorm_28_1LayerNorm1 × 512 × 4544
167Attention_28Multi-Head Attention1 × 512 × 4544
168Add_28_attnAdd1 × 512 × 4544
169LayerNorm_28_2LayerNorm1 × 512 × 4544
170FFN_28Feed Forward1 × 512 × 4544
171Add_28_ffnAdd1 × 512 × 4544
172LayerNorm_29_1LayerNorm1 × 512 × 4544
173Attention_29Multi-Head Attention1 × 512 × 4544
174Add_29_attnAdd1 × 512 × 4544
175LayerNorm_29_2LayerNorm1 × 512 × 4544
176FFN_29Feed Forward1 × 512 × 4544
177Add_29_ffnAdd1 × 512 × 4544
178LayerNorm_30_1LayerNorm1 × 512 × 4544
179Attention_30Multi-Head Attention1 × 512 × 4544
180Add_30_attnAdd1 × 512 × 4544
181LayerNorm_30_2LayerNorm1 × 512 × 4544
182FFN_30Feed Forward1 × 512 × 4544
183Add_30_ffnAdd1 × 512 × 4544
184LayerNorm_31_1LayerNorm1 × 512 × 4544
185Attention_31Multi-Head Attention1 × 512 × 4544
186Add_31_attnAdd1 × 512 × 4544
187LayerNorm_31_2LayerNorm1 × 512 × 4544
188FFN_31Feed Forward1 × 512 × 4544
189Add_31_ffnAdd1 × 512 × 4544
190LayerNorm_32_1LayerNorm1 × 512 × 4544
191Attention_32Multi-Head Attention1 × 512 × 4544
192Add_32_attnAdd1 × 512 × 4544
193LayerNorm_32_2LayerNorm1 × 512 × 4544
194FFN_32Feed Forward1 × 512 × 4544
195Add_32_ffnAdd1 × 512 × 4544
196OutputOutput1 × 512 × 4544

What the verifier says

info32 attention layers at embedDim 4544 cache full per-head K/V: about 568 KB per token at fp16, which dominates memory at long context. Grouped-query attention (e.g. 8:1) would cut this ~8×; multi-head latent attention (MLA) shrinks it ~10× or more. This is the move production LLMs make; it does not change the parameter count. Fix: Switch attention to groupedQueryAttention (set numKVHeads below numHeads, e.g. numHeads/4) or mla (a low-rank cached latent).
full-mha-serving-cost
infoAt 32 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init
warnAcross 32 attention layers this design caches 568 KB per token, so a single 8,192-token sequence needs ~4.8 GB of KV cache before weights or activations. That exceeds the 4 GB budget this rule assumes for serving headroom. Fix: Cut KV width: raise the GQA ratio (fewer numKVHeads), switch to MLA, reduce depth or embedDim, or accept a shorter serving context.
kv-cache-context-budget

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace tiiuae/falcon-7b --plan --share