N Neurarch Architectures Models Checks Data Docs Open the app

Models / florence2

Florence-2-large

Reconstructed from its own config.json with no weights read. 618K downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
208M
207,844,352 parameters
In the published checkpoint
777M
776,721,497 scalars · safetensors.total, read 2025-08-04
Delta
-73.2%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `configuration_florence2.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
50
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$30.91
CardMemory
T4 (16GB)weights + activationsfits
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

52 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 4096
2EmbeddingEmbedding1 × 4096 × 1024
3Positional_EmbeddingLearned Pos Embed1 × 4096 × 1024
4Attention_1Multi-Head Attention1 × 4096 × 1024
5Add_1Add1 × 4096 × 1024
6LayerNorm_1_1LayerNorm1 × 4096 × 1024
7FFN_1Feed Forward1 × 4096 × 1024
8Attention_2Multi-Head Attention1 × 4096 × 1024
9Add_2Add1 × 4096 × 1024
10LayerNorm_2_1LayerNorm1 × 4096 × 1024
11FFN_2Feed Forward1 × 4096 × 1024
12Attention_3Multi-Head Attention1 × 4096 × 1024
13Add_3Add1 × 4096 × 1024
14LayerNorm_3_1LayerNorm1 × 4096 × 1024
15FFN_3Feed Forward1 × 4096 × 1024
16Attention_4Multi-Head Attention1 × 4096 × 1024
17Add_4Add1 × 4096 × 1024
18LayerNorm_4_1LayerNorm1 × 4096 × 1024
19FFN_4Feed Forward1 × 4096 × 1024
20Attention_5Multi-Head Attention1 × 4096 × 1024
21Add_5Add1 × 4096 × 1024
22LayerNorm_5_1LayerNorm1 × 4096 × 1024
23FFN_5Feed Forward1 × 4096 × 1024
24Attention_6Multi-Head Attention1 × 4096 × 1024
25Add_6Add1 × 4096 × 1024
26LayerNorm_6_1LayerNorm1 × 4096 × 1024
27FFN_6Feed Forward1 × 4096 × 1024
28Attention_7Multi-Head Attention1 × 4096 × 1024
29Add_7Add1 × 4096 × 1024
30LayerNorm_7_1LayerNorm1 × 4096 × 1024
31FFN_7Feed Forward1 × 4096 × 1024
32Attention_8Multi-Head Attention1 × 4096 × 1024
33Add_8Add1 × 4096 × 1024
34LayerNorm_8_1LayerNorm1 × 4096 × 1024
35FFN_8Feed Forward1 × 4096 × 1024
36Attention_9Multi-Head Attention1 × 4096 × 1024
37Add_9Add1 × 4096 × 1024
38LayerNorm_9_1LayerNorm1 × 4096 × 1024
39FFN_9Feed Forward1 × 4096 × 1024
40Attention_10Multi-Head Attention1 × 4096 × 1024
41Add_10Add1 × 4096 × 1024
42LayerNorm_10_1LayerNorm1 × 4096 × 1024
43FFN_10Feed Forward1 × 4096 × 1024
44Attention_11Multi-Head Attention1 × 4096 × 1024
45Add_11Add1 × 4096 × 1024
46LayerNorm_11_1LayerNorm1 × 4096 × 1024
47FFN_11Feed Forward1 × 4096 × 1024
48Attention_12Multi-Head Attention1 × 4096 × 1024
49Add_12Add1 × 4096 × 1024
50LayerNorm_12_1LayerNorm1 × 4096 × 1024
51FFN_12Feed Forward1 × 4096 × 1024
52OutputOutput1 × 4096 × 1024

What the verifier says

infoAt 12 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace microsoft/Florence-2-large --plan --share

Other florence2 checkpoints

Florence-2-base
82.7M derived · -64.3% against the checkpoint
Florence-2-base-ft
82.7M derived · -64.3% against the checkpoint
Florence-2-VQAJP2
82.7M derived · -69.5% against the checkpoint