N Neurarch Architectures Models Checks Data Docs Open the app

Models / unlimited-ocr

Unlimited-OCR

Reconstructed from its own config.json with no weights read. 2.7M downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
2.81B
2,810,339,008 parameters
In the published checkpoint
3.34B
3,336,106,240 scalars · safetensors.total, read 2026-07-29
Delta
-15.8%

custom-code This repository ships its own modeling code (`auto_map`, e.g. `modeling_unlimitedocr.py`), so `config.json` names a class in the repo rather than an architecture `transformers` defines. The graph below is what those config keys mean under `transformers` semantics, which is not necessarily what the repo's own file builds. A gap here is a statement about what we read, not about the checkpoint.

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
79
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$1333.30
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsdoes not fit
H100 (80GB)weights + activationsfits

Structure

82 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 32768
2EmbeddingEmbedding1 × 32768 × 1280
3Positional_EmbeddingLearned Pos Embed1 × 32768 × 1280
4Vision inputInput3 × 1024 × 1024
5PatchEmbedPatch Embed4096 × 1280
6Patch_Position_EmbeddingLearned Pos Embed4096 × 1280
7Vision encoder (internals not in config)Projection4096 × 1280
8Vision tokensReshape1 × 4096 × 1280
9Multimodal fusion (concat tokens)Concatenate1 × 36864 × 1280
10LayerNorm_1_1LayerNorm1 × 36864 × 1280
11Attention_1Grouped Query Attn1 × 36864 × 1280
12Add_1_attnAdd1 × 36864 × 1280
13LayerNorm_1_2LayerNorm1 × 36864 × 1280
14FFN_1Feed Forward1 × 36864 × 1280
15Add_1_ffnAdd1 × 36864 × 1280
16LayerNorm_2_1LayerNorm1 × 36864 × 1280
17Attention_2Grouped Query Attn1 × 36864 × 1280
18Add_2_attnAdd1 × 36864 × 1280
19LayerNorm_2_2LayerNorm1 × 36864 × 1280
20MoE_2Shared-Expert MoE1 × 36864 × 1280
21Add_2_ffnAdd1 × 36864 × 1280
22LayerNorm_3_1LayerNorm1 × 36864 × 1280
23Attention_3Grouped Query Attn1 × 36864 × 1280
24Add_3_attnAdd1 × 36864 × 1280
25LayerNorm_3_2LayerNorm1 × 36864 × 1280
26MoE_3Shared-Expert MoE1 × 36864 × 1280
27Add_3_ffnAdd1 × 36864 × 1280
28LayerNorm_4_1LayerNorm1 × 36864 × 1280
29Attention_4Grouped Query Attn1 × 36864 × 1280
30Add_4_attnAdd1 × 36864 × 1280
31LayerNorm_4_2LayerNorm1 × 36864 × 1280
32MoE_4Shared-Expert MoE1 × 36864 × 1280
33Add_4_ffnAdd1 × 36864 × 1280
34LayerNorm_5_1LayerNorm1 × 36864 × 1280
35Attention_5Grouped Query Attn1 × 36864 × 1280
36Add_5_attnAdd1 × 36864 × 1280
37LayerNorm_5_2LayerNorm1 × 36864 × 1280
38MoE_5Shared-Expert MoE1 × 36864 × 1280
39Add_5_ffnAdd1 × 36864 × 1280
40LayerNorm_6_1LayerNorm1 × 36864 × 1280
41Attention_6Grouped Query Attn1 × 36864 × 1280
42Add_6_attnAdd1 × 36864 × 1280
43LayerNorm_6_2LayerNorm1 × 36864 × 1280
44MoE_6Shared-Expert MoE1 × 36864 × 1280
45Add_6_ffnAdd1 × 36864 × 1280
46LayerNorm_7_1LayerNorm1 × 36864 × 1280
47Attention_7Grouped Query Attn1 × 36864 × 1280
48Add_7_attnAdd1 × 36864 × 1280
49LayerNorm_7_2LayerNorm1 × 36864 × 1280
50MoE_7Shared-Expert MoE1 × 36864 × 1280
51Add_7_ffnAdd1 × 36864 × 1280
52LayerNorm_8_1LayerNorm1 × 36864 × 1280
53Attention_8Grouped Query Attn1 × 36864 × 1280
54Add_8_attnAdd1 × 36864 × 1280
55LayerNorm_8_2LayerNorm1 × 36864 × 1280
56MoE_8Shared-Expert MoE1 × 36864 × 1280
57Add_8_ffnAdd1 × 36864 × 1280
58LayerNorm_9_1LayerNorm1 × 36864 × 1280
59Attention_9Grouped Query Attn1 × 36864 × 1280
60Add_9_attnAdd1 × 36864 × 1280
61LayerNorm_9_2LayerNorm1 × 36864 × 1280
62MoE_9Shared-Expert MoE1 × 36864 × 1280
63Add_9_ffnAdd1 × 36864 × 1280
64LayerNorm_10_1LayerNorm1 × 36864 × 1280
65Attention_10Grouped Query Attn1 × 36864 × 1280
66Add_10_attnAdd1 × 36864 × 1280
67LayerNorm_10_2LayerNorm1 × 36864 × 1280
68MoE_10Shared-Expert MoE1 × 36864 × 1280
69Add_10_ffnAdd1 × 36864 × 1280
70LayerNorm_11_1LayerNorm1 × 36864 × 1280
71Attention_11Grouped Query Attn1 × 36864 × 1280
72Add_11_attnAdd1 × 36864 × 1280
73LayerNorm_11_2LayerNorm1 × 36864 × 1280
74MoE_11Shared-Expert MoE1 × 36864 × 1280
75Add_11_ffnAdd1 × 36864 × 1280
76LayerNorm_12_1LayerNorm1 × 36864 × 1280
77Attention_12Grouped Query Attn1 × 36864 × 1280
78Add_12_attnAdd1 × 36864 × 1280
79LayerNorm_12_2LayerNorm1 × 36864 × 1280
80MoE_12Shared-Expert MoE1 × 36864 × 1280
81Add_12_ffnAdd1 × 36864 × 1280
82OutputOutput1 × 36864 × 1280

What the verifier says

warn"Positional_Embedding" receives input but its output is not connected. This layer will be unreachable in the forward pass. Fix: Connect the output forward, or add an Output node if this is the final layer.
dead-end
infoAt 12 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace baidu/Unlimited-OCR --plan --share

Other unlimited-ocr checkpoints

Unlimited-OCR-AWQ
2.94B derived · -12.3% against the checkpoint