N Neurarch Architectures Models Checks Data Docs Open the app

Models / whisper

whisper-large-v3-turbo

Reconstructed from its own config.json with no weights read. 7.1M downloads on Hugging Face.

Our count against the checkpoint

The left number comes from the graph. The right one is the number of scalars in the published weight files. Nothing on this page was tuned to make them agree.

Derived from structure
801M
801,484,800 parameters
In the published checkpoint
809M
808,878,080 scalars · safetensors.total, read 2026-09-06
Delta
-0.91%

What it costs to run

Cost is a roofline estimate on the priced GPU for 10 epochs at batch 32 over 50,000 samples (assumed; no dataset attached). GPU fit is fp32 weights plus gradients plus two Adam moments (16 bytes per parameter) with 1.3x headroom; activations are not included and grow with batch size.

Layers
121
Will it forward-pass
Yes
Priced on
A10G (24GB)
Est. one run
$2.54
CardMemory
T4 (16GB)weights + activationsdoes not fit
A100 (40GB)weights + activationsfits
H100 (80GB)weights + activationsfits

Structure

124 nodes. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeOutput shape
1InputInput1 × 512
2MelSpectrogramMel Spectrogram1 × 80 × 1
3Enc_Conv1Conv1D1 × 1280 × 1
4Enc_Conv2Conv1D1 × 1280 × 1
5Enc_ToTokensPermute1 × 1 × 1280
6Enc_PosEncPositional Encoding1 × 1 × 1280
7Enc_Attn_1Multi-Head Attention1 × 1 × 1280
8Enc_Norm_1LayerNorm1 × 1 × 1280
9Enc_FFN_1Feed Forward1 × 1 × 1280
10Enc_Attn_2Multi-Head Attention1 × 1 × 1280
11Enc_Norm_2LayerNorm1 × 1 × 1280
12Enc_FFN_2Feed Forward1 × 1 × 1280
13Enc_Attn_3Multi-Head Attention1 × 1 × 1280
14Enc_Norm_3LayerNorm1 × 1 × 1280
15Enc_FFN_3Feed Forward1 × 1 × 1280
16Enc_Attn_4Multi-Head Attention1 × 1 × 1280
17Enc_Norm_4LayerNorm1 × 1 × 1280
18Enc_FFN_4Feed Forward1 × 1 × 1280
19Enc_Attn_5Multi-Head Attention1 × 1 × 1280
20Enc_Norm_5LayerNorm1 × 1 × 1280
21Enc_FFN_5Feed Forward1 × 1 × 1280
22Enc_Attn_6Multi-Head Attention1 × 1 × 1280
23Enc_Norm_6LayerNorm1 × 1 × 1280
24Enc_FFN_6Feed Forward1 × 1 × 1280
25Enc_Attn_7Multi-Head Attention1 × 1 × 1280
26Enc_Norm_7LayerNorm1 × 1 × 1280
27Enc_FFN_7Feed Forward1 × 1 × 1280
28Enc_Attn_8Multi-Head Attention1 × 1 × 1280
29Enc_Norm_8LayerNorm1 × 1 × 1280
30Enc_FFN_8Feed Forward1 × 1 × 1280
31Enc_Attn_9Multi-Head Attention1 × 1 × 1280
32Enc_Norm_9LayerNorm1 × 1 × 1280
33Enc_FFN_9Feed Forward1 × 1 × 1280
34Enc_Attn_10Multi-Head Attention1 × 1 × 1280
35Enc_Norm_10LayerNorm1 × 1 × 1280
36Enc_FFN_10Feed Forward1 × 1 × 1280
37Enc_Attn_11Multi-Head Attention1 × 1 × 1280
38Enc_Norm_11LayerNorm1 × 1 × 1280
39Enc_FFN_11Feed Forward1 × 1 × 1280
40Enc_Attn_12Multi-Head Attention1 × 1 × 1280
41Enc_Norm_12LayerNorm1 × 1 × 1280
42Enc_FFN_12Feed Forward1 × 1 × 1280
43Enc_Attn_13Multi-Head Attention1 × 1 × 1280
44Enc_Norm_13LayerNorm1 × 1 × 1280
45Enc_FFN_13Feed Forward1 × 1 × 1280
46Enc_Attn_14Multi-Head Attention1 × 1 × 1280
47Enc_Norm_14LayerNorm1 × 1 × 1280
48Enc_FFN_14Feed Forward1 × 1 × 1280
49Enc_Attn_15Multi-Head Attention1 × 1 × 1280
50Enc_Norm_15LayerNorm1 × 1 × 1280
51Enc_FFN_15Feed Forward1 × 1 × 1280
52Enc_Attn_16Multi-Head Attention1 × 1 × 1280
53Enc_Norm_16LayerNorm1 × 1 × 1280
54Enc_FFN_16Feed Forward1 × 1 × 1280
55Enc_Attn_17Multi-Head Attention1 × 1 × 1280
56Enc_Norm_17LayerNorm1 × 1 × 1280
57Enc_FFN_17Feed Forward1 × 1 × 1280
58Enc_Attn_18Multi-Head Attention1 × 1 × 1280
59Enc_Norm_18LayerNorm1 × 1 × 1280
60Enc_FFN_18Feed Forward1 × 1 × 1280
61Enc_Attn_19Multi-Head Attention1 × 1 × 1280
62Enc_Norm_19LayerNorm1 × 1 × 1280
63Enc_FFN_19Feed Forward1 × 1 × 1280
64Enc_Attn_20Multi-Head Attention1 × 1 × 1280
65Enc_Norm_20LayerNorm1 × 1 × 1280
66Enc_FFN_20Feed Forward1 × 1 × 1280
67Enc_Attn_21Multi-Head Attention1 × 1 × 1280
68Enc_Norm_21LayerNorm1 × 1 × 1280
69Enc_FFN_21Feed Forward1 × 1 × 1280
70Enc_Attn_22Multi-Head Attention1 × 1 × 1280
71Enc_Norm_22LayerNorm1 × 1 × 1280
72Enc_FFN_22Feed Forward1 × 1 × 1280
73Enc_Attn_23Multi-Head Attention1 × 1 × 1280
74Enc_Norm_23LayerNorm1 × 1 × 1280
75Enc_FFN_23Feed Forward1 × 1 × 1280
76Enc_Attn_24Multi-Head Attention1 × 1 × 1280
77Enc_Norm_24LayerNorm1 × 1 × 1280
78Enc_FFN_24Feed Forward1 × 1 × 1280
79Enc_Attn_25Multi-Head Attention1 × 1 × 1280
80Enc_Norm_25LayerNorm1 × 1 × 1280
81Enc_FFN_25Feed Forward1 × 1 × 1280
82Enc_Attn_26Multi-Head Attention1 × 1 × 1280
83Enc_Norm_26LayerNorm1 × 1 × 1280
84Enc_FFN_26Feed Forward1 × 1 × 1280
85Enc_Attn_27Multi-Head Attention1 × 1 × 1280
86Enc_Norm_27LayerNorm1 × 1 × 1280
87Enc_FFN_27Feed Forward1 × 1 × 1280
88Enc_Attn_28Multi-Head Attention1 × 1 × 1280
89Enc_Norm_28LayerNorm1 × 1 × 1280
90Enc_FFN_28Feed Forward1 × 1 × 1280
91Enc_Attn_29Multi-Head Attention1 × 1 × 1280
92Enc_Norm_29LayerNorm1 × 1 × 1280
93Enc_FFN_29Feed Forward1 × 1 × 1280
94Enc_Attn_30Multi-Head Attention1 × 1 × 1280
95Enc_Norm_30LayerNorm1 × 1 × 1280
96Enc_FFN_30Feed Forward1 × 1 × 1280
97Enc_Attn_31Multi-Head Attention1 × 1 × 1280
98Enc_Norm_31LayerNorm1 × 1 × 1280
99Enc_FFN_31Feed Forward1 × 1 × 1280
100Enc_Attn_32Multi-Head Attention1 × 1 × 1280
101Enc_Norm_32LayerNorm1 × 1 × 1280
102Enc_FFN_32Feed Forward1 × 1 × 1280
103decoder_tokensInput1 × 448
104Dec_Token_EmbeddingEmbedding1 × 448 × 1280
105Dec_PosEmbeddingLearned Pos Embed1 × 448 × 1280
106Dec_SelfAttn_1Causal Attention1 × 448 × 1280
107Dec_CrossAttn_1Cross-Attention1 × 448 × 1280
108Dec_Norm_1LayerNorm1 × 448 × 1280
109Dec_FFN_1Feed Forward1 × 448 × 1280
110Dec_SelfAttn_2Causal Attention1 × 448 × 1280
111Dec_CrossAttn_2Cross-Attention1 × 448 × 1280
112Dec_Norm_2LayerNorm1 × 448 × 1280
113Dec_FFN_2Feed Forward1 × 448 × 1280
114Dec_SelfAttn_3Causal Attention1 × 448 × 1280
115Dec_CrossAttn_3Cross-Attention1 × 448 × 1280
116Dec_Norm_3LayerNorm1 × 448 × 1280
117Dec_FFN_3Feed Forward1 × 448 × 1280
118Dec_SelfAttn_4Causal Attention1 × 448 × 1280
119Dec_CrossAttn_4Cross-Attention1 × 448 × 1280
120Dec_Norm_4LayerNorm1 × 448 × 1280
121Dec_FFN_4Feed Forward1 × 448 × 1280
122Dec_Norm_FinalLayerNorm1 × 448 × 1280
123LM_HeadLinear1 × 448 × 51866
124OutputOutput1 × 448 × 51866

What the verifier says

infoAt 36 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / √(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers))
deep-attention-default-init

Do this to your own model

Same numbers, on a model in your repo, in one command. No account.

pip install neurarch-trace
neurarch-trace openai/whisper-large-v3-turbo --plan --share