N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / Whisper Small vs T5 Small

Whisper Small vs T5 Small

Speech-to-text against text-to-text, both encoder-decoder.

T5 Small has 9.7M more parameters than Whisper Small: 14 layers added, 9 removed, 8 changed.

Baseline

Whisper Small

Layers
15
Parameters
47M
Input
80 × 3000
Output
1 × 448 × 51865
Forward-passes
yes
Est. train cost
$0.705
Compared

T5 Small

Layers
20
Parameters
57M
Input
1 × 512
Output
1 × 128 × 32128
Forward-passes
yes
Est. train cost
$0.207

The deltas

Every number is T5 Small relative to Whisper Small.

Parameters
+9.7M (+20.7%)
Layers
+5
Added
14
Removed
9
Changed
8
Unchanged
1

Which GPUs each one fits

Each side is measured at its own declared input (80 × 3000 against 1 × 512). Both columns are right about their own model; the difference between them is not a fact about the designs.

GPUWhisper SmallT5 Small
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 1 of 32 rows are the same layer with the same parameters.

Hide all 32 rows
Whisper SmallT5 Small
LayerParamsOutputLayerParamsOutput
1changed
shape
mel_features
Input
80 × 3000encoder_ids
Input
1 × 512
2changed
shape
decoder_tokens
Input
1 × 448decoder_ids
Input
1 × 128
3removedconv1
Audio Conv
1.5K384 × 3000
4changed
numEmbeddings, embeddingDim
token_embed
Embedding
20M1 × 448 × 384shared_embed
Embedding
16M1 × 512 × 512
5removedgelu_1
Gelu
384 × 3000
6removeddec_pos_emb
Positional Encoding
1 × 448 × 384
7removedconv2
Audio Conv
1.5K384 × 1500
8removedgelu_2
Gelu
384 × 1500
9removedto_tokens
Permute
1500 × 384
10removedenc_pos_emb
Positional Encoding
1500 × 384
11removedenc_block_1
Transformer Block
1.8M1500 × 384
12addeddec_embed
Embedding
16M1 × 128 × 512
13addedenc_norm
Rms Norm
5121 × 512 × 512
14addeddec_sa_norm
Rms Norm
5121 × 128 × 512
15changed
type, embedDim, numHeads, ffDim
enc_block_2
Transformer Block
1.8M1500 × 384enc_self_attn
Multi Head Attention
1.1M1 × 512 × 512
16addeddec_self_attn
Causal Attention
1.0M1 × 128 × 512
17addedenc_residual
Add
1 × 512 × 512
18addeddec_sa_residual
Add
1 × 128 × 512
19addedenc_ffn_norm
Rms Norm
5121 × 512 × 512
20addeddec_ca_norm
Rms Norm
5121 × 128 × 512
21addedenc_ffn
Feed Forward
2.1M1 × 512 × 512
22addedenc_ffn_residual
Add
1 × 512 × 512
23changed
normalizedShape
enc_norm
Layer Norm
7681500 × 384enc_out_norm
Layer Norm
1.0K1 × 512 × 512
24removeddec_block_1
Transformer Block
1.8M1 × 448 × 384
25changed
type, embedDim, numHeads, ffDim
dec_block_2
Transformer Block
1.8M1 × 448 × 384cross_attn
Multi Head Attention
1.1M1 × 128 × 512
26addeddec_ca_residual
Add
1 × 128 × 512
27addeddec_ffn_norm
Rms Norm
5121 × 128 × 512
28addeddec_ffn
Feed Forward
2.1M1 × 128 × 512
29addeddec_ffn_residual
Add
1 × 128 × 512
30changed
normalizedShape
dec_norm
Layer Norm
7681 × 448 × 384dec_out_norm
Layer Norm
1.0K1 × 128 × 512
31changed
outFeatures
lm_head
Linear
1 × 448 × 51865lm_head
Linear
1 × 128 × 32128
32sametoken_logits
Output
1 × 448 × 51865logits
Output
1 × 128 × 32128

What this is not

Take it further

Open either graph in the editor, change it, and check it again: Whisper Small · T5 Small

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

BERT Base vs T5 Small
Encoder-only against encoder-decoder.