N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / BERT Base vs T5 Small

BERT Base vs T5 Small

Encoder-only against encoder-decoder.

T5 Small has 26M more parameters than BERT Base: 14 layers added, 2 removed, 8 changed.

Baseline

BERT Base

Layers
9
Parameters
31M
Input
1 × 512
Output
1 × 512 × 768
Forward-passes
yes
Est. train cost
$0.197
Compared

T5 Small

Layers
20
Parameters
57M
Input
1 × 512
Output
1 × 128 × 32128
Forward-passes
yes
Est. train cost
$0.207

The deltas

Every number is T5 Small relative to BERT Base.

Parameters
+26M (+82.3%)
Layers
+11
Added
14
Removed
2
Changed
8
Unchanged
1

Which GPUs each one fits

Memory for the graph at the input shape both declare. A highlighted row is a card one of them fits and the other does not, which is the difference that decides a purchase.

GPUBERT BaseT5 Small
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 1 of 25 rows are the same layer with the same parameters.

Hide all 25 rows
BERT BaseT5 Small
LayerParamsOutputLayerParamsOutput
1addedencoder_ids
Input
1 × 512
2changed
shape
input_ids
Input
1 × 512decoder_ids
Input
1 × 128
3changed
numEmbeddings, embeddingDim
word_embed
Embedding
23M1 × 512 × 768shared_embed
Embedding
16M1 × 512 × 512
4removedpos_embed
Positional Encoding
1 × 512 × 768
5addeddec_embed
Embedding
16M1 × 128 × 512
6addedenc_norm
Rms Norm
5121 × 512 × 512
7addeddec_sa_norm
Rms Norm
5121 × 128 × 512
8addedenc_self_attn
Multi Head Attention
1.1M1 × 512 × 512
9addeddec_self_attn
Causal Attention
1.0M1 × 128 × 512
10addedenc_residual
Add
1 × 512 × 512
11addeddec_sa_residual
Add
1 × 128 × 512
12addedenc_ffn_norm
Rms Norm
5121 × 512 × 512
13addeddec_ca_norm
Rms Norm
5121 × 128 × 512
14addedenc_ffn
Feed Forward
2.1M1 × 512 × 512
15addedenc_ffn_residual
Add
1 × 512 × 512
16changed
normalizedShape
embed_norm
Layer Norm
1.5K1 × 512 × 768enc_out_norm
Layer Norm
1.0K1 × 512 × 512
17removedembed_drop
Dropout
1 × 512 × 768
18changed
embedDim, numHeads
self_attn
Multi Head Attention
2.4M1 × 512 × 768cross_attn
Multi Head Attention
1.1M1 × 128 × 512
19addeddec_ca_residual
Add
1 × 128 × 512
20changed
type, normalizedShape
norm
Layer Norm
1.5K1 × 512 × 768dec_ffn_norm
Rms Norm
5121 × 128 × 512
21changed
embedDim, ffDim
dense
Feed Forward
4.7M1 × 512 × 768dec_ffn
Feed Forward
2.1M1 × 128 × 512
22addeddec_ffn_residual
Add
1 × 128 × 512
23changed
normalizedShape
norm
Layer Norm
1.5K1 × 512 × 768dec_out_norm
Layer Norm
1.0K1 × 128 × 512
24changed
inFeatures, outFeatures
dense
Linear
591K1 × 512 × 768lm_head
Linear
1 × 128 × 32128
25samecls_embedding
Output
1 × 512 × 768logits
Output
1 × 128 × 32128

What this is not

Take it further

Open either graph in the editor, change it, and check it again: BERT Base · T5 Small

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

GPT-2 vs BERT Base
Decoder-only against encoder-only, same era, same size class.
BERT Base vs ViT-B/16
The same transformer applied to text and to images.
Whisper Small vs T5 Small
Speech-to-text against text-to-text, both encoder-decoder.