N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / GPT-2 vs BERT Base

GPT-2 vs BERT Base

Decoder-only against encoder-only, same era, same size class.

BERT Base has 53M fewer parameters than GPT-2: 2 layers added, 3 removed, 5 changed.

Baseline

GPT-2

Layers
10
Parameters
84M
Input
1 × 1024
Output
1 × 1024 × 50257
Forward-passes
yes
Est. train cost
$1.82
Compared

BERT Base

Layers
9
Parameters
31M
Input
1 × 512
Output
1 × 512 × 768
Forward-passes
yes
Est. train cost
$0.197

The deltas

Every number is BERT Base relative to GPT-2.

Parameters
-53M (-63.1%)
Layers
-1
Added
2
Removed
3
Changed
5
Unchanged
4

Which GPUs each one fits

Each side is measured at its own declared input (1 × 1024 against 1 × 512). Both columns are right about their own model; the difference between them is not a fact about the designs.

GPUGPT-2BERT Base
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 4 of 14 rows are the same layer with the same parameters.

Hide all 14 rows
GPT-2BERT Base
LayerParamsOutputLayerParamsOutput
1changed
shape
tokens
Input
1 × 1024input_ids
Input
1 × 512
2changed
numEmbeddings
token_embed
Embedding
39M1 × 1024 × 768word_embed
Embedding
23M1 × 512 × 768
3changed
maxLen
pos_embed
Positional Encoding
1 × 1024 × 768pos_embed
Positional Encoding
1 × 512 × 768
4sameln_1
Layer Norm
1.5K1 × 1024 × 768embed_norm
Layer Norm
1.5K1 × 512 × 768
5removedattn
Causal Attention
2.4M1 × 1024 × 768
6removedresidual_1
Add
1 × 1024 × 768
7addedembed_drop
Dropout
1 × 512 × 768
8addedself_attn
Multi Head Attention
2.4M1 × 512 × 768
9sameln_2
Layer Norm
1.5K1 × 1024 × 768norm
Layer Norm
1.5K1 × 512 × 768
10changed
hiddenDim, embedDim
mlp
Feed Forward
4.7M1 × 1024 × 768dense
Feed Forward
4.7M1 × 512 × 768
11removedresidual_2
Add
1 × 1024 × 768
12sameln_f
Layer Norm
1.5K1 × 1024 × 768norm
Layer Norm
1.5K1 × 512 × 768
13changed
outFeatures, inFeatures
lm_head
Linear
1 × 1024 × 50257dense
Linear
591K1 × 512 × 768
14samelogits
Output
1 × 1024 × 50257cls_embedding
Output
1 × 512 × 768

What this is not

Take it further

Open either graph in the editor, change it, and check it again: GPT-2 · BERT Base

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

GPT-2 vs Qwen3-8B
What seven years of scaling actually changed inside the decoder.
BERT Base vs T5 Small
Encoder-only against encoder-decoder.
BERT Base vs ViT-B/16
The same transformer applied to text and to images.