Comparisons / GPT-2 vs BERT Base
GPT-2 vs BERT Base
Decoder-only against encoder-only, same era, same size class.
BERT Base has 53M fewer parameters than GPT-2: 2 layers added, 3 removed, 5 changed.
GPT-2
- Layers
- 10
- Parameters
- 84M
- Input
- 1 × 1024
- Output
- 1 × 1024 × 50257
- Forward-passes
- yes
- Est. train cost
- $1.82
BERT Base
- Layers
- 9
- Parameters
- 31M
- Input
- 1 × 512
- Output
- 1 × 512 × 768
- Forward-passes
- yes
- Est. train cost
- $0.197
The deltas
Every number is BERT Base relative to GPT-2.
Which GPUs each one fits
Each side is measured at its own declared input (1 × 1024 against 1 × 512). Both columns are right about their own model; the difference between them is not a fact about the designs.
| GPU | GPT-2 | BERT Base |
|---|---|---|
| T4 16GB | fits | fits |
| A100 40GB | fits | fits |
| H100 80GB | fits | fits |
Layer by layer
Aligned in topological order. 4 of 14 rows are the same layer with the same parameters.
Hide all 14 rows
| GPT-2 | BERT Base | ||||||
|---|---|---|---|---|---|---|---|
| Layer | Params | Output | Layer | Params | Output | ||
| 1 | changed shape | tokens Input | 1 × 1024 | input_ids Input | 1 × 512 | ||
| 2 | changed numEmbeddings | token_embed Embedding | 39M | 1 × 1024 × 768 | word_embed Embedding | 23M | 1 × 512 × 768 |
| 3 | changed maxLen | pos_embed Positional Encoding | 1 × 1024 × 768 | pos_embed Positional Encoding | 1 × 512 × 768 | ||
| 4 | same | ln_1 Layer Norm | 1.5K | 1 × 1024 × 768 | embed_norm Layer Norm | 1.5K | 1 × 512 × 768 |
| 5 | removed | attn Causal Attention | 2.4M | 1 × 1024 × 768 | — | ||
| 6 | removed | residual_1 Add | 1 × 1024 × 768 | — | |||
| 7 | added | — | embed_drop Dropout | 1 × 512 × 768 | |||
| 8 | added | — | self_attn Multi Head Attention | 2.4M | 1 × 512 × 768 | ||
| 9 | same | ln_2 Layer Norm | 1.5K | 1 × 1024 × 768 | norm Layer Norm | 1.5K | 1 × 512 × 768 |
| 10 | changed hiddenDim, embedDim | mlp Feed Forward | 4.7M | 1 × 1024 × 768 | dense Feed Forward | 4.7M | 1 × 512 × 768 |
| 11 | removed | residual_2 Add | 1 × 1024 × 768 | — | |||
| 12 | same | ln_f Layer Norm | 1.5K | 1 × 1024 × 768 | norm Layer Norm | 1.5K | 1 × 512 × 768 |
| 13 | changed outFeatures, inFeatures | lm_head Linear | 1 × 1024 × 50257 | dense Linear | 591K | 1 × 512 × 768 | |
| 14 | same | logits Output | 1 × 1024 × 50257 | cls_embedding Output | 1 × 512 × 768 | ||
What this is not
- The two are priced at different declared inputs (1 × 1024 against 1 × 512), so memory, cost and GPU fit are each right about their own model and are not a comparison between them. The layer and parameter deltas are unaffected.
- Parameter counts are derived from the graph, not read from a checkpoint. They are exact for a graph that is fully specified and approximate for one that is not.
- Cost and GPU fit are estimates from the graph under one set of assumptions, not measurements of a run.
Take it further
Open either graph in the editor, change it, and check it again: GPT-2 · BERT Base
Compare any two models of your own, including anything on Hugging Face: the comparison tool.
Machine-readable: this page as markdown ·
the pair index ·
POST https://www.neurarch.com/api/v1/plan for a graph of your own.