N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / BERT Base vs ViT-B/16

BERT Base vs ViT-B/16

The same transformer applied to text and to images.

ViT-B/16 has 23M fewer parameters than BERT Base: 4 layers added, 2 removed, 5 changed.

Baseline

BERT Base

Layers
9
Parameters
31M
Input
1 × 512
Output
1 × 512 × 768
Forward-passes
yes
Est. train cost
$0.197
Compared

ViT-B/16

Layers
11
Parameters
8.4M
Input
3 × 224 × 224
Output
196 × 1000
Forward-passes
yes
Est. train cost
$0.105

The deltas

Every number is ViT-B/16 relative to BERT Base.

Parameters
-23M (-72.9%)
Layers
+2
Added
4
Removed
2
Changed
5
Unchanged
4

Which GPUs each one fits

Each side is measured at its own declared input (1 × 512 against 3 × 224 × 224). Both columns are right about their own model; the difference between them is not a fact about the designs.

GPUBERT BaseViT-B/16
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 4 of 15 rows are the same layer with the same parameters.

Hide all 15 rows
BERT BaseViT-B/16
LayerParamsOutputLayerParamsOutput
1changed
shape
input_ids
Input
1 × 512image
Input
3 × 224 × 224
2removedword_embed
Embedding
23M1 × 512 × 768
3addedpatch_embed
Patch Embed
591K196 × 768
4changed
maxLen
pos_embed
Positional Encoding
1 × 512 × 768pos_embed
Positional Encoding
196 × 768
5removedembed_norm
Layer Norm
1.5K1 × 512 × 768
6changed
p
embed_drop
Dropout
1 × 512 × 768dropout
Dropout
196 × 768
7addednorm_1
Layer Norm
1.5K196 × 768
8sameself_attn
Multi Head Attention
2.4M1 × 512 × 768attn
Multi Head Attention
2.4M196 × 768
9addedresidual_1
Add
196 × 768
10samenorm
Layer Norm
1.5K1 × 512 × 768norm_2
Layer Norm
1.5K196 × 768
11changed
embedDim, hiddenDim
dense
Feed Forward
4.7M1 × 512 × 768mlp
Feed Forward
4.7M196 × 768
12addedresidual_2
Add
196 × 768
13samenorm
Layer Norm
1.5K1 × 512 × 768norm_final
Layer Norm
1.5K196 × 768
14changed
inFeatures, outFeatures
dense
Linear
591K1 × 512 × 768head
Linear
196 × 1000
15samecls_embedding
Output
1 × 512 × 768class_logits
Output
196 × 1000

What this is not

Take it further

Open either graph in the editor, change it, and check it again: BERT Base · ViT-B/16

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

GPT-2 vs BERT Base
Decoder-only against encoder-only, same era, same size class.
BERT Base vs T5 Small
Encoder-only against encoder-decoder.
ResNet Block vs ViT-B/16
Convolution against attention for images.
ViT-B/16 vs Swin-Tiny
A flat vision transformer against a hierarchical one.
ViT-B/16 vs DiT-XL/2
A vision transformer against a diffusion transformer.