N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / ViT-B/16 vs Swin-Tiny

ViT-B/16 vs Swin-Tiny

A flat vision transformer against a hierarchical one.

Swin-Tiny has 20M more parameters than ViT-B/16: 74 layers added, 2 removed, 4 changed.

Baseline

ViT-B/16

Layers
11
Parameters
8.4M
Input
3 × 224 × 224
Output
196 × 1000
Forward-passes
yes
Est. train cost
$0.105
Compared

Swin-Tiny

Layers
83
Parameters
28M
Input
3 × 224 × 224
Output
1000
Forward-passes
yes
Est. train cost
$0.288

The deltas

Every number is Swin-Tiny relative to ViT-B/16.

Parameters
+20M (+234%)
Layers
+72
Added
74
Removed
2
Changed
4
Unchanged
7

Which GPUs each one fits

Memory for the graph at the input shape both declare. A highlighted row is a card one of them fits and the other does not, which is the difference that decides a purchase.

GPUViT-B/16Swin-Tiny
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 7 of 87 rows are the same layer with the same parameters.

Show all 87 rows
ViT-B/16Swin-Tiny
LayerParamsOutputLayerParamsOutput
1sameimage
Input
3 × 224 × 224image
Input
3 × 224 × 224
2changed
patchSize, embedDim
patch_embed
Patch Embed
591K196 × 768patch_embed_4x4
Patch Embed
4.7K3136 × 96
3removedpos_embed
Positional Encoding
196 × 768
4removeddropout
Dropout
196 × 768
5addeds1b1_norm1
Layer Norm
1923136 × 96
6addeds1b1_window_attn
Multi Head Attention
37K3136 × 96
7addeds1b1_res1
Add
3136 × 96
8addeds1b1_norm2
Layer Norm
1923136 × 96
9addeds1b1_mlp
Feed Forward
74K3136 × 96
10addeds1b1_res2
Add
3136 × 96
11addeds1b2_norm1
Layer Norm
1923136 × 96
12addeds1b2_shifted_window_attn
Multi Head Attention
37K3136 × 96
13addeds1b2_res1
Add
3136 × 96
14addeds1b2_norm2
Layer Norm
1923136 × 96
15addeds1b2_mlp
Feed Forward
74K3136 × 96
16addeds1b2_res2
Add
3136 × 96
17addedmerge_patches_n15
Reshape
784 × 384
18addedpatch_merging_2
Linear
74K784 × 192
19addeds2b1_norm1
Layer Norm
384784 × 192
20addeds2b1_window_attn
Multi Head Attention
148K784 × 192
21addeds2b1_res1
Add
784 × 192
22addeds2b1_norm2
Layer Norm
384784 × 192
23addeds2b1_mlp
Feed Forward
296K784 × 192
24addeds2b1_res2
Add
784 × 192
25addeds2b2_norm1
Layer Norm
384784 × 192
26addeds2b2_shifted_window_attn
Multi Head Attention
148K784 × 192
27addeds2b2_res1
Add
784 × 192
28addeds2b2_norm2
Layer Norm
384784 × 192
29addeds2b2_mlp
Feed Forward
296K784 × 192
30addeds2b2_res2
Add
784 × 192
31addedmerge_patches_n28
Reshape
196 × 768
32addedpatch_merging_3
Linear
295K196 × 384
33addeds3b1_norm1
Layer Norm
768196 × 384
34addeds3b1_window_attn
Multi Head Attention
591K196 × 384
35addeds3b1_res1
Add
196 × 384
36addeds3b1_norm2
Layer Norm
768196 × 384
37addeds3b1_mlp
Feed Forward
1.2M196 × 384
38addeds3b1_res2
Add
196 × 384
39addeds3b2_norm1
Layer Norm
768196 × 384
40addeds3b2_shifted_window_attn
Multi Head Attention
591K196 × 384
41addeds3b2_res1
Add
196 × 384
42addeds3b2_norm2
Layer Norm
768196 × 384
43addeds3b2_mlp
Feed Forward
1.2M196 × 384
44addeds3b2_res2
Add
196 × 384
45addeds3b3_norm1
Layer Norm
768196 × 384
46addeds3b3_window_attn
Multi Head Attention
591K196 × 384
47addeds3b3_res1
Add
196 × 384
48addeds3b3_norm2
Layer Norm
768196 × 384
49addeds3b3_mlp
Feed Forward
1.2M196 × 384
50addeds3b3_res2
Add
196 × 384
51addeds3b4_norm1
Layer Norm
768196 × 384
52addeds3b4_shifted_window_attn
Multi Head Attention
591K196 × 384
53addeds3b4_res1
Add
196 × 384
54addeds3b4_norm2
Layer Norm
768196 × 384
55addeds3b4_mlp
Feed Forward
1.2M196 × 384
56addeds3b4_res2
Add
196 × 384
57addeds3b5_norm1
Layer Norm
768196 × 384
58addeds3b5_window_attn
Multi Head Attention
591K196 × 384
59addeds3b5_res1
Add
196 × 384
60addeds3b5_norm2
Layer Norm
768196 × 384
61addeds3b5_mlp
Feed Forward
1.2M196 × 384
62addeds3b5_res2
Add
196 × 384
63addeds3b6_norm1
Layer Norm
768196 × 384
64addeds3b6_shifted_window_attn
Multi Head Attention
591K196 × 384
65addeds3b6_res1
Add
196 × 384
66addeds3b6_norm2
Layer Norm
768196 × 384
67addeds3b6_mlp
Feed Forward
1.2M196 × 384
68addeds3b6_res2
Add
196 × 384
69addedmerge_patches_n65
Reshape
49 × 1536
70addedpatch_merging_4
Linear
1.2M49 × 768
71addeds4b1_norm1
Layer Norm
1.5K49 × 768
72addeds4b1_window_attn
Multi Head Attention
2.4M49 × 768
73addeds4b1_res1
Add
49 × 768
74addeds4b1_norm2
Layer Norm
1.5K49 × 768
75addeds4b1_mlp
Feed Forward
4.7M49 × 768
76addeds4b1_res2
Add
49 × 768
77samenorm_1
Layer Norm
1.5K196 × 768s4b2_norm1
Layer Norm
1.5K49 × 768
78changed
numHeads
attn
Multi Head Attention
2.4M196 × 768s4b2_shifted_window_attn
Multi Head Attention
2.4M49 × 768
79sameresidual_1
Add
196 × 768s4b2_res1
Add
49 × 768
80samenorm_2
Layer Norm
1.5K196 × 768s4b2_norm2
Layer Norm
1.5K49 × 768
81changed
hiddenDim, embedDim
mlp
Feed Forward
4.7M196 × 768s4b2_mlp
Feed Forward
4.7M49 × 768
82sameresidual_2
Add
196 × 768s4b2_res2
Add
49 × 768
83samenorm_final
Layer Norm
1.5K196 × 768final_norm
Layer Norm
1.5K49 × 768
84addedto_channels
Permute
768 × 49
85addedavgpool
Global Avg Pool1d
768
86changed
inFeatures
head
Linear
196 × 1000classifier
Linear
769K1000
87sameclass_logits
Output
196 × 1000class_logits
Output
1000

What this is not

Take it further

Open either graph in the editor, change it, and check it again: ViT-B/16 · Swin-Tiny

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

ResNet Block vs ViT-B/16
Convolution against attention for images.
BERT Base vs ViT-B/16
The same transformer applied to text and to images.
ViT-B/16 vs DiT-XL/2
A vision transformer against a diffusion transformer.