N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / U-Net vs Diffusion UNet

U-Net vs Diffusion UNet

A segmentation U-Net against the one a diffusion model uses.

Diffusion UNet has 6.0M more parameters than U-Net: 7 layers added, 12 removed, 11 changed.

Baseline

U-Net

Layers
22
Parameters
721K
Input
3 × 256 × 256
Output
1 × 256 × 768
Forward-passes
yes
Est. train cost
$0.593
Compared

Diffusion UNet

Layers
17
Parameters
6.7M
Input
4 × 64 × 64
Output
4 × 64 × 64
Forward-passes
yes
Est. train cost
$0.541

The deltas

Every number is Diffusion UNet relative to U-Net.

Parameters
+6.0M (9.3× the size)
Layers
-5
Added
7
Removed
12
Changed
11
Unchanged
1

Which GPUs each one fits

Each side is measured at its own declared input (3 × 256 × 256 against 4 × 64 × 64). Both columns are right about their own model; the difference between them is not a fact about the designs.

GPUU-NetDiffusion UNet
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 1 of 31 rows are the same layer with the same parameters.

Hide all 31 rows
U-NetDiffusion UNet
LayerParamsOutputLayerParamsOutput
1changed
shape
image
Input
3 × 256 × 256noisy_latent
Input
4 × 64 × 64
2removedenc1_conv
Conv2d
64064 × 256 × 256
3removedenc1_bn
Batch Norm
12864 × 256 × 256
4removedenc1_relu
Relu
64 × 256 × 256
5removedenc1_pool
Maxpool2d
64 × 128 × 128
6changed
outChannels
enc2_conv
Conv2d
1.3K128 × 128 × 128conv_in
Conv2d
3.2K320 × 64 × 64
7changed
type, numFeatures, numGroups, numChannels
enc2_bn
Batch Norm
256128 × 128 × 128down1_norm
Group Norm
640320 × 64 × 64
8removedenc2_relu
Relu
128 × 128 × 128
9removedenc2_pool
Maxpool2d
128 × 64 × 64
10changed
outChannels
bottleneck_conv
Conv2d
2.6K256 × 64 × 64down1_conv
Conv2d
3.2K320 × 64 × 64
11removedbottleneck_bn
Batch Norm
512256 × 64 × 64
12changed
type
bottleneck_relu
Relu
256 × 64 × 64down1_silu
Swish
320 × 64 × 64
13removedup2
Transpose Conv2d
640128 × 128 × 128
14removeddec2_skip
Concatenate
128 × 128 × 256
15addedto_tokens
Reshape
4096 × 320
16addeddown1_text_attn
Cross Attention
411K4096 × 320
17addedto_feature_map
Reshape
320 × 64 × 64
18changed
outChannels, stride
dec2_conv
Conv2d
1.3K128 × 128 × 256downsample_1
Conv2d
6.4K640 × 32 × 32
19changed
type, numFeatures, numGroups, numChannels
dec2_bn
Batch Norm
256128 × 128 × 256mid_norm
Group Norm
1.3K640 × 32 × 32
20removeddec2_relu
Relu
128 × 128 × 256
21removedup1
Transpose Conv2d
32064 × 256 × 512
22removeddec1_skip
Concatenate
64 × 256 × 768
23addedto_tokens
Reshape
1024 × 640
24addedmid_text_attn
Cross Attention
1.6M1024 × 640
25addedto_feature_map
Reshape
640 × 32 × 32
26addedupsample_1
Upsample
640 × 64 × 64
27changed
outChannels
dec1_conv
Conv2d
64064 × 256 × 768up1_conv
Conv2d
3.2K320 × 64 × 64
28changed
type, numFeatures, numGroups, numChannels
dec1_bn
Batch Norm
12864 × 256 × 768conv_out_norm
Group Norm
640320 × 64 × 64
29changed
type
dec1_relu
Relu
64 × 256 × 768up1_silu
Swish
320 × 64 × 64
30changed
outChannels, kernelSize, padding
output_conv
Conv2d
21 × 256 × 768conv_out
Conv2d
404 × 64 × 64
31samesegmentation_mask
Output
1 × 256 × 768predicted_noise
Output
4 × 64 × 64

What this is not

Take it further

Open either graph in the editor, change it, and check it again: U-Net · Diffusion UNet

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

Diffusion UNet vs DiT-XL/2
Convolutional against transformer backbones for diffusion.