N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / Transformer Block vs Mamba SSM Block

Transformer Block vs Mamba SSM Block

Attention against a state-space layer for the same job.

Mamba SSM Block has 161M more parameters than Transformer Block: 12 layers added, 2 removed, 4 changed.

Baseline

Transformer Block

Layers
6
Parameters
7.1M
Input
512 × 768
Output
512 × 768
Forward-passes
yes
Est. train cost
$0.185
Compared

Mamba SSM Block

Layers
16
Parameters
168M
Input
1 × 1024
Output
1 × 1024 × 50280
Forward-passes
yes
Est. train cost
$2.58

The deltas

Every number is Mamba SSM Block relative to Transformer Block.

Parameters
+161M (24× the size)
Layers
+10
Added
12
Removed
2
Changed
4
Unchanged
2

Which GPUs each one fits

Each side is measured at its own declared input (512 × 768 against 1 × 1024). Both columns are right about their own model; the difference between them is not a fact about the designs.

GPUTransformer BlockMamba SSM Block
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 2 of 20 rows are the same layer with the same parameters.

Hide all 20 rows
Transformer BlockMamba SSM Block
LayerParamsOutputLayerParamsOutput
1changed
shape
Input
Input
512 × 768tokens
Input
1 × 1024
2removedMultiHeadAttention
Multi Head Attention
2.4M512 × 768
3removedAdd_1
Add
512 × 768
4addedembed
Embedding
51M1 × 1024 × 1024
5changed
type, normalizedShape
LayerNorm_1
Layer Norm
512 × 768norm_ssm
Rms Norm
1.0K1 × 1024 × 1024
6addedin_proj
Linear
1 × 1024 × 4096
7addedto_channels
Permute
1 × 4096 × 1024
8addedz_gate
Swish
1 × 1024 × 4096
9addedcausal_conv
Conv1d
16K1 × 4096 × 1024
10addedto_tokens
Permute
1 × 1024 × 4096
11addedsilu_x
Swish
1 × 1024 × 4096
12addedssm_scan
Mamba
51M1 × 1024 × 4096
13addedgate_out
Multiply
1 × 1024 × 4096
14changed
type, hiddenDim, ffDim, outFeatures
FeedForward
Feed Forward
4.7M512 × 768out_proj
Linear
1 × 1024 × 1024
15sameAdd_2
Add
512 × 768residual_1
Add
1 × 1024 × 1024
16changed
type, normalizedShape
LayerNorm_2
Layer Norm
512 × 768norm_ffn
Rms Norm
1.0K1 × 1024 × 1024
17addedffn
Swiglu
6.3M1 × 1024 × 1024
18addedresidual_2
Add
1 × 1024 × 1024
19addedlm_head
Linear
1 × 1024 × 50280
20sameOutput
Output
512 × 768output
Output
1 × 1024 × 50280

What this is not

Take it further

Open either graph in the editor, change it, and check it again: Transformer Block · Mamba SSM Block

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

Transformer Block vs LLaMA-3 Block
What changed in the transformer block between 2017 and now.
Mamba SSM Block vs Jamba
Pure state space against a hybrid that keeps some attention.
Simple RNN vs Mamba SSM Block
The recurrent layer everyone started with against its modern replacement.