Comparisons / Transformer Block vs Mamba SSM Block
Transformer Block vs Mamba SSM Block
Attention against a state-space layer for the same job.
Mamba SSM Block has 161M more parameters than Transformer Block: 12 layers added, 2 removed, 4 changed.
Transformer Block
- Layers
- 6
- Parameters
- 7.1M
- Input
- 512 × 768
- Output
- 512 × 768
- Forward-passes
- yes
- Est. train cost
- $0.185
Mamba SSM Block
- Layers
- 16
- Parameters
- 168M
- Input
- 1 × 1024
- Output
- 1 × 1024 × 50280
- Forward-passes
- yes
- Est. train cost
- $2.58
The deltas
Every number is Mamba SSM Block relative to Transformer Block.
Which GPUs each one fits
Each side is measured at its own declared input (512 × 768 against 1 × 1024). Both columns are right about their own model; the difference between them is not a fact about the designs.
| GPU | Transformer Block | Mamba SSM Block |
|---|---|---|
| T4 16GB | fits | fits |
| A100 40GB | fits | fits |
| H100 80GB | fits | fits |
Layer by layer
Aligned in topological order. 2 of 20 rows are the same layer with the same parameters.
Hide all 20 rows
| Transformer Block | Mamba SSM Block | ||||||
|---|---|---|---|---|---|---|---|
| Layer | Params | Output | Layer | Params | Output | ||
| 1 | changed shape | Input Input | 512 × 768 | tokens Input | 1 × 1024 | ||
| 2 | removed | MultiHeadAttention Multi Head Attention | 2.4M | 512 × 768 | — | ||
| 3 | removed | Add_1 Add | 512 × 768 | — | |||
| 4 | added | — | embed Embedding | 51M | 1 × 1024 × 1024 | ||
| 5 | changed type, normalizedShape | LayerNorm_1 Layer Norm | 512 × 768 | norm_ssm Rms Norm | 1.0K | 1 × 1024 × 1024 | |
| 6 | added | — | in_proj Linear | 1 × 1024 × 4096 | |||
| 7 | added | — | to_channels Permute | 1 × 4096 × 1024 | |||
| 8 | added | — | z_gate Swish | 1 × 1024 × 4096 | |||
| 9 | added | — | causal_conv Conv1d | 16K | 1 × 4096 × 1024 | ||
| 10 | added | — | to_tokens Permute | 1 × 1024 × 4096 | |||
| 11 | added | — | silu_x Swish | 1 × 1024 × 4096 | |||
| 12 | added | — | ssm_scan Mamba | 51M | 1 × 1024 × 4096 | ||
| 13 | added | — | gate_out Multiply | 1 × 1024 × 4096 | |||
| 14 | changed type, hiddenDim, ffDim, outFeatures | FeedForward Feed Forward | 4.7M | 512 × 768 | out_proj Linear | 1 × 1024 × 1024 | |
| 15 | same | Add_2 Add | 512 × 768 | residual_1 Add | 1 × 1024 × 1024 | ||
| 16 | changed type, normalizedShape | LayerNorm_2 Layer Norm | 512 × 768 | norm_ffn Rms Norm | 1.0K | 1 × 1024 × 1024 | |
| 17 | added | — | ffn Swiglu | 6.3M | 1 × 1024 × 1024 | ||
| 18 | added | — | residual_2 Add | 1 × 1024 × 1024 | |||
| 19 | added | — | lm_head Linear | 1 × 1024 × 50280 | |||
| 20 | same | Output Output | 512 × 768 | output Output | 1 × 1024 × 50280 | ||
What this is not
- The two are priced at different declared inputs (512 × 768 against 1 × 1024), so memory, cost and GPU fit are each right about their own model and are not a comparison between them. The layer and parameter deltas are unaffected.
- Parameter counts are derived from the graph, not read from a checkpoint. They are exact for a graph that is fully specified and approximate for one that is not.
- Cost and GPU fit are estimates from the graph under one set of assumptions, not measurements of a run.
Take it further
Open either graph in the editor, change it, and check it again: Transformer Block · Mamba SSM Block
Compare any two models of your own, including anything on Hugging Face: the comparison tool.
Machine-readable: this page as markdown ·
the pair index ·
POST https://www.neurarch.com/api/v1/plan for a graph of your own.