Comparisons / Simple RNN vs Mamba SSM Block
Simple RNN vs Mamba SSM Block
The recurrent layer everyone started with against its modern replacement.
Mamba SSM Block has 167M more parameters than Simple RNN: 15 layers added, 1 removed, 2 changed.
Simple RNN
- Layers
- 2
- Parameters
- 1.1M
- Input
- 128 × 300
- Output
- 128 × 10
- Forward-passes
- yes
- Est. train cost
- $0.047
Mamba SSM Block
- Layers
- 16
- Parameters
- 168M
- Input
- 1 × 1024
- Output
- 1 × 1024 × 50280
- Forward-passes
- yes
- Est. train cost
- $2.58
The deltas
Every number is Mamba SSM Block relative to Simple RNN.
Which GPUs each one fits
Each side is measured at its own declared input (128 × 300 against 1 × 1024). Both columns are right about their own model; the difference between them is not a fact about the designs.
| GPU | Simple RNN | Mamba SSM Block |
|---|---|---|
| T4 16GB | fits | fits |
| A100 40GB | fits | fits |
| H100 80GB | fits | fits |
Layer by layer
Aligned in topological order. 1 of 19 rows are the same layer with the same parameters.
Hide all 19 rows
| Simple RNN | Mamba SSM Block | ||||||
|---|---|---|---|---|---|---|---|
| Layer | Params | Output | Layer | Params | Output | ||
| 1 | changed shape | Input Input | 128 × 300 | tokens Input | 1 × 1024 | ||
| 2 | removed | LSTM Lstm | 791K | 128 × 256 | — | ||
| 3 | added | — | embed Embedding | 51M | 1 × 1024 × 1024 | ||
| 4 | added | — | norm_ssm Rms Norm | 1.0K | 1 × 1024 × 1024 | ||
| 5 | added | — | in_proj Linear | 1 × 1024 × 4096 | |||
| 6 | added | — | to_channels Permute | 1 × 4096 × 1024 | |||
| 7 | added | — | z_gate Swish | 1 × 1024 × 4096 | |||
| 8 | added | — | causal_conv Conv1d | 16K | 1 × 4096 × 1024 | ||
| 9 | added | — | to_tokens Permute | 1 × 1024 × 4096 | |||
| 10 | added | — | silu_x Swish | 1 × 1024 × 4096 | |||
| 11 | added | — | ssm_scan Mamba | 51M | 1 × 1024 × 4096 | ||
| 12 | added | — | gate_out Multiply | 1 × 1024 × 4096 | |||
| 13 | added | — | out_proj Linear | 1 × 1024 × 1024 | |||
| 14 | added | — | residual_1 Add | 1 × 1024 × 1024 | |||
| 15 | added | — | norm_ffn Rms Norm | 1.0K | 1 × 1024 × 1024 | ||
| 16 | added | — | ffn Swiglu | 6.3M | 1 × 1024 × 1024 | ||
| 17 | added | — | residual_2 Add | 1 × 1024 × 1024 | |||
| 18 | changed outFeatures | Linear Linear | 128 × 10 | lm_head Linear | 1 × 1024 × 50280 | ||
| 19 | same | Output Output | 128 × 10 | output Output | 1 × 1024 × 50280 | ||
What this is not
- The two are priced at different declared inputs (128 × 300 against 1 × 1024), so memory, cost and GPU fit are each right about their own model and are not a comparison between them. The layer and parameter deltas are unaffected.
- Parameter counts are derived from the graph, not read from a checkpoint. They are exact for a graph that is fully specified and approximate for one that is not.
- Cost and GPU fit are estimates from the graph under one set of assumptions, not measurements of a run.
Take it further
Open either graph in the editor, change it, and check it again: Simple RNN · Mamba SSM Block
Compare any two models of your own, including anything on Hugging Face: the comparison tool.
Machine-readable: this page as markdown ·
the pair index ·
POST https://www.neurarch.com/api/v1/plan for a graph of your own.