N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / Simple RNN vs Mamba SSM Block

Simple RNN vs Mamba SSM Block

The recurrent layer everyone started with against its modern replacement.

Mamba SSM Block has 167M more parameters than Simple RNN: 15 layers added, 1 removed, 2 changed.

Baseline

Simple RNN

Layers
2
Parameters
1.1M
Input
128 × 300
Output
128 × 10
Forward-passes
yes
Est. train cost
$0.047
Compared

Mamba SSM Block

Layers
16
Parameters
168M
Input
1 × 1024
Output
1 × 1024 × 50280
Forward-passes
yes
Est. train cost
$2.58

The deltas

Every number is Mamba SSM Block relative to Simple RNN.

Parameters
+167M (153× the size)
Layers
+14
Added
15
Removed
1
Changed
2
Unchanged
1

Which GPUs each one fits

Each side is measured at its own declared input (128 × 300 against 1 × 1024). Both columns are right about their own model; the difference between them is not a fact about the designs.

GPUSimple RNNMamba SSM Block
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 1 of 19 rows are the same layer with the same parameters.

Hide all 19 rows
Simple RNNMamba SSM Block
LayerParamsOutputLayerParamsOutput
1changed
shape
Input
Input
128 × 300tokens
Input
1 × 1024
2removedLSTM
Lstm
791K128 × 256
3addedembed
Embedding
51M1 × 1024 × 1024
4addednorm_ssm
Rms Norm
1.0K1 × 1024 × 1024
5addedin_proj
Linear
1 × 1024 × 4096
6addedto_channels
Permute
1 × 4096 × 1024
7addedz_gate
Swish
1 × 1024 × 4096
8addedcausal_conv
Conv1d
16K1 × 4096 × 1024
9addedto_tokens
Permute
1 × 1024 × 4096
10addedsilu_x
Swish
1 × 1024 × 4096
11addedssm_scan
Mamba
51M1 × 1024 × 4096
12addedgate_out
Multiply
1 × 1024 × 4096
13addedout_proj
Linear
1 × 1024 × 1024
14addedresidual_1
Add
1 × 1024 × 1024
15addednorm_ffn
Rms Norm
1.0K1 × 1024 × 1024
16addedffn
Swiglu
6.3M1 × 1024 × 1024
17addedresidual_2
Add
1 × 1024 × 1024
18changed
outFeatures
Linear
Linear
128 × 10lm_head
Linear
1 × 1024 × 50280
19sameOutput
Output
128 × 10output
Output
1 × 1024 × 50280

What this is not

Take it further

Open either graph in the editor, change it, and check it again: Simple RNN · Mamba SSM Block

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

Transformer Block vs Mamba SSM Block
Attention against a state-space layer for the same job.
Mamba SSM Block vs Jamba
Pure state space against a hybrid that keeps some attention.