N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / Mamba SSM Block vs Jamba

Mamba SSM Block vs Jamba

Pure state space against a hybrid that keeps some attention.

Jamba has 13B more parameters than Mamba SSM Block: 42 layers added, 7 removed, 8 changed.

Baseline

Mamba SSM Block

Layers
16
Parameters
168M
Input
1 × 1024
Output
1 × 1024 × 50280
Forward-passes
yes
Est. train cost
$2.58
Compared

Jamba

Layers
51
Parameters
13B
Input
1 × 4096
Output
1 × 4096 × 65536
Forward-passes
yes
Est. train cost
$338.33

The deltas

Every number is Jamba relative to Mamba SSM Block.

Parameters
+13B (77× the size)
Layers
+35
Added
42
Removed
7
Changed
8
Unchanged
3

Which GPUs each one fits

Each side is measured at its own declared input (1 × 1024 against 1 × 4096). Both columns are right about their own model; the difference between them is not a fact about the designs.

GPUMamba SSM BlockJamba
T4 16GBfitsno
A100 40GBfitsno
H100 80GBfitsno

Layer by layer

Aligned in topological order. 3 of 60 rows are the same layer with the same parameters.

Hide all 60 rows
Mamba SSM BlockJamba
LayerParamsOutputLayerParamsOutput
1changed
shape
tokens
Input
1 × 1024tokens
Input
1 × 4096
2changed
numEmbeddings, embeddingDim
embed
Embedding
51M1 × 1024 × 1024token_embed
Embedding
268M1 × 4096 × 4096
3addedmixer_norm_1
Rms Norm
4.1K1 × 4096 × 4096
4addedmamba_1
Mamba
101M1 × 4096 × 4096
5addedmixer_residual_1
Add
1 × 4096 × 4096
6changed
normalizedShape
norm_ssm
Rms Norm
1.0K1 × 1024 × 1024ffn_norm_1
Rms Norm
4.1K1 × 4096 × 4096
7changed
type, outFeatures, embedDim, ffDim
in_proj
Linear
1 × 1024 × 4096mlp_1
Feed Forward
117M1 × 4096 × 4096
8removedto_channels
Permute
1 × 4096 × 1024
9removedz_gate
Swish
1 × 1024 × 4096
10removedcausal_conv
Conv1d
16K1 × 4096 × 1024
11removedto_tokens
Permute
1 × 1024 × 4096
12removedsilu_x
Swish
1 × 1024 × 4096
13addedffn_residual_1
Add
1 × 4096 × 4096
14addedmixer_norm_2
Rms Norm
4.1K1 × 4096 × 4096
15changed
dConv, expand
ssm_scan
Mamba
51M1 × 1024 × 4096mamba_2
Mamba
101M1 × 4096 × 4096
16removedgate_out
Multiply
1 × 1024 × 4096
17addedmixer_residual_2
Add
1 × 4096 × 4096
18addedffn_norm_2
Rms Norm
4.1K1 × 4096 × 4096
19addedmoe_2
Moe Layer
2.8B1 × 4096 × 4096
20addedffn_residual_2
Add
1 × 4096 × 4096
21addedmixer_norm_3
Rms Norm
4.1K1 × 4096 × 4096
22addedmamba_3
Mamba
101M1 × 4096 × 4096
23addedmixer_residual_3
Add
1 × 4096 × 4096
24addedffn_norm_3
Rms Norm
4.1K1 × 4096 × 4096
25changed
type, outFeatures, embedDim, ffDim
out_proj
Linear
1 × 1024 × 1024mlp_3
Feed Forward
117M1 × 4096 × 4096
26sameresidual_1
Add
1 × 1024 × 1024ffn_residual_3
Add
1 × 4096 × 4096
27changed
normalizedShape
norm_ffn
Rms Norm
1.0K1 × 1024 × 1024mixer_norm_4
Rms Norm
4.1K1 × 4096 × 4096
28removedffn
Swiglu
6.3M1 × 1024 × 1024
29addedmamba_4
Mamba
101M1 × 4096 × 4096
30addedmixer_residual_4
Add
1 × 4096 × 4096
31addedffn_norm_4
Rms Norm
4.1K1 × 4096 × 4096
32addedmoe_4
Moe Layer
2.8B1 × 4096 × 4096
33addedffn_residual_4
Add
1 × 4096 × 4096
34addedmixer_norm_5
Rms Norm
4.1K1 × 4096 × 4096
35addedattn_5
Grouped Query Attention
42M1 × 4096 × 4096
36addedmixer_residual_5
Add
1 × 4096 × 4096
37addedffn_norm_5
Rms Norm
4.1K1 × 4096 × 4096
38addedmlp_5
Feed Forward
117M1 × 4096 × 4096
39addedffn_residual_5
Add
1 × 4096 × 4096
40addedmixer_norm_6
Rms Norm
4.1K1 × 4096 × 4096
41addedmamba_6
Mamba
101M1 × 4096 × 4096
42addedmixer_residual_6
Add
1 × 4096 × 4096
43addedffn_norm_6
Rms Norm
4.1K1 × 4096 × 4096
44addedmoe_6
Moe Layer
2.8B1 × 4096 × 4096
45addedffn_residual_6
Add
1 × 4096 × 4096
46addedmixer_norm_7
Rms Norm
4.1K1 × 4096 × 4096
47addedmamba_7
Mamba
101M1 × 4096 × 4096
48addedmixer_residual_7
Add
1 × 4096 × 4096
49addedffn_norm_7
Rms Norm
4.1K1 × 4096 × 4096
50addedmlp_7
Feed Forward
117M1 × 4096 × 4096
51addedffn_residual_7
Add
1 × 4096 × 4096
52addedmixer_norm_8
Rms Norm
4.1K1 × 4096 × 4096
53addedmamba_8
Mamba
101M1 × 4096 × 4096
54addedmixer_residual_8
Add
1 × 4096 × 4096
55addedffn_norm_8
Rms Norm
4.1K1 × 4096 × 4096
56addedmoe_8
Moe Layer
2.8B1 × 4096 × 4096
57sameresidual_2
Add
1 × 1024 × 1024ffn_residual_8
Add
1 × 4096 × 4096
58addedfinal_norm
Rms Norm
4.1K1 × 4096 × 4096
59changed
outFeatures, inFeatures
lm_head
Linear
1 × 1024 × 50280lm_head
Linear
269M1 × 4096 × 65536
60sameoutput
Output
1 × 1024 × 50280logits
Output
1 × 4096 × 65536

What this is not

Take it further

Open either graph in the editor, change it, and check it again: Mamba SSM Block · Jamba

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

Transformer Block vs Mamba SSM Block
Attention against a state-space layer for the same job.
Simple RNN vs Mamba SSM Block
The recurrent layer everyone started with against its modern replacement.