Comparisons / Mamba SSM Block vs Jamba
Mamba SSM Block vs Jamba
Pure state space against a hybrid that keeps some attention.
Jamba has 13B more parameters than Mamba SSM Block: 42 layers added, 7 removed, 8 changed.
Mamba SSM Block
- Layers
- 16
- Parameters
- 168M
- Input
- 1 × 1024
- Output
- 1 × 1024 × 50280
- Forward-passes
- yes
- Est. train cost
- $2.58
Jamba
- Layers
- 51
- Parameters
- 13B
- Input
- 1 × 4096
- Output
- 1 × 4096 × 65536
- Forward-passes
- yes
- Est. train cost
- $338.33
The deltas
Every number is Jamba relative to Mamba SSM Block.
Which GPUs each one fits
Each side is measured at its own declared input (1 × 1024 against 1 × 4096). Both columns are right about their own model; the difference between them is not a fact about the designs.
| GPU | Mamba SSM Block | Jamba |
|---|---|---|
| T4 16GB | fits | no |
| A100 40GB | fits | no |
| H100 80GB | fits | no |
Layer by layer
Aligned in topological order. 3 of 60 rows are the same layer with the same parameters.
Hide all 60 rows
| Mamba SSM Block | Jamba | ||||||
|---|---|---|---|---|---|---|---|
| Layer | Params | Output | Layer | Params | Output | ||
| 1 | changed shape | tokens Input | 1 × 1024 | tokens Input | 1 × 4096 | ||
| 2 | changed numEmbeddings, embeddingDim | embed Embedding | 51M | 1 × 1024 × 1024 | token_embed Embedding | 268M | 1 × 4096 × 4096 |
| 3 | added | — | mixer_norm_1 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 4 | added | — | mamba_1 Mamba | 101M | 1 × 4096 × 4096 | ||
| 5 | added | — | mixer_residual_1 Add | 1 × 4096 × 4096 | |||
| 6 | changed normalizedShape | norm_ssm Rms Norm | 1.0K | 1 × 1024 × 1024 | ffn_norm_1 Rms Norm | 4.1K | 1 × 4096 × 4096 |
| 7 | changed type, outFeatures, embedDim, ffDim | in_proj Linear | 1 × 1024 × 4096 | mlp_1 Feed Forward | 117M | 1 × 4096 × 4096 | |
| 8 | removed | to_channels Permute | 1 × 4096 × 1024 | — | |||
| 9 | removed | z_gate Swish | 1 × 1024 × 4096 | — | |||
| 10 | removed | causal_conv Conv1d | 16K | 1 × 4096 × 1024 | — | ||
| 11 | removed | to_tokens Permute | 1 × 1024 × 4096 | — | |||
| 12 | removed | silu_x Swish | 1 × 1024 × 4096 | — | |||
| 13 | added | — | ffn_residual_1 Add | 1 × 4096 × 4096 | |||
| 14 | added | — | mixer_norm_2 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 15 | changed dConv, expand | ssm_scan Mamba | 51M | 1 × 1024 × 4096 | mamba_2 Mamba | 101M | 1 × 4096 × 4096 |
| 16 | removed | gate_out Multiply | 1 × 1024 × 4096 | — | |||
| 17 | added | — | mixer_residual_2 Add | 1 × 4096 × 4096 | |||
| 18 | added | — | ffn_norm_2 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 19 | added | — | moe_2 Moe Layer | 2.8B | 1 × 4096 × 4096 | ||
| 20 | added | — | ffn_residual_2 Add | 1 × 4096 × 4096 | |||
| 21 | added | — | mixer_norm_3 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 22 | added | — | mamba_3 Mamba | 101M | 1 × 4096 × 4096 | ||
| 23 | added | — | mixer_residual_3 Add | 1 × 4096 × 4096 | |||
| 24 | added | — | ffn_norm_3 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 25 | changed type, outFeatures, embedDim, ffDim | out_proj Linear | 1 × 1024 × 1024 | mlp_3 Feed Forward | 117M | 1 × 4096 × 4096 | |
| 26 | same | residual_1 Add | 1 × 1024 × 1024 | ffn_residual_3 Add | 1 × 4096 × 4096 | ||
| 27 | changed normalizedShape | norm_ffn Rms Norm | 1.0K | 1 × 1024 × 1024 | mixer_norm_4 Rms Norm | 4.1K | 1 × 4096 × 4096 |
| 28 | removed | ffn Swiglu | 6.3M | 1 × 1024 × 1024 | — | ||
| 29 | added | — | mamba_4 Mamba | 101M | 1 × 4096 × 4096 | ||
| 30 | added | — | mixer_residual_4 Add | 1 × 4096 × 4096 | |||
| 31 | added | — | ffn_norm_4 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 32 | added | — | moe_4 Moe Layer | 2.8B | 1 × 4096 × 4096 | ||
| 33 | added | — | ffn_residual_4 Add | 1 × 4096 × 4096 | |||
| 34 | added | — | mixer_norm_5 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 35 | added | — | attn_5 Grouped Query Attention | 42M | 1 × 4096 × 4096 | ||
| 36 | added | — | mixer_residual_5 Add | 1 × 4096 × 4096 | |||
| 37 | added | — | ffn_norm_5 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 38 | added | — | mlp_5 Feed Forward | 117M | 1 × 4096 × 4096 | ||
| 39 | added | — | ffn_residual_5 Add | 1 × 4096 × 4096 | |||
| 40 | added | — | mixer_norm_6 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 41 | added | — | mamba_6 Mamba | 101M | 1 × 4096 × 4096 | ||
| 42 | added | — | mixer_residual_6 Add | 1 × 4096 × 4096 | |||
| 43 | added | — | ffn_norm_6 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 44 | added | — | moe_6 Moe Layer | 2.8B | 1 × 4096 × 4096 | ||
| 45 | added | — | ffn_residual_6 Add | 1 × 4096 × 4096 | |||
| 46 | added | — | mixer_norm_7 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 47 | added | — | mamba_7 Mamba | 101M | 1 × 4096 × 4096 | ||
| 48 | added | — | mixer_residual_7 Add | 1 × 4096 × 4096 | |||
| 49 | added | — | ffn_norm_7 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 50 | added | — | mlp_7 Feed Forward | 117M | 1 × 4096 × 4096 | ||
| 51 | added | — | ffn_residual_7 Add | 1 × 4096 × 4096 | |||
| 52 | added | — | mixer_norm_8 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 53 | added | — | mamba_8 Mamba | 101M | 1 × 4096 × 4096 | ||
| 54 | added | — | mixer_residual_8 Add | 1 × 4096 × 4096 | |||
| 55 | added | — | ffn_norm_8 Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 56 | added | — | moe_8 Moe Layer | 2.8B | 1 × 4096 × 4096 | ||
| 57 | same | residual_2 Add | 1 × 1024 × 1024 | ffn_residual_8 Add | 1 × 4096 × 4096 | ||
| 58 | added | — | final_norm Rms Norm | 4.1K | 1 × 4096 × 4096 | ||
| 59 | changed outFeatures, inFeatures | lm_head Linear | 1 × 1024 × 50280 | lm_head Linear | 269M | 1 × 4096 × 65536 | |
| 60 | same | output Output | 1 × 1024 × 50280 | logits Output | 1 × 4096 × 65536 | ||
What this is not
- The two are priced at different declared inputs (1 × 1024 against 1 × 4096), so memory, cost and GPU fit are each right about their own model and are not a comparison between them. The layer and parameter deltas are unaffected.
- Parameter counts are derived from the graph, not read from a checkpoint. They are exact for a graph that is fully specified and approximate for one that is not.
- Cost and GPU fit are estimates from the graph under one set of assumptions, not measurements of a run.
Take it further
Open either graph in the editor, change it, and check it again: Mamba SSM Block · Jamba
Compare any two models of your own, including anything on Hugging Face: the comparison tool.
Machine-readable: this page as markdown ·
the pair index ·
POST https://www.neurarch.com/api/v1/plan for a graph of your own.