N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / Transformer Block vs LLaMA-3 Block

Transformer Block vs LLaMA-3 Block

What changed in the transformer block between 2017 and now.

LLaMA-3 Block has 695M more parameters than Transformer Block: 5 layers added, 3 removed, 2 changed.

Baseline

Transformer Block

Layers
6
Parameters
7.1M
Input
512 × 768
Output
512 × 768
Forward-passes
yes
Est. train cost
$0.185
Compared

LLaMA-3 Block

Layers
8
Parameters
703M
Input
1 × 2048
Output
1 × 2048 × 4096
Forward-passes
yes
Est. train cost
$14.27

The deltas

Every number is LLaMA-3 Block relative to Transformer Block.

Parameters
+695M (99× the size)
Layers
+2
Added
5
Removed
3
Changed
2
Unchanged
3

Which GPUs each one fits

Each side is measured at its own declared input (512 × 768 against 1 × 2048). Both columns are right about their own model; the difference between them is not a fact about the designs.

GPUTransformer BlockLLaMA-3 Block
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 3 of 13 rows are the same layer with the same parameters.

Hide all 13 rows
Transformer BlockLLaMA-3 Block
LayerParamsOutputLayerParamsOutput
1changed
shape
Input
Input
512 × 768tokens
Input
1 × 2048
2removedMultiHeadAttention
Multi Head Attention
2.4M512 × 768
3addedrope
Rope
4addedembed
Embedding
525M1 × 2048 × 4096
5addedattn_norm
Rms Norm
4.1K1 × 2048 × 4096
6addedgqa
Grouped Query Attention
42M1 × 2048 × 4096
7sameAdd_1
Add
512 × 768residual_1
Add
1 × 2048 × 4096
8changed
type, normalizedShape
LayerNorm_1
Layer Norm
512 × 768ffn_norm
Rms Norm
4.1K1 × 2048 × 4096
9removedFeedForward
Feed Forward
4.7M512 × 768
10addedswiglu_ffn
Swiglu
135M1 × 2048 × 4096
11sameAdd_2
Add
512 × 768residual_2
Add
1 × 2048 × 4096
12removedLayerNorm_2
Layer Norm
512 × 768
13sameOutput
Output
512 × 768hidden_state
Output
1 × 2048 × 4096

What this is not

Take it further

Open either graph in the editor, change it, and check it again: Transformer Block · LLaMA-3 Block

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

LLaMA-3 Block vs Mixtral MoE Block
A dense feed-forward against a mixture of experts, at the block level.
LLaMA-3 Block vs DeepSeek-V3
What a frontier open model adds to the block everyone started from.
Transformer Block vs Mamba SSM Block
Attention against a state-space layer for the same job.
LLaMA-3 Block vs Phi-3 Mini Block
A block from a large model against a whole small one.