N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / LLaMA-3 Block vs Phi-3 Mini Block

LLaMA-3 Block vs Phi-3 Mini Block

A block from a large model against a whole small one.

Phi-3 Mini Block has 392M fewer parameters than LLaMA-3 Block: 2 layers added, 6 changed.

Baseline

LLaMA-3 Block

Layers
8
Parameters
703M
Input
1 × 2048
Output
1 × 2048 × 4096
Forward-passes
yes
Est. train cost
$14.27
Compared

Phi-3 Mini Block

Layers
10
Parameters
310M
Input
1 × 2048
Output
1 × 2048 × 32064
Forward-passes
yes
Est. train cost
$16.76

The deltas

Every number is Phi-3 Mini Block relative to LLaMA-3 Block.

Parameters
-392M (-55.8%)
Layers
+2
Added
2
Removed
0
Changed
6
Unchanged
4

Which GPUs each one fits

Memory for the graph at the input shape both declare. A highlighted row is a card one of them fits and the other does not, which is the difference that decides a purchase.

GPULLaMA-3 BlockPhi-3 Mini Block
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 4 of 12 rows are the same layer with the same parameters.

Hide all 12 rows
LLaMA-3 BlockPhi-3 Mini Block
LayerParamsOutputLayerParamsOutput
1sametokens
Input
1 × 2048tokens
Input
1 × 2048
2changed
dim, maxSeqLen, headDim
rope
Rope
rope
Rope
3changed
numEmbeddings, embeddingDim
embed
Embedding
525M1 × 2048 × 4096embed
Embedding
99M1 × 2048 × 3072
4changed
normalizedShape
attn_norm
Rms Norm
4.1K1 × 2048 × 4096norm_attn
Rms Norm
3.1K1 × 2048 × 3072
5changed
embedDim, numKVHeads
gqa
Grouped Query Attention
42M1 × 2048 × 4096attn
Grouped Query Attention
38M1 × 2048 × 3072
6sameresidual_1
Add
1 × 2048 × 4096residual_1
Add
1 × 2048 × 3072
7changed
normalizedShape
ffn_norm
Rms Norm
4.1K1 × 2048 × 4096norm_ffn
Rms Norm
3.1K1 × 2048 × 3072
8changed
embedDim, intermediateSize
swiglu_ffn
Swiglu
135M1 × 2048 × 4096ffn
Swiglu
75M1 × 2048 × 3072
9sameresidual_2
Add
1 × 2048 × 4096residual_2
Add
1 × 2048 × 3072
10addednorm_out
Rms Norm
3.1K1 × 2048 × 3072
11addedlm_head
Linear
1 × 2048 × 32064
12samehidden_state
Output
1 × 2048 × 4096output
Output
1 × 2048 × 32064

What this is not

Take it further

Open either graph in the editor, change it, and check it again: LLaMA-3 Block · Phi-3 Mini Block

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.

Related comparisons

LLaMA-3 Block vs Mixtral MoE Block
A dense feed-forward against a mixture of experts, at the block level.
LLaMA-3 Block vs DeepSeek-V3
What a frontier open model adds to the block everyone started from.
Transformer Block vs LLaMA-3 Block
What changed in the transformer block between 2017 and now.
Qwen3-8B vs Phi-3 Mini Block
What gets cut to make a small model small.