Comparisons / LLaMA-3 Block vs Phi-3 Mini Block
LLaMA-3 Block vs Phi-3 Mini Block
A block from a large model against a whole small one.
Phi-3 Mini Block has 392M fewer parameters than LLaMA-3 Block: 2 layers added, 6 changed.
LLaMA-3 Block
- Layers
- 8
- Parameters
- 703M
- Input
- 1 × 2048
- Output
- 1 × 2048 × 4096
- Forward-passes
- yes
- Est. train cost
- $14.27
Phi-3 Mini Block
- Layers
- 10
- Parameters
- 310M
- Input
- 1 × 2048
- Output
- 1 × 2048 × 32064
- Forward-passes
- yes
- Est. train cost
- $16.76
The deltas
Every number is Phi-3 Mini Block relative to LLaMA-3 Block.
Which GPUs each one fits
Memory for the graph at the input shape both declare. A highlighted row is a card one of them fits and the other does not, which is the difference that decides a purchase.
| GPU | LLaMA-3 Block | Phi-3 Mini Block |
|---|---|---|
| T4 16GB | fits | fits |
| A100 40GB | fits | fits |
| H100 80GB | fits | fits |
Layer by layer
Aligned in topological order. 4 of 12 rows are the same layer with the same parameters.
Hide all 12 rows
| LLaMA-3 Block | Phi-3 Mini Block | ||||||
|---|---|---|---|---|---|---|---|
| Layer | Params | Output | Layer | Params | Output | ||
| 1 | same | tokens Input | 1 × 2048 | tokens Input | 1 × 2048 | ||
| 2 | changed dim, maxSeqLen, headDim | rope Rope | rope Rope | ||||
| 3 | changed numEmbeddings, embeddingDim | embed Embedding | 525M | 1 × 2048 × 4096 | embed Embedding | 99M | 1 × 2048 × 3072 |
| 4 | changed normalizedShape | attn_norm Rms Norm | 4.1K | 1 × 2048 × 4096 | norm_attn Rms Norm | 3.1K | 1 × 2048 × 3072 |
| 5 | changed embedDim, numKVHeads | gqa Grouped Query Attention | 42M | 1 × 2048 × 4096 | attn Grouped Query Attention | 38M | 1 × 2048 × 3072 |
| 6 | same | residual_1 Add | 1 × 2048 × 4096 | residual_1 Add | 1 × 2048 × 3072 | ||
| 7 | changed normalizedShape | ffn_norm Rms Norm | 4.1K | 1 × 2048 × 4096 | norm_ffn Rms Norm | 3.1K | 1 × 2048 × 3072 |
| 8 | changed embedDim, intermediateSize | swiglu_ffn Swiglu | 135M | 1 × 2048 × 4096 | ffn Swiglu | 75M | 1 × 2048 × 3072 |
| 9 | same | residual_2 Add | 1 × 2048 × 4096 | residual_2 Add | 1 × 2048 × 3072 | ||
| 10 | added | — | norm_out Rms Norm | 3.1K | 1 × 2048 × 3072 | ||
| 11 | added | — | lm_head Linear | 1 × 2048 × 32064 | |||
| 12 | same | hidden_state Output | 1 × 2048 × 4096 | output Output | 1 × 2048 × 32064 | ||
What this is not
- Parameter counts are derived from the graph, not read from a checkpoint. They are exact for a graph that is fully specified and approximate for one that is not.
- Cost and GPU fit are estimates from the graph under one set of assumptions, not measurements of a run.
Take it further
Open either graph in the editor, change it, and check it again: LLaMA-3 Block · Phi-3 Mini Block
Compare any two models of your own, including anything on Hugging Face: the comparison tool.
Machine-readable: this page as markdown ·
the pair index ·
POST https://www.neurarch.com/api/v1/plan for a graph of your own.