N Neurarch Architectures Models Checks Data Docs Open the app

Comparisons / 1D CNN + LSTM vs PatchTST

1D CNN + LSTM vs PatchTST

Convolution-plus-recurrence against patched attention for time series.

PatchTST has 84K more parameters than 1D CNN + LSTM: 14 layers added, 9 removed, 4 changed.

Baseline

1D CNN + LSTM

Layers
12
Parameters
312K
Input
12 × 5000
Output
5
Forward-passes
yes
Est. train cost
$0.059
Compared

PatchTST

Layers
17
Parameters
395K
Input
1 × 22 × 1000
Output
4
Forward-passes
yes
Est. train cost
$0.043

The deltas

Every number is PatchTST relative to 1D CNN + LSTM.

Parameters
+84K (+26.8%)
Layers
+5
Added
14
Removed
9
Changed
4
Unchanged
1

Which GPUs each one fits

Each side is measured at its own declared input (12 × 5000 against 1 × 22 × 1000). Both columns are right about their own model; the difference between them is not a fact about the designs.

GPU1D CNN + LSTMPatchTST
T4 16GBfitsfits
A100 40GBfitsfits
H100 80GBfitsfits

Layer by layer

Aligned in topological order. 1 of 28 rows are the same layer with the same parameters.

Hide all 28 rows
1D CNN + LSTMPatchTST
LayerParamsOutputLayerParamsOutput
1changed
shape
ts_window
Input
12 × 5000ts_window
Input
1 × 22 × 1000
2removedconv1
Conv1d
51264 × 5000
3addedpatch_embed
Patch Embed
98K62 × 128
4addedpos_embed
Positional Encoding
62 × 128
5changed
type, normalizedShape
bn
Batch Norm
12864 × 5000norm
Layer Norm
25662 × 128
6removedact
Relu
64 × 5000
7removedpool
Maxpool1d
64 × 2500
8removedconv2
Conv1d
768128 × 2500
9addedself_attn
Multi Head Attention
66K62 × 128
10addedresidual
Add
62 × 128
11changed
type
bn
Batch Norm
256128 × 2500norm
Layer Norm
25662 × 128
12removedact
Relu
128 × 2500
13removedpool
Maxpool1d
128 × 1250
14removedto_timesteps
Permute
1250 × 128
15removedlstm
Lstm
199K128
16removeddrop
Dropout
128
17addeddense
Feed Forward
66K62 × 128
18addedresidual
Add
62 × 128
19addednorm
Layer Norm
25662 × 128
20addedself_attn
Multi Head Attention
66K62 × 128
21addedresidual
Add
62 × 128
22addednorm
Layer Norm
25662 × 128
23addeddense
Feed Forward
66K62 × 128
24addedresidual
Add
62 × 128
25addednorm
Layer Norm
25662 × 128
26addedflatten
Flatten
7936
27changed
outFeatures
classifier
Linear
5classifier
Linear
4
28samelogits
Output
5logits
Output
4

What this is not

Take it further

Open either graph in the editor, change it, and check it again: 1D CNN + LSTM · PatchTST

Compare any two models of your own, including anything on Hugging Face: the comparison tool.

Machine-readable: this page as markdown · the pair index · POST https://www.neurarch.com/api/v1/plan for a graph of your own.