N Neurarch Architectures Checks Docs Open the app

Architectures / Generative

๐ŸŒ€ DiT-XL/2

Diffusion Transformer โ€” replaces the UNet denoiser with a ViT backbone conditioned on timestep + class via adaLN-Zero (Peebles 2023)

Layers
204
Parameters
670.68M
Input
4 ร— 32 ร— 32
Output
256 ร— 32
Verifier
Clean

Every number on this page is computed from the graph by the same functions the app runs, not written by hand.

Open DiT-XL/2 on the canvas Free, no account needed

When to pick it

Pick when you want the modern transformer-based image-generation backbone (used by Sora/SD3-era models) instead of a convolutional UNet denoiser.

Structure

204 layers. Output shapes are propagated from the input shape, batch dimension excluded.

LayerTypeParametersOutput shape
1noisy_latentInputshape=[4, 32, 32]4 ร— 32 ร— 32
2timestep_+_classInputshape=[1, 2]1 ร— 2
3patchify_2x2Patch EmbedembedDim=1152, patchSize=2256 ร— 1152
4cond_embedEmbedding1 ร— 2 ร— 1152
5pos_embedPositional EncodingembedDim=1152, maxLen=256256 ร— 1152
6adaLN_1LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
7norm1_1LayerNormnormalizedShape=1152256 ร— 1152
8self_attn_1Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
9residual1_1Add256 ร— 1152
10norm2_1LayerNormnormalizedShape=1152256 ร— 1152
11mlp_1Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
12residual2_1Add256 ร— 1152
13adaLN_2LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
14norm1_2LayerNormnormalizedShape=1152256 ร— 1152
15self_attn_2Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
16residual1_2Add256 ร— 1152
17norm2_2LayerNormnormalizedShape=1152256 ร— 1152
18mlp_2Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
19residual2_2Add256 ร— 1152
20adaLN_3LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
21norm1_3LayerNormnormalizedShape=1152256 ร— 1152
22self_attn_3Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
23residual1_3Add256 ร— 1152
24norm2_3LayerNormnormalizedShape=1152256 ร— 1152
25mlp_3Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
26residual2_3Add256 ร— 1152
27adaLN_4LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
28norm1_4LayerNormnormalizedShape=1152256 ร— 1152
29self_attn_4Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
30residual1_4Add256 ร— 1152
31norm2_4LayerNormnormalizedShape=1152256 ร— 1152
32mlp_4Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
33residual2_4Add256 ร— 1152
34adaLN_5LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
35norm1_5LayerNormnormalizedShape=1152256 ร— 1152
36self_attn_5Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
37residual1_5Add256 ร— 1152
38norm2_5LayerNormnormalizedShape=1152256 ร— 1152
39mlp_5Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
40residual2_5Add256 ร— 1152
41adaLN_6LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
42norm1_6LayerNormnormalizedShape=1152256 ร— 1152
43self_attn_6Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
44residual1_6Add256 ร— 1152
45norm2_6LayerNormnormalizedShape=1152256 ร— 1152
46mlp_6Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
47residual2_6Add256 ร— 1152
48adaLN_7LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
49norm1_7LayerNormnormalizedShape=1152256 ร— 1152
50self_attn_7Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
51residual1_7Add256 ร— 1152
52norm2_7LayerNormnormalizedShape=1152256 ร— 1152
53mlp_7Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
54residual2_7Add256 ร— 1152
55adaLN_8LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
56norm1_8LayerNormnormalizedShape=1152256 ร— 1152
57self_attn_8Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
58residual1_8Add256 ร— 1152
59norm2_8LayerNormnormalizedShape=1152256 ร— 1152
60mlp_8Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
61residual2_8Add256 ร— 1152
62adaLN_9LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
63norm1_9LayerNormnormalizedShape=1152256 ร— 1152
64self_attn_9Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
65residual1_9Add256 ร— 1152
66norm2_9LayerNormnormalizedShape=1152256 ร— 1152
67mlp_9Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
68residual2_9Add256 ร— 1152
69adaLN_10LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
70norm1_10LayerNormnormalizedShape=1152256 ร— 1152
71self_attn_10Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
72residual1_10Add256 ร— 1152
73norm2_10LayerNormnormalizedShape=1152256 ร— 1152
74mlp_10Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
75residual2_10Add256 ร— 1152
76adaLN_11LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
77norm1_11LayerNormnormalizedShape=1152256 ร— 1152
78self_attn_11Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
79residual1_11Add256 ร— 1152
80norm2_11LayerNormnormalizedShape=1152256 ร— 1152
81mlp_11Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
82residual2_11Add256 ร— 1152
83adaLN_12LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
84norm1_12LayerNormnormalizedShape=1152256 ร— 1152
85self_attn_12Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
86residual1_12Add256 ร— 1152
87norm2_12LayerNormnormalizedShape=1152256 ร— 1152
88mlp_12Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
89residual2_12Add256 ร— 1152
90adaLN_13LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
91norm1_13LayerNormnormalizedShape=1152256 ร— 1152
92self_attn_13Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
93residual1_13Add256 ร— 1152
94norm2_13LayerNormnormalizedShape=1152256 ร— 1152
95mlp_13Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
96residual2_13Add256 ร— 1152
97adaLN_14LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
98norm1_14LayerNormnormalizedShape=1152256 ร— 1152
99self_attn_14Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
100residual1_14Add256 ร— 1152
101norm2_14LayerNormnormalizedShape=1152256 ร— 1152
102mlp_14Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
103residual2_14Add256 ร— 1152
104adaLN_15LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
105norm1_15LayerNormnormalizedShape=1152256 ร— 1152
106self_attn_15Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
107residual1_15Add256 ร— 1152
108norm2_15LayerNormnormalizedShape=1152256 ร— 1152
109mlp_15Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
110residual2_15Add256 ร— 1152
111adaLN_16LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
112norm1_16LayerNormnormalizedShape=1152256 ร— 1152
113self_attn_16Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
114residual1_16Add256 ร— 1152
115norm2_16LayerNormnormalizedShape=1152256 ร— 1152
116mlp_16Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
117residual2_16Add256 ร— 1152
118adaLN_17LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
119norm1_17LayerNormnormalizedShape=1152256 ร— 1152
120self_attn_17Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
121residual1_17Add256 ร— 1152
122norm2_17LayerNormnormalizedShape=1152256 ร— 1152
123mlp_17Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
124residual2_17Add256 ร— 1152
125adaLN_18LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
126norm1_18LayerNormnormalizedShape=1152256 ร— 1152
127self_attn_18Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
128residual1_18Add256 ร— 1152
129norm2_18LayerNormnormalizedShape=1152256 ร— 1152
130mlp_18Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
131residual2_18Add256 ร— 1152
132adaLN_19LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
133norm1_19LayerNormnormalizedShape=1152256 ร— 1152
134self_attn_19Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
135residual1_19Add256 ร— 1152
136norm2_19LayerNormnormalizedShape=1152256 ร— 1152
137mlp_19Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
138residual2_19Add256 ร— 1152
139adaLN_20LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
140norm1_20LayerNormnormalizedShape=1152256 ร— 1152
141self_attn_20Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
142residual1_20Add256 ร— 1152
143norm2_20LayerNormnormalizedShape=1152256 ร— 1152
144mlp_20Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
145residual2_20Add256 ร— 1152
146adaLN_21LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
147norm1_21LayerNormnormalizedShape=1152256 ร— 1152
148self_attn_21Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
149residual1_21Add256 ร— 1152
150norm2_21LayerNormnormalizedShape=1152256 ร— 1152
151mlp_21Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
152residual2_21Add256 ร— 1152
153adaLN_22LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
154norm1_22LayerNormnormalizedShape=1152256 ร— 1152
155self_attn_22Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
156residual1_22Add256 ร— 1152
157norm2_22LayerNormnormalizedShape=1152256 ร— 1152
158mlp_22Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
159residual2_22Add256 ร— 1152
160adaLN_23LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
161norm1_23LayerNormnormalizedShape=1152256 ร— 1152
162self_attn_23Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
163residual1_23Add256 ร— 1152
164norm2_23LayerNormnormalizedShape=1152256 ร— 1152
165mlp_23Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
166residual2_23Add256 ร— 1152
167adaLN_24LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
168norm1_24LayerNormnormalizedShape=1152256 ร— 1152
169self_attn_24Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
170residual1_24Add256 ร— 1152
171norm2_24LayerNormnormalizedShape=1152256 ร— 1152
172mlp_24Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
173residual2_24Add256 ร— 1152
174adaLN_25LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
175norm1_25LayerNormnormalizedShape=1152256 ร— 1152
176self_attn_25Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
177residual1_25Add256 ร— 1152
178norm2_25LayerNormnormalizedShape=1152256 ร— 1152
179mlp_25Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
180residual2_25Add256 ร— 1152
181adaLN_26LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
182norm1_26LayerNormnormalizedShape=1152256 ร— 1152
183self_attn_26Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
184residual1_26Add256 ร— 1152
185norm2_26LayerNormnormalizedShape=1152256 ร— 1152
186mlp_26Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
187residual2_26Add256 ร— 1152
188adaLN_27LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
189norm1_27LayerNormnormalizedShape=1152256 ร— 1152
190self_attn_27Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
191residual1_27Add256 ร— 1152
192norm2_27LayerNormnormalizedShape=1152256 ร— 1152
193mlp_27Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
194residual2_27Add256 ร— 1152
195adaLN_28LinearoutFeatures=6912, inFeatures=11521 ร— 2 ร— 6912
196norm1_28LayerNormnormalizedShape=1152256 ร— 1152
197self_attn_28Multi-Head AttentionembedDim=1152, numHeads=16256 ร— 1152
198residual1_28Add256 ร— 1152
199norm2_28LayerNormnormalizedShape=1152256 ร— 1152
200mlp_28Feed ForwardembedDim=1152, ffDim=4608256 ร— 1152
201residual2_28Add256 ร— 1152
202final_normLayerNormnormalizedShape=1152256 ร— 1152
203unpatchifyLinearoutFeatures=32, inFeatures=1152256 ร— 32
204predicted_noiseOutput256 ร— 32

What the verifier says

The same 41 structural checks that run on every edit in the app, on this graph.

infoAt 28 stacked attention layers, residual-branch outputs add up; unscaled init lets activation variance grow with depth. GPT-2/LLaMA-family models scale the residual projections by depth (N(0, 0.02 / โˆš(2L))). Fix: Scale residual output projections by depth: nn.init.normal_(w, std=0.02 / math.sqrt(2 * n_layers)) (self_attn_1)
deep-attention-default-init

The PyTorch it exports

Generated from the graph above. First 46 lines; the app exports the whole file, plus the training loop, the data contract and a deploy bundle.

# Architecture designed with Neurarch: https://neurarch.com
# PyTorch: compatible with Python 3.8+ and torch>=1.12
# Colab: pip install torch torchvision  (usually pre-installed)

import torch
import torch.nn as nn
import torch.nn.functional as F
from typing import Tuple

class DiT_XL2(nn.Module):
    def __init__(self):
        super().__init__()

        self.patchEmbed_1 = nn.Conv2d(4, 1152, kernel_size=2, stride=2)  # Patch embedding (ViT-style)
        self.embedding_1 = nn.Embedding(1000, 1152)
        self.linear_1 = nn.Linear(1152, 6912)
        self.layerNorm_1 = nn.LayerNorm(1152)
        self.multiHeadAttention_1 = nn.MultiheadAttention(embed_dim=1152, num_heads=16, batch_first=True)
        self.layerNorm_2 = nn.LayerNorm(1152)
        self.feedForward_1 = nn.Sequential(
            nn.Linear(1152, 4608),
            nn.ReLU(),
            nn.Linear(4608, 1152)
        )
        self.linear_2 = nn.Linear(1152, 6912)
        self.layerNorm_3 = nn.LayerNorm(1152)
        self.multiHeadAttention_2 = nn.MultiheadAttention(embed_dim=1152, num_heads=16, batch_first=True)
        self.layerNorm_4 = nn.LayerNorm(1152)
        self.feedForward_2 = nn.Sequential(
            nn.Linear(1152, 4608),
            nn.ReLU(),
            nn.Linear(4608, 1152)
        )
        self.linear_3 = nn.Linear(1152, 6912)
        self.layerNorm_5 = nn.LayerNorm(1152)
        self.multiHeadAttention_3 = nn.MultiheadAttention(embed_dim=1152, num_heads=16, batch_first=True)
        self.layerNorm_6 = nn.LayerNorm(1152)
        self.feedForward_3 = nn.Sequential(
            nn.Linear(1152, 4608),
            nn.ReLU(),
            nn.Linear(4608, 1152)
        )
        self.linear_4 = nn.Linear(1152, 6912)
        self.layerNorm_7 = nn.LayerNorm(1152)
        self.multiHeadAttention_4 = nn.MultiheadAttention(embed_dim=1152, num_heads=16, batch_first=True)
        self.layerNorm_8 = nn.LayerNorm(1152)

For agents

This architecture is machine-readable end to end. An agent can list the set, fetch this graph, edit it, and have the edit verified before any GPU time is spent.

Also in Generative

๐ŸŽจ Diffusion UNet
Stable-Diffusion-style noise predictor โ€” latent UNet with cross-attention to a text embedding
15 layers ยท 6.69M