Linum has released JiT-DDT, a 2.5B-parameter pixel-space text-to-image diffusion transformer. The research artifact uses an encoder-decoder architecture to predict clean images directly, eliminating the need for a separate Variational Autoencoder (VAE) typically used in Latent Diffusion Models (LDMs).
Linum Releases JiT-DDT: Pixel-Space Text-to-Image Training with 3.6× Faster Speed
Compared to the company's Linum v2 baseline, JiT-DDT trains 3.6× faster while generating images at 512×512 resolution, which is 4× the pixel count of the 256×256 baseline. To manage high-dimensional inputs, the architecture employs a single-stream Diffusion Transformer (DiT) that processes visual and text tokens together, and utilizes PixelREPA—a representation-alignment technique that aligns the encoder's hidden states with DINOv3 features.
The model incorporates a specialized noise schedule to improve the recovery of fine-grained details and uses Qwen3.5-4B as the text encoder. The weights and code are available under the Apache 2.0 license. The implementation is compatible with CUDA GPUs with approximately 24 GB of memory.
Sources
- Training Text-to-Image Models 3.6× Faster (Hacker News Frontpage, 2026-09-16)
- Model code