English

Model ReleasesLinumJiT-DDT

Linum Releases JiT-DDT: Pixel-Space Text-to-Image Training with 3.6× Faster Speed

Linum has released JiT-DDT, a 2.5B-parameter pixel-space text-to-image diffusion transformer. The research artifact uses an encoder-decoder architecture to predict clean images directly, eliminating the need for a separate Variational Autoencoder (VAE) typically used in Latent Diffusion Models (LDMs).

Compared to the company's Linum v2 baseline, JiT-DDT trains 3.6× faster while generating images at 512×512 resolution, which is 4× the pixel count of the 256×256 baseline. To manage high-dimensional inputs, the architecture employs a single-stream Diffusion Transformer (DiT) that processes visual and text tokens together, and utilizes PixelREPA—a representation-alignment technique that aligns the encoder's hidden states with DINOv3 features.

The model incorporates a specialized noise schedule to improve the recovery of fine-grained details and uses Qwen3.5-4B as the text encoder. The weights and code are available under the Apache 2.0 license. The implementation is compatible with CUDA GPUs with approximately 24 GB of memory.

Sources

  1. Training Text-to-Image Models 3.6× Faster (Hacker News Frontpage, 2026-09-16)
  2. Model code