PSSA is a small language model written from scratch in Rust without the use of machine learning frameworks like PyTorch or TensorFlow. Rather than using the attention mechanisms found in standard Transformer architectures, PSSA employs a recurrent state-space layer and a bank of episodic memories to manage context.
PSSA: A Non-Transformer Language Model Written in Rust
According to the project documentation, PSSA shows performance advantages when compared to a Transformer using matched parameters and the same corpus. PSSA learns faster and generates text approximately twelve times quicker on the same CPU. This efficiency is attributed to its architectural design: while a Transformer's computational cost grows quadratically with sequence length due to re-reading the entire context at every step, PSSA maintains a fixed-size state that allows costs to grow linearly.
In experiments using 12.7M tokens of WikiText-103, PSSA achieved a training cross-entropy of 3.98, compared to 4.43 for the baseline Transformer. On held-out, unseen text, PSSA also demonstrated a lower perplexity of 53.7, while the Transformer scored 83.7. The project documentation notes that PSSA reaches specific loss thresholds significantly faster than the Transformer.
The current implementation includes a custom CLI for training, checkpointing, and dataset management. The developers noted that generation quality remains limited at the current scale and requires further testing with more significant compute resources to validate performance at larger parameter sizes.
Sources
- PSSA: A non-transformer language model written from scratch in Rust (Hacker News Frontpage, 2026-09-30)