New small-scale decision models, named Jeff, have been released to perform zero-shot classification with high speed and low latency. Based on fine-tuned versions of Qwen3.5 (0.8B and 2B) and Gemma 4 (2B), these models are designed to be integrated into local code using the same request format as Jev.
Model ReleasesAlibaba CloudGemmaJeff
Jeff: Small, Fast Decision Models Based on Qwen and Gemma 4 for Zero-Shot Classification
Unlike generative models, Jeff returns calibrated probabilities for specified options through a single forward pass without generating text or requiring parsing. In performance tests, the models achieved decision speeds of approximately 22 ms on an NVIDIA RTX PRO 6000 and 28 ms on an Apple M4 Max (using MLX).
The models support zero-shot capabilities, meaning they can classify arbitrary categories—such as user intent, moderation labels, or game moves—by simple description, even if those categories were not present in the training data. While the small size means reasoning capabilities are lower than much larger models like Jev, the developers noted that short fine-tuning on specific examples can significantly increase accuracy; for instance, a voice-navigation fine-tune improved accuracy from 31.7% to 95.8% in under half an hour on a single GPU.
The project was developed entirely on local hardware using synthetic training data generated by an open model. The models are available on Hugging Face under the Apache 2.0 license, with the code released under the MIT license.
Sources
- Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms (Hacker News Frontpage, 2026-09-28)