JevBench has been introduced as a reproducible benchmark designed to evaluate typed decision models. The benchmark measures four key axes—Intelligence, Calibration, Speed, and Cost—applying a 25% weight to each via a geometric mean to calculate the official JevBench Score.
JevBench introduced as a reproducible benchmark for typed decision models
To prevent models that are fast and cheap but perform no better than random guessing from ranking highly, the scoring system includes a penalty for Intelligence scores below 50. Specifically, if Intelligence falls below this threshold, the final score is multiplied by (I / 50)^2.
The benchmark also implements specific scoping for different difficulty tiers. While the "Easy" scope focuses on Intelligence, the "Hard" tier measures all four metrics on a subset of 220 high-difficulty decisions. The evaluation is designed to account for regional hosting differences, such as EU-based inference, to provide transparency in deployment environments.
The project is open source and available via GitHub. In recent tests, the classifier.dev "fast tier" (using the Jev 1.13.0 model) scored 97.3% on the judge tier and 70.5% on the hard tier in specific comparison metrics.
Sources
- Show HN: JevBench, a reproducible benchmark for typed decision models (Hacker News Frontpage, 2026-09-22)