A new paper from Ornn Data, "The Economics of Open-Weight Inference," argues that the rise of open-weight models is challenging the assumption that newer hardware generations immediately render older GPUs obsolete.
The study finds that open-weight models can be significantly more cost-effective than closed-model alternatives. When standardized by intelligence, the cheapest qualifying open-weight model completes tasks at approximately one-fifth the cost of comparable closed models. For certain workloads, such as the sparse mixture-of-experts model gpt-oss-120b, self-hosting on rented hardware can make older NVIDIA A100 GPUs more cost-efficient than the newer H100 for specific task completions.
This cost advantage is particularly relevant for compute-intensive workloads that can tolerate higher latency, including long-running agents, batch evaluation, and reinforcement learning. Because these workloads are hardware agnostic, they can be routed to the most cost-efficient hardware available, providing sustained demand for older capacity.
Ornn's rental data highlights this trend, showing that the five-year term price for the A100 retains about 80 percent of its one-month term price. In contrast, the Hopper and Blackwell families show steeper depreciation curves, with retention rates between 44 and 60 percent. While newer GPUs offer superior performance for demanding tasks, the economic utility of older generations is being preserved by the growing demand for efficient, open-weight inference.
Sources
- The Economics of Open-Weight Inference (Hacker News Frontpage, 2026-09-22)
- Ornn