Cumulus Labs has released "IonRouter," an inference API for open-source and fine-tuned models. It is compatible with existing OpenAI client code and can be connected simply by changing the base URL.
Product LaunchesCumulus LabsIonRouter
Cumulus Labs Releases IonRouter API Service Featuring GH200-Optimized Inference Engine
This article is a translation. Read the Japanese original
The company pointed out a problem where inference providers were limited to a choice between "fast but expensive" or "cheap but cumbersome to configure." To address this, they developed "IonAttention," a C++ inference runtime specifically optimized for the GH200 memory architecture.
This runtime effectively utilizes the 900GB/s CPU-GPU link and 452GB of LPDDR5X memory. Key features include the use of hardware cache coherence and technology that allows CUDA graphs to operate at zero cost as if they had dynamic parameters.
In multimodal processing, the service recorded a performance of 588 tok/s, compared to 298 tok/s from Together AI. However, the company acknowledged that the p50 latency of 1.46 seconds is inferior to competitors and stated that they are currently working on improvements.
The pricing model is based on token usage, with no idle costs. For GPT-OSS-120B, the input token price is set at $0.02 and the output token price at $0.095.
Source: Launch HN: IonRouter (YC W26) – High-throughput, low-cost inference(HN 72pt・37コメント) (HN Search (backfill), 2026-03-13)