WebLLM is an engine that leverages WebGPU for hardware acceleration to execute large language model (LLM) inference directly within the browser. Since all processing is completed locally in the browser, no server-side support is required.
WebLLM Released: An Engine for In-Browser LLMinference Leveraging WebGPU
This article is a translation. Read the Japanese original
This project is positioned as a companion project to MLC LLM, which enables the deployment of LLMs across various hardware environments.
WebLLM offers full compatibility with the OpenAI API. In addition to streaming and JSON mode, it provides features such as logit-level control and seeding. It also supports advanced JSON generation capabilities optimized using WebAssembly.
Supported models include Llama 3, Phi 3, Gemma, Mistral, and Qwen. Users can integrate the engine as a package via NPM, Yarn, or CDN. Furthermore, it supports Web Workers, Service Workers, and Chrome extensions.
Source: WebLLM: high-performance in-browser LLM inference engine (Hacker News Frontpage, 2026-09-02)