English

Product LaunchesCueGemma 4 E4B

Cue AI Reduces Voice Input Latency by 44% with Local Execution of Gemma 4 E4B

This article is a translation. Read the Japanese original

Cue is a voice-operated AI agent for desktops. When a user presses a hotkey and speaks, the agent performs tasks such as text input in any application or selecting tools by reading the screen. The goal is to enable natural computer operation via voice, making the keyboard optional.

Previously, Cue utilized a cloud-based alignment step, which introduced a delay of 800–900ms per interaction. Furthermore, other tested models tended to over-edit casual spoken language into a formal style, erasing the speaker's personality. The team sought text that could be read in the user's own tone rather than the model's style.

With the integration of Gemma 4 E4B, latency was reduced by 44%. The median latency dropped from 876ms to 488ms. As a result, the alignment step now fits within the perceptual budget of "feeling faster than typing."

Measurements were conducted on Apple Silicon (M-series) via Ollama. The benchmark consisted of 227 real-voice samples, including English and multilingual mixes. The data is from May 2026, based on approximately four weeks of data from beta users.

Cue's pipeline is simple. When a user speaks, the audio is converted to raw text via cloud STT. This text is then sent to a local LLM running Gemma 4, which applies formatting rules using a system prompt of approximately 400 tokens. The aligned text is then inserted at the cursor position via native OS APIs.

The entire process is completed in under one second. This improvement led to an approximately 30% increase in the amount of text input per person. Users who previously only entered short messages began inputting longer sentences, and more users continued using voice mode even for time-constrained tasks.

Cue offers unlimited text input for all users, including those on the free plan. Because Gemma 4 E4B runs entirely on-device, the marginal cost of the alignment step has become zero. This makes features that previously required limits or monetization economically viable even for the free plan.

Initially, the team planned to use Gemma as an offline fallback, assuming that larger models would offer higher accuracy. However, the 227-sample benchmark revealed that Gemma 4 possesses the appropriate capacity to handle formatting and corrections without over-editing. In the context of text alignment, this characteristic is not a limitation but exactly the desired behavior.

In Cue's deployment, only the base model and prompt engineering are used. The context of the active application (app name, field type, placeholder text) is injected into the prompt, allowing the model to recognize, for example, whether it is drafting a Slack message.


Source: Cue AI (HN 32pt, 2 comments) (HN Search (backfill), 2026-07-21)