English

Product LaunchesCanary

Canary Releases AI QA Tool That Automatically Generates Tests from PR Diffs, Outperforming Major LLMs in Proprietary Benchmark

This article is a translation. Read the Japanese original

After connecting to a codebase, Canary understands the application structure, including routes, controllers, and validation logic. When a PR is pushed, the tool reads the diff, interprets the intent of the changes, and executes end-to-end tests against a preview environment.

Results are reflected as comments on the PR, featuring screen recordings of the changes and flags for areas that do not behave as expected. It is also possible to trigger specific user flow tests via PR comments.

Generated tests can be moved into a regression suite, and the tool includes functionality to generate and schedule test suites for continuous execution based on natural language prompts. In one instance at a construction tech company, the tool detected a discrepancy of approximately $1,600 in an invoicing flow.

The company stated that because QA spans diverse modalities—such as source code, DOM, and device emulators—it is difficult to achieve with a single foundation model. They also noted the necessity of infrastructure such as custom browser fleets and ephemeral environments.

To evaluate performance, the company released "QA-Bench v0," a benchmark for code verification. This benchmark compares Canary against models such as GPT 5.4 and Claude Code (Opus 4.6) using 35 real PRs from four repositories, including Grafana and Mattermost.

In terms of coverage, Canary scored 11 points higher than GPT 5.4, 18 points higher than Claude Code, and 26 points higher than Sonnet 4.6. Detailed methodologies and breakdowns by repository are summarized in the benchmark report.


Source: Launch HN: Canary (YC W26) – AI QA that understands your code(HN 58pt・26コメント) (HN Search (backfill), 2026-03-20)