English

NewsReal-SWE

New Benchmark "Real-SWE" Evaluates AI Models on Private, Real-World Enterprise Codebases

Real-SWE has been released as a new benchmark designed to evaluate frontier AI models on private, real-world, enterprise-grade codebases. Unlike benchmarks that rely on expert-generated or synthetic tasks, the tasks in Real-SWE are lifted verbatim from private production environments licensed from real-world companies.

The benchmark focuses on the complexity of actual software engineering, where changes often span multiple parts of an application. Each task requires agents to understand existing business logic, architecture, and coding patterns within a real operational context. According to the release, a typical instruction in Real-SWE is approximately 1,742 characters long and involves changes across multiple files, with an average of 11 files per task.

The evaluation is conducted in isolated sandboxes using native harnesses to reflect how engineers work in practice. This approach assesses the combination of the model and the harness rather than the model in isolation. The tasks are designed to be slightly underspecified, similar to benchmarks like DeepSWE and Terminal Bench, requiring agents to discover implementation details within the codebase and surrounding tools.

Sources

  1. Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases (Hacker News Frontpage, 2026-09-12)