RADAR ·
Best model solves 38.8 percent of real company coding tasks in Real-SWE
Real-SWE, a benchmark by Snagnik Das, Siddhant Paliwal and Janak Sunil, measures AI models on the private codebases of real companies. Tasks come from licensed production codebases and cover changes that affect how a business runs, such as invoicing, tax calculation and customer migration. Because the code and its solutions are not on the public internet, models have to work out company-specific rules from the codebase and connected tools. The codebases include an events app with more than 200,000 users and a consumer fintech platform.
Each model was tested with its own native harness, and every task was run eight times. Fable 5.1 with Claude Code ranked first, resolving 38.8 percent of tasks. GPT-6 Astra with Codex CLI reached 33.8 percent and Gemini 3.8 Flash with Gemini CLI 31.2 percent. Kimi K3 scored 18.8 percent and GPT-5.6 Sol 16.2 percent.
Source: Hacker News