A composite benchmark of real software-engineering ability — SWE-bench-style repo fixes, SciCode and terminal tasks — used to rank models for coding.
The Coding Index aggregates evaluations that resemble actual programming work: resolving real GitHub issues in existing repositories (SWE-bench Verified), writing scientific computing code (SciCode), and completing tasks in a live terminal. That makes it a far better predictor of coding usefulness than general-knowledge benchmarks.
Scores cluster tightly at the frontier — the top models often sit within a point or two of each other — so in practice speed, price, context window and tool-calling reliability decide which coding model to deploy.
The index is produced by Artificial Analysis, an independent evaluation lab; modelgrep refreshes scores daily and ranks every benchmarked model by it on the coding leaderboard.