modelgrep

Best AI Agents for Full-Stack Apps

Match · Updated August 2026

The best LLM for building full-stack apps is Claude Fable 5, rated 1297 in Design Arena's full-stack agent arena — given a brief and a tool loop, and judged on what it actually shipped. Claude Opus 4.8 (1286) and GLM 5.2 (1276) round out the top three.

1297Agent Elo
59.9Intelligence
64 t/sSpeed
$10.00Input /M
1MContext
  1. 1anthropic logo
    claude-fable-5
    ReasoningToolsJSON+159.9 intel · $10.00/M · 64 t/s
    1297
    Agent Elo
  2. 2anthropic logo
    claude-opus-4.8
    ReasoningToolsJSON+155.7 intel · $5.00/M · 72 t/s
    1286
    Agent Elo
  3. 3z-ai logo
    glm-5.2
    ReasoningToolsJSON51.1 intel · $0.760/M · 135 t/s
    1276
    Agent Elo
  4. 4anthropic logo
    claude-sonnet-5
    ReasoningToolsJSON+153.4 intel · $2.00/M · 79 t/s
    1272
    Agent Elo
  5. 5anthropic logo
    claude-opus-4.6
    ReasoningToolsJSON+137.8 intel · $5.00/M · 1M ctx
    1263
    Agent Elo
  6. 6anthropic logo
    claude-sonnet-4.6
    ReasoningToolsJSON+135.9 intel · $3.00/M · 1M ctx
    1253
    Agent Elo
  7. 7qwen logo
    qwen3.7-max
    ReasoningToolsJSON46.0 intel · $1.48/M · 45 t/s
    1240
    Agent Elo
  8. 8google logo
    gemini-3.5-flash
    ReasoningToolsJSON+250.2 intel · $1.50/M · 270 t/s
    1235
    Agent Elo
  9. 9minimax logo
    minimax-m3
    ReasoningToolsJSON+144.4 intel · $0.300/M · 113 t/s
    1223
    Agent Elo
  10. 10moonshotai logo
    kimi-k2.7-code
    ReasoningToolsJSON+141.9 intel · $0.730/M · 154 t/s
    1219
    Agent Elo
  11. 11z-ai logo
    glm-5.1
    ReasoningToolsJSON40.2 intel · $0.966/M · 205K ctx
    1209
    Agent Elo
  12. 12moonshotai logo
    kimi-k2.6
    ReasoningToolsJSON+144.2 intel · $0.589/M · 262K ctx
    1199
    Agent Elo
  13. 13z-ai logo
    glm-5v-turbo
    ReasoningToolsJSON+134.5 intel · $1.20/M · 203K ctx
    1190
    Agent Elo
  14. 14z-ai logo
    glm-5
    ReasoningToolsJSON39.5 intel · $0.950/M · 205K ctx
    1164
    Agent Elo
  15. 15moonshotai logo
    kimi-k2.5
    ReasoningToolsJSON+135.4 intel · $0.570/M · 262K ctx
    1157
    Agent Elo
  16. 16openai logo
    gpt-5.5
    ReasoningToolsJSON+154.8 intel · $5.00/M · 1.1M ctx
    1128
    Agent Elo
  17. 17google logo
    gemini-3.1-pro-preview
    ReasoningToolsJSON+246.5 intel · $2.00/M · 136 t/s
    1110
    Agent Elo
  18. 18google logo
    gemini-3-flash-preview
    ReasoningToolsJSON+2$0.500/M · 1.0M ctx
    1104
    Agent Elo
  19. 19openai logo
    gpt-5.6-terra
    ReasoningToolsJSON+155.0 intel · $1.00/M · 83 t/s
    1096
    Agent Elo
  20. 20z-ai logo
    glm-4.7
    ReasoningToolsJSON33.7 intel · $0.400/M · 205K ctx
    1092
    Agent Elo
  21. 21openai logo
    gpt-5.2
    ReasoningToolsJSON+142.2 intel · $1.75/M · 81 t/s
    1084
    Agent Elo
  22. 22z-ai logo
    glm-4.6v
    ReasoningToolsJSON+111.0 intel · $0.300/M · 26 t/s
    1073
    Agent Elo
  23. 23z-ai logo
    glm-4.6
    ReasoningToolsJSON23.0 intel · $0.500/M · 36 t/s
    1073
    Agent Elo
  24. 24openai logo
    gpt-5.4
    ReasoningToolsJSON+151.4 intel · $2.50/M · 1.1M ctx
    1060
    Agent Elo
  25. 25x-ai logo
    grok-4.3
    ReasoningToolsJSON+137.6 intel · $1.25/M · 1M ctx
    1046
    Agent Elo

How this is ranked

Models ranked by Design Arena Elo in the full-stack agent arena — each one is given a brief and a tool loop and has to build a working application, with humans judging the results head to head. This measures something a static coding benchmark cannot: whether a model can hold a multi-file project together, recover from its own mistakes, and finish.

Frequently asked

Which LLM builds the best full-stack app?

The best LLM for building full-stack apps is Claude Fable 5, rated 1297 in Design Arena's full-stack agent arena — given a brief and a tool loop, and judged on what it actually shipped. Claude Opus 4.8 (1286) and GLM 5.2 (1276) round out the top three.

Which AI actually builds a working web app?

The best AI model for building full-stack apps is Claude Fable 5, rated 1297 in Design Arena's full-stack agent arena — given a brief and a tool loop, and judged on what it actually shipped. Claude Opus 4.8 (1286) and GLM 5.2 (1276) round out the top three.

What's a good alternative to Claude Fable 5?

Claude Opus 4.8 (1286) is the closest alternative on this metric, followed by GLM 5.2 (1276). See the full ranking above for the tradeoffs.

By maker

All rankings