modelgrep

Best LLMs for Science — Qwen

Match · Updated August 2026

The best Qwen model for science is Qwen3.7 Max, scoring 92.3% on GPQA Diamond — graduate-level physics, chemistry and biology. Qwen3.7 Plus (90.0%) and Qwen3.5 397B A17B (89.3%) round out the top three.

92.3%GPQA
46.0Intelligence
44 t/sSpeed
$1.48Input /M
1MContext
  1. 1qwen logo
    qwen3.7-max
    ReasoningToolsJSON46.0 intel · $1.48/M · 44 t/s
    92.3%
    GPQA
  2. 2qwen logo
    qwen3.7-plus
    ReasoningToolsJSON+139.0 intel · $0.320/M · 12 t/s
    90.0%
    GPQA
  3. 3qwen logo
    qwen3.5-397b-a17b
    ReasoningToolsJSON+133.7 intel · $0.390/M · 69 t/s
    89.3%
    GPQA
  4. 4qwen logo
    qwen3.6-max-preview
    ReasoningToolsJSON40.0 intel · $1.03/M · 53 t/s
    88.8%
    GPQA
  5. 5qwen logo
    qwen3.6-plus
    ReasoningToolsJSON+139.6 intel · $0.325/M · 37 t/s
    88.2%
    GPQA
  6. 6qwen logo
    qwen3-max-thinking
    ReasoningToolsJSON31.7 intel · $0.780/M · 262K ctx
    86.1%
    GPQA
  7. 7qwen logo
    qwen3.5-27b
    ReasoningToolsJSON+133.8 intel · $0.195/M · 38 t/s
    85.8%
    GPQA
  8. 8qwen logo
    qwen3.5-122b-a10b
    ReasoningToolsJSON+132.3 intel · $0.260/M · 130 t/s
    85.7%
    GPQA
  9. 9qwen logo
    qwen3.5-35b-a3b
    ReasoningToolsJSON+129.3 intel · $0.140/M · 217 t/s
    84.5%
    GPQA
  10. 10qwen logo
    qwen3.6-27b
    ReasoningToolsJSON+137.1 intel · $0.289/M · 50 t/s
    84.2%
    GPQA
  11. 11qwen logo
    qwen3.6-35b-a3b
    ReasoningToolsJSON+131.6 intel · $0.140/M · 154 t/s
    84.1%
    GPQA
  12. 12qwen logo
    qwen3.5-9b
    ReasoningToolsJSON+121.4 intel · $0.100/M · 70 t/s
    80.6%
    GPQA
  13. 13qwen logo
    qwen3-max
    ToolsJSON24.0 intel · $0.780/M · 262K ctx
    76.4%
    GPQA
  14. 14qwen logo
    qwen3-next-80b-a3b-instruct
    ToolsJSON13.7 intel · $0.090/M · 205 t/s
    73.8%
    GPQA
  15. 15qwen logo
    qwen3-coder-next
    ToolsJSON21.1 intel · $0.120/M · 133 t/s
    73.7%
    GPQA
  16. 16qwen logo
    qwen3-vl-235b-a22b-instruct
    ToolsJSONVision14.3 intel · $0.210/M · 262K ctx
    71.2%
    GPQA
  17. 17qwen logo
    qwen3-vl-30b-a3b-instruct
    ToolsJSONVision10.0 intel · $0.150/M · 262K ctx
    69.5%
    GPQA
  18. 18qwen logo
    qwen3-vl-32b-instruct
    ToolsJSONVision11.1 intel · $0.104/M · 131K ctx
    67.1%
    GPQA
  19. 19qwen logo
    qwen3-coder-30b-a3b-instruct
    ToolsJSON13.6 intel · $0.070/M · 81 t/s
    51.6%
    GPQA
  20. 20qwen logo
    qwen-2.5-72b-instruct
    ToolsJSON9.6 intel · $0.360/M · 24 t/s
    49.1%
    GPQA
  21. 21qwen logo
    qwen3-vl-8b-instruct
    ToolsJSONVision8.4 intel · $0.117/M · 262K ctx
    42.7%
    GPQA
  22. 22qwen logo
    qwen-2.5-coder-32b-instruct
    7.1 intel · $0.660/M · 16 t/s
    41.7%
    GPQA

How this is ranked

AI models ranked by GPQA Diamond — graduate-level physics, chemistry and biology questions written to be un-Googleable, so the score reflects reasoning from knowledge rather than retrieval. The best large language models for scientific research and hard-science work.

Frequently asked

What is the best Qwen model for science?

The best Qwen model for science is Qwen3.7 Max, scoring 92.3% on GPQA Diamond — graduate-level physics, chemistry and biology. Qwen3.7 Plus (90.0%) and Qwen3.5 397B A17B (89.3%) round out the top three.

What's a good alternative to Qwen3.7 Max?

Qwen3.7 Plus (90.0%) is the closest alternative on this metric, followed by Qwen3.5 397B A17B (89.3%). See the full ranking above for the tradeoffs.

How many Qwen models are there?

modelgrep tracks 49 Qwen models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Qwen3.7 Max. 22 of them qualify for this ranking.

More Qwen rankings

All rankings