The best LLM for RAG is Claude Fable 5 — 59.9 intelligence with 1M tokens of context for retrieved passages. GPT-5.6 Sol (58.9) and Kimi K3 (57.1) round out the top three.
The best AI models for retrieval-augmented generation, ranked by intelligence among models with at least 128K tokens of context — enough to hold retrieved passages plus conversation. RAG pipelines run at volume, so weigh the price and speed columns as hard as the score.
The best LLM for RAG is Claude Fable 5 — 59.9 intelligence with 1M tokens of context for retrieved passages. GPT-5.6 Sol (58.9) and Kimi K3 (57.1) round out the top three.
The best AI model for RAG is Claude Fable 5 — 59.9 intelligence with 1M tokens of context for retrieved passages. GPT-5.6 Sol (58.9) and Kimi K3 (57.1) round out the top three.
GPT-5.6 Sol (58.9) is the closest alternative on this metric, followed by Kimi K3 (57.1). See the full ranking above for the tradeoffs.