modelgrep

Artificial Analysis Agentic Index

A composite benchmark of agent ability — multi-step planning, tool calling and task completion (including Tau²-Bench) — used to rank models for agentic work.

The Agentic Index measures what agent frameworks actually demand: choosing the right tool, chaining many calls without losing the plot, recovering from errors, and finishing multi-step tasks (benchmarks like Tau²-Bench simulate customer-service and operations workflows with real tool APIs).

Agentic ability diverges from raw intelligence more than any other axis — some models that score brilliantly on knowledge benchmarks are unreliable at sustained tool use, and vice versa. If you're building agents, rank by this index rather than the general one.

modelgrep refreshes Agentic Index scores daily and ranks every benchmarked model by it on the agents leaderboard.