modelgrep

Best Local inclusionAI Models

Match · Updated August 2026

The best local inclusionAI model is Ling-3.0-flash (37.4 intelligence) — open weights on Hugging Face, small enough to run on your own hardware. Ling-3.0-flash (free) (37.4) is next.

37.4Intelligence
168 t/sSpeed
$0.075Input /M
131KContext
  1. 1I
    ling-3.0-flash
    ReasoningTools37.4 intel · $0.075/M · 168 t/s
    37.4
    Intelligence
  2. 2I
    ling-3.0-flash:free
    ReasoningTools37.4 intel · Free/M · 98 t/s
    37.4
    Intelligence

How this is ranked

The best local LLMs — open-weight models small enough to run on your own hardware with Ollama, llama.cpp or vLLM — ranked by intelligence. Frontier-scale open models (400B+) are excluded: downloadable isn't the same as runnable. Every model here is on Hugging Face.

Frequently asked

What is the best local inclusionAI model?

The best local inclusionAI model is Ling-3.0-flash (37.4 intelligence) — open weights on Hugging Face, small enough to run on your own hardware. Ling-3.0-flash (free) (37.4) is next.

What's a good alternative to Ling-3.0-flash?

Ling-3.0-flash (free) (37.4) is the closest alternative on this metric. See the full ranking above for the tradeoffs.

How many inclusionAI models are there?

modelgrep tracks 5 inclusionAI models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Ling-3.0-flash. 2 of them qualify for this ranking.

More inclusionAI rankings

All rankings