The best local inclusionAI model is Ling-3.0-flash (37.4 intelligence) — open weights on Hugging Face, small enough to run on your own hardware. Ling-3.0-flash (free) (37.4) is next.
The best local LLMs — open-weight models small enough to run on your own hardware with Ollama, llama.cpp or vLLM — ranked by intelligence. Frontier-scale open models (400B+) are excluded: downloadable isn't the same as runnable. Every model here is on Hugging Face.
The best local inclusionAI model is Ling-3.0-flash (37.4 intelligence) — open weights on Hugging Face, small enough to run on your own hardware. Ling-3.0-flash (free) (37.4) is next.
Ling-3.0-flash (free) (37.4) is the closest alternative on this metric. See the full ranking above for the tradeoffs.
modelgrep tracks 5 inclusionAI models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Ling-3.0-flash. 2 of them qualify for this ranking.