Topic Hub

Best Model Benchmarking Open Source Projects

This hub collects the best open source projects in Model Benchmarking and ranks them by both momentum and authority.

Data window: Last 7 days (with 24h tie-breakers for Trending now)

Last updated: Sep 18, 2026

Projects

38

GitHub stars

44.9K

Hugging Face likes

27.5K

Replicate runs

51

Authority projects

Projects with the strongest long-term signal in Model Benchmarking, ranked by total stars, likes, or runs on their primary platform.

llmfit

Find the best LLM models for your hardware with one command.

GitHub
28.7Kstars

open-llm-leaderboard/open_llm_leaderboard

Evaluate and compare Large Language Models with an interactive leaderboard.

Hugging Face Space
14.1Klikes

mteb/leaderboard

Compare and evaluate text and image embedding models across various tasks and languages.

Hugging Face Space
7.7Klikes

DeepSpec

GitHub

GitHub
6.6Kstars

lmarena-ai/lmarena-leaderboard

Compare and evaluate AI models with a crowdsourced benchmarking platform.

Hugging Face Space
4.7Klikes

bullshit-benchmark

Evaluating AI models' ability to detect nonsense and challenge invalid assumptions.

GitHub
1.1Kstars

RealReplicaBench

GitHub

GitHub
1Kstars

RealReplicaBench

GitHub

GitHub
1Kstars

agent-memory-leaderboard/leaderboard

Hugging Face Space

Hugging Face Space
767likes

meta-harness-tbench2-artifact

Achieve high scores on terminal-based benchmarks with an advanced AI agent.

GitHub
694stars

chainreason

Evaluating language models on Ethereum and DeFi tasks.

GitHub
623stars

ProgramBench

Rebuilding programs from scratch using language models.

GitHub
574stars

Hidden gems

Smaller projects with unusually strong momentum. We look for lower total metrics plus positive 7-day growth.

nordef-matrix-open-control

GitHub

GitHub
247stars
2477d

augustus

Scan large language models for vulnerabilities and adversarial attacks.

GitHub
196stars
1967d

tts-prosody-probe

Analyze speech synthesis naturalness beyond spectral fidelity.

GitHub
218stars
2187d

asr-rescore-bench

Benchmark for evaluating LLM-based ASR rescoring strategies to improve speech recognition accuracy.

GitHub
205stars
2027d

codex-candy-eval

GitHub

GitHub
190stars
1757d

diagnostic

Identifies AI misalignment risks through 32 tests across fabrication, manipulation, deception, unpredictability, and opacity categories.

GitHub
142stars
1427d

AgentHarness

Evaluate AI model performance on deep-research benchmarks with this open-source harness.

GitHub
133stars
927d

speech-tokenizer-arena

Benchmarking tool for discrete speech tokenizers to compare performance and choose the best fit for specific use cases.

GitHub
168stars
1687d

WildClawBench

Evaluating AI agents in real-world scenarios to improve their performance and capabilities.

GitHub
183stars
1477d

FINAL-Bench/all-bench-leaderboard

Compare AI model scores across multiple modalities in one view.

Hugging Face Space
51likes
307d

joelniklaus/harness-optimization

Hugging Face Space

Hugging Face Space
52likes
237d

gamecraft-bench

GitHub

GitHub
135stars
717d