Topic Hub
Best Model Benchmarking Open Source Projects
This hub collects the best open source projects in Model Benchmarking and ranks them by both momentum and authority.
Data window: Last 7 days (with 24h tie-breakers for Trending now)
Last updated: Sep 18, 2026
Projects
38
GitHub stars
44.9K
Hugging Face likes
27.5K
Replicate runs
51
Trending now
Fastest-growing projects in Model Benchmarking over the last 7 days, with 24h change as a tie-breaker.
meta-harness-tbench2-artifact
Achieve high scores on terminal-based benchmarks with an advanced AI agent.
Video-MME-v2
Evaluating video understanding models with a comprehensive benchmark.
bullshit-benchmark
Evaluating AI models' ability to detect nonsense and challenge invalid assumptions.
asr-rescore-bench
Benchmark for evaluating LLM-based ASR rescoring strategies to improve speech recognition accuracy.
Hidden gems
Smaller projects with unusually strong momentum. We look for lower total metrics plus positive 7-day growth.
asr-rescore-bench
Benchmark for evaluating LLM-based ASR rescoring strategies to improve speech recognition accuracy.
diagnostic
Identifies AI misalignment risks through 32 tests across fabrication, manipulation, deception, unpredictability, and opacity categories.
AgentHarness
Evaluate AI model performance on deep-research benchmarks with this open-source harness.
speech-tokenizer-arena
Benchmarking tool for discrete speech tokenizers to compare performance and choose the best fit for specific use cases.
WildClawBench
Evaluating AI agents in real-world scenarios to improve their performance and capabilities.
FINAL-Bench/all-bench-leaderboard
Compare AI model scores across multiple modalities in one view.