AgentDish directory
benchmark
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#19
↓ -3
gemma4.c
A pure C runtime for Gemma 4 E2B CPU inference, with export, benchmarking, and numerical validation tools included. |
Developer Tools / Machine Learning / Inference Runtime | 91 | ↓ -3 | 13 days ago | Details |
|
#138
→ 0
ReactBench
ReactBench is a benchmark for evaluating coding agents on realistic React work, with published scores, cost comparisons, and example tasks focused on production-grade frontend issues. |
Developer Tools / Code Assistant | 89 | → 0 | 57 days ago | Details |
|
#206
↓ -3
Robot Football League (RFL)
An open robot football league where frontier AI models and open-source club code manage simulated Unitree G1 humanoids in daily matches, with live streams, match logs, and a public engine for teams to join. |
Developer Tools / Code Assistant | 88 | ↓ -3 | 16 days ago | Details |
|
#265
↓ -3
Genesys
Open-source causal-graph memory for AI agents with MCP support, scoring, active forgetting, and benchmark results published in the repo. |
AI Development / Memory / Agent Infrastructure | 88 | ↓ -3 | 51 days ago | Details |
|
#340
↓ -3
CAD-Bench
An open benchmark and leaderboard for AI CAD agents, with 308 prompts across 20 categories and layered scoring for geometry, engineering, manufacturability, and cognition. |
Research / Knowledge Work | 88 | ↓ -3 | 125 days ago | Details |
|
#495
↑ +2
Modded-NanoGPT
Open-source repository for a NanoGPT training speedrun, focused on getting a 124M model to target loss on 8x H100 GPUs in under 90 seconds. The README includes run commands, Docker instructions, and a detailed list of model and systems optimizations. |
Developer Tools / Machine Learning / Training Optimization | 86 | ↑ +2 | 19 days ago | Details |
|
#583
↑ +2
The Banana Test
A visual benchmark that asks AI coding agents to generate a single-file Three.js animation of a banana plant’s full life cycle, then compares the live results side by side. |
AI Tools / Benchmarking | 86 | ↑ +2 | 66 days ago | Details |
|
#809
↓ -6
ObserverBench
ObserverBench is an interactive benchmark and guided demo for testing whether internal AI monitoring methods help make safer decisions, not just better predictions. It includes real-model results, benchmark tasks, and a no-login walkthrough showing how warning scores affect review and harm that slips through. |
Developer Tools / AI Evaluation | 84 | ↓ -6 | 17 hours ago | Details |
|
#834
↓ -6
Agentic Determinism Index
Open-source harness for measuring how consistently hosted LLM APIs return the same output across repeated runs and over time, with raw transcripts, scoring, and a static leaderboard generator. |
Developer Tools / Code Assistant | 84 | ↓ -6 | 10 days ago | Details |
|
#836
↓ -6
Preseason
Open-source benchmark that measures which developer tools LLMs recommend when asked to build real web apps. It runs fixed prompts against a panel of models, parses the tool recommendations, and publishes rankings and head-to-head comparisons. |
Developer Tools / AI Benchmarking | 84 | ↓ -6 | 11 days ago | Details |
|
#869
↓ -6
Agent Memory Leaderboard
Public benchmark for comparing agent memory systems, with separate textual and coding tracks, submission flow, documentation, and API guidance. |
Developer Tools / Benchmark | 84 | ↓ -6 | 29 days ago | Details |
|
#872
↓ -6
LLM-Tests
An open-source benchmark for testing whether LLMs can follow long arithmetic derivations without using tools. It compares models on recall-proof inputs, logs digit accuracy, and reports results across multiple endpoints and model families. |
Developer Tools / Code Assistant | 84 | ↓ -6 | 30 days ago | Details |
|
#1046
↓ -6
DeepSWE
DeepSWE is a benchmark for measuring frontier coding agents on original, long-horizon software engineering tasks. The page shows a leaderboard, methodology overview, task examples, and a full blog explaining the benchmark design and results. |
Developer Tools / AI Benchmarking | 84 | ↓ -6 | 106 days ago | Details |
|
A GitHub repository that publishes a week of controlled runs comparing seven Google models on the same software task, with transcripts, telemetry, screenshots, and scored benchmark artifacts. |
AI Development / Benchmarking | 83 | ↓ -3 | 18 days ago | Details |
|
A black-box benchmark report on how AI-generated tests detect functional bugs in live APIs across 20 scenarios and 7 systems. |
Developer Tools / Code Assistant | 83 | ↓ -3 | 99 days ago | Details |
|
#1493
↑ +2
Vestige – Silent Rotation benchmark
A GitHub repo section documenting a benchmark for Vestige, a local-first Rust MCP memory layer for multi-agent coding fleets. The page explains the benchmark setup, what is being measured, how to reproduce the main claim, and where to find transcripts, evidence, and harness code. |
Developer Tools / AI / ML Infrastructure | 81 | ↑ +2 | 52 days ago | Details |
|
A workbench report comparing MiniMax M3 and GLM 5.2 on autonomous coding tasks, with scored results, latency and cost data, task-type breakdowns, and examples of where each model performed better. |
Developer Tools / Code Assistant | 81 | ↑ +2 | 83 days ago | Details |
|
#1561
↓ -1
AI World Bakeoff
A showcase of 27 explorable Three.js worlds built autonomously by 9 coding models from the same brief. |
Design / Creative Tools | 80 | ↓ -1 | 52 days ago | Details |
|
#1692
↑ +6
44% on ARC-AGI-1 in 67 cents
A research blog post describing an open-source small transformer trained from scratch to reach 44% on ARC-AGI-1 at very low cost, with architecture notes, ablations, and pointers to code. |
Research / ML Experiment / Benchmark Result | 78 | ↑ +6 | 9 days ago | Details |
|
A Topos case study showing how a guided refactor of a synthetic healthcare claims engine reduced token usage, wall time, and estimated cost for later Gemini feature sessions. |
Developer Tools / Code Assistant | 78 | ↑ +6 | 77 days ago | Details |
|
#1801
→ 0
agent-memory-leaderboard
Public evaluation code for an Agent Memory Leaderboard, including answer-generation and scoring contracts for comparing LLM agent memory systems. |
Developer Tool / AI Evaluation / Benchmarking | 77 | → 0 | 40 days ago | Details |
|
An in-depth Medium post from IronBee about how a verification loop affects AI coding agents, using Web-Bench and comparing DeepSeek with Claude Opus on a real web app task. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 66 days ago | Details |
|
#1978
↑ +1
WifeBench
A playful benchmark dashboard that ranks LLMs based on one person's 10-question scoring process. |
Writing / Copywriting | 73 | ↑ +1 | 68 days ago | Details |