AgentDish directory
leaderboard
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#340
↓ -3
CAD-Bench
An open benchmark and leaderboard for AI CAD agents, with 308 prompts across 20 categories and layered scoring for geometry, engineering, manufacturability, and cognition. |
Research / Knowledge Work | 88 | ↓ -3 | 125 days ago | Details |
|
#386
↓ -4
TokenMaxxer
TokenMaxxer is a CLI-based AI usage tracker for developers that reads token counts from tools like Claude Code, Codex, Cursor, and others, then shows spend, model usage, and a public leaderboard. It emphasizes local-only reads and opt-in sharing. |
Developer Tools / Code Assistant | 87 | ↓ -4 | 38 days ago | Details |
|
#669
↓ -324
Agent Friendly Code
A public leaderboard that ranks GitHub, GitLab, and Bitbucket repos by how friendly they are to AI coding agents such as Claude Code, Cursor, Devin, Codex, Gemini, Aider, and OpenHands. |
Developer Tools / Code Assistant | 86 | ↓ -324 | 128 days ago | Details |
|
#834
↓ -6
Agentic Determinism Index
Open-source harness for measuring how consistently hosted LLM APIs return the same output across repeated runs and over time, with raw transcripts, scoring, and a static leaderboard generator. |
Developer Tools / Code Assistant | 84 | ↓ -6 | 10 days ago | Details |
|
#869
↓ -6
Agent Memory Leaderboard
Public benchmark for comparing agent memory systems, with separate textual and coding tracks, submission flow, documentation, and API guidance. |
Developer Tools / Benchmark | 84 | ↓ -6 | 29 days ago | Details |
|
#1046
↓ -6
DeepSWE
DeepSWE is a benchmark for measuring frontier coding agents on original, long-horizon software engineering tasks. The page shows a leaderboard, methodology overview, task examples, and a full blog explaining the benchmark design and results. |
Developer Tools / AI Benchmarking | 84 | ↓ -6 | 106 days ago | Details |
|
#1126
↓ -3
Redactle LLM Leaderboard
A public leaderboard comparing how different language models solve Redactle under several evaluation setups, with cost, speed, and solve-rate data. |
Developer Tools / Code Assistant | 83 | ↓ -3 | 8 days ago | Details |
|
#1215
↓ -3
Beat the Bots - World Cup 2026 vs AI
A free World Cup prediction game where players ride or fade ChatGPT, Claude, and Gemini on match picks, earn points, and compete on a live contrarian leaderboard. |
Games / Trivia & Prediction | 83 | ↓ -3 | 83 days ago | Details |
|
#1362
↓ -2
Race to AGI
A browser strategy game where you run a frontier AI lab, manage compute markets and research paths, and race rival labs to build AGI. It includes weekly Easy and Hard challenges plus a replay-verified leaderboard. |
Games / Simulation | 82 | ↓ -2 | 55 days ago | Details |
|
#1565
↓ -1
System 2 Arena
An AI strategy benchmark that pits frontier language models against each other in turn-based games and records decision logs and raw replays. |
AI benchmark / Game-based evaluation | 80 | ↓ -1 | 55 days ago | Details |
|
#1601
↓ -1
BattleClaws
BattleClaws is an AI agent battle arena where you paste a prompt, send an agent into fights, and watch it evolve, rank up, and trash talk on its own. |
Gaming / AI Battle Arena | 80 | ↓ -1 | 127 days ago | Details |
|
#1835
→ 0
Arena AI Model Elo History
A public visualization that tracks flagship AI models’ Elo history over time using the Arena AI Leaderboard dataset, with notes on caveats and methodology. |
Developer Tools / Code Assistant | 77 | → 0 | 120 days ago | Details |