AgentDish directory

agent-evaluation

Accepted listings with this tag.

Listing Category Score Trend Checked
#391 ↓ -4
Mirrors

Mirrors is a staging and replay environment for AI agents. It rebuilds the systems an agent calls, replays real sessions against prompts, tools, and model changes, and surfaces regressions before deployment.

Developer Tools / AI Development Tools 87 ↓ -4 43 days ago Details
#435 ↓ -4
commensa-audit

A local-first Python tool that audits Git history to quantify rework in AI-generated engineering work, including PR correction rates, churn clusters, superseded work, and line survival.

Developer Tools / AI Analytics 87 ↓ -4 90 days ago Details
#471 ↑ +2
Egma

Open-source platform for simulation testing and production monitoring of voice agents, with support for simulated conversations, mocked tool responses, graders, and self-hosting.

Developer Tools / Testing & QA 86 ↑ +2 17 hours ago Details
#564 ↑ +2
AI Arcade

An interactive site showcasing short AI-built arcade games and a model bench that compares coding models across the same tiny game constraints.

Developer Tools / Code Assistant 86 ↑ +2 60 days ago Details
#949 ↓ -6
ClaySeal Arena

A prompt-injection capture-the-flag arena for testing how AI agents behave under attack. Users can play challenges, create guarded agents, add optional guards and sandboxed tools, and track scores on a leaderboard.

Developer Tools / Code Assistant 84 ↓ -6 57 days ago Details

A research post from Prime Intellect comparing 153 autonomous runs across 18 frontier models on a nanoGPT optimizer speedrun. It presents the setup, harness, results, and discussion around autonomous AI research performance.

Research / AI Research Evaluation 82 ↓ -2 26 days ago Details

A DeepSeek Harness plugin that uses an LLM to verify and rank candidate outputs with select, compare, track, and rollout tools.

Developer Tools / AI DevTools 78 ↑ +6 22 days ago Details

A blog post describing an evaluation harness for comparing agentic coding tools and prompts across realistic SWE tasks, with token-cost results and model-specific behavior notes.

Developer Tools / Code Assistant 74 ↓ -1 89 days ago Details