AgentDish directory

evaluation

Accepted listings with this tag.

Listing Category Score Trend Checked

A Supabase blog post about testing documentation with an eval, finding agent-driven RLS setup failures, and revising the guide to work better for coding agents.

Writing / Copywriting 77 → 0 3 days ago Details

Public evaluation code for an Agent Memory Leaderboard, including answer-generation and scoring contracts for comparing LLM agent memory systems.

Developer Tool / AI Evaluation / Benchmarking 77 → 0 40 days ago Details
#1895 → 0
Clusy

Clusy is an agent-native notebook platform for ML and data science work in the cloud. The page says it can source data, inspect it, choose architecture and compute, and run end-to-end workflows, with a demo showing a finetuning task and follow-up work queued while the notebook runs.

Research / Knowledge Work 75 → 0 72 days ago Details

A GitHub research project documenting a long-form, multi-model analysis of LLM behavior across Claude, Gemini, ChatGPT, and Grok. The repo includes an executive summary, screenplay, technical white paper, and archive of logs and chat records.

AI Research / LLM Evaluation & Analysis 75 → 0 108 days ago Details
#1927 ↓ -1
Jekyll-Hyde

A Hermes plugin that uses adversarial LLM clones to confront sandbagging and reward-hacking behavior during agent sessions.

Developer Tools / Code Assistant 74 ↓ -1 33 days ago Details
#1978 ↑ +1
WifeBench

A playful benchmark dashboard that ranks LLMs based on one person's 10-question scoring process.

Writing / Copywriting 73 ↑ +1 68 days ago Details

A research-style blog post comparing grep with LSP-backed navigation in coding-agent workflows, showing how tool interface shape affects agent behavior and task success.

Writing / Copywriting 72 ↑ +1 7 days ago Details