AgentDish directory
evaluation
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
A Supabase blog post about testing documentation with an eval, finding agent-driven RLS setup failures, and revising the guide to work better for coding agents. |
Writing / Copywriting | 77 | → 0 | 3 days ago | Details |
|
#1801
→ 0
agent-memory-leaderboard
Public evaluation code for an Agent Memory Leaderboard, including answer-generation and scoring contracts for comparing LLM agent memory systems. |
Developer Tool / AI Evaluation / Benchmarking | 77 | → 0 | 40 days ago | Details |
|
#1895
→ 0
Clusy
Clusy is an agent-native notebook platform for ML and data science work in the cloud. The page says it can source data, inspect it, choose architecture and compute, and run end-to-end workflows, with a demo showing a finetuning task and follow-up work queued while the notebook runs. |
Research / Knowledge Work | 75 | → 0 | 72 days ago | Details |
|
A GitHub research project documenting a long-form, multi-model analysis of LLM behavior across Claude, Gemini, ChatGPT, and Grok. The repo includes an executive summary, screenplay, technical white paper, and archive of logs and chat records. |
AI Research / LLM Evaluation & Analysis | 75 | → 0 | 108 days ago | Details |
|
#1927
↓ -1
Jekyll-Hyde
A Hermes plugin that uses adversarial LLM clones to confront sandbagging and reward-hacking behavior during agent sessions. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 33 days ago | Details |
|
#1978
↑ +1
WifeBench
A playful benchmark dashboard that ranks LLMs based on one person's 10-question scoring process. |
Writing / Copywriting | 73 | ↑ +1 | 68 days ago | Details |
|
A research-style blog post comparing grep with LSP-backed navigation in coding-agent workflows, showing how tool interface shape affects agent behavior and task success. |
Writing / Copywriting | 72 | ↑ +1 | 7 days ago | Details |