AgentDish directory

ai-evaluation

Accepted listings with this tag.

Listing Category Score Trend Checked
#365 ↓ -4
Wayfinder

Reference implementation for building and understanding AI evaluation systems, including rule-based evaluation, human review, LLM-as-a-Judge, online evaluation, and experiment comparison for a flight search app.

Developer Tools / AI Evaluation 87 ↓ -4 5 days ago Details
#809 ↓ -6
ObserverBench

ObserverBench is an interactive benchmark and guided demo for testing whether internal AI monitoring methods help make safer decisions, not just better predictions. It includes real-model results, benchmark tasks, and a no-login walkthrough showing how warning scores affect review and harm that slips through.

Developer Tools / AI Evaluation 84 ↓ -6 17 hours ago Details

ARF defines a JSON record format for AI evaluation runs with canonicalization and reproducible SHA-256 digests, so scores can be traced back to the underlying question, model, inputs, and judgment.

Developer Tool / AI Evaluation / Data Format 83 ↓ -3 35 days ago Details
#1638 ↑ +2
Lagotto Meter

Lagotto Meter is a beta web tool that checks how an AI agent perceives a website versus what the site claims about itself. It positions itself as a semantic truth/readability analyzer with a reproducible analysis method and public tools.

Developer Tool / AI Evaluation 79 ↑ +2 52 days ago Details
#1770 ↑ +6
LLM INQUISITOR

A GitHub repository that proposes a practical methodology for evaluating how AI systems behave during real work, with quick-start, practitioner, and methodology guides included.

Developer Tools / AI Evaluation 78 ↑ +6 114 days ago Details
#1838 ↑ +129
Agent Eval

A GitHub repo for evaluating agentic AI pipeline systems, with guidance for defining metrics, building eval cases, running repeatable tests, and tracking regressions.

Developer Tools / Copywriting 77 ↑ +129 128 days ago Details