AgentDish directory

mechanistic interpretability

Accepted listings with this tag.

Listing Category Score Trend Checked
#809 ↓ -6
ObserverBench

ObserverBench is an interactive benchmark and guided demo for testing whether internal AI monitoring methods help make safer decisions, not just better predictions. It includes real-model results, benchmark tasks, and a no-login walkthrough showing how warning scores affect review and harm that slips through.

Developer Tools / AI Evaluation 84 ↓ -6 17 hours ago Details
#1705 ↑ +6
HyperSAE

HyperSAE is a Python package for mechanistic interpretability that trains hyperbolic sparse autoencoders on LLM activations. The page includes installation steps, a quickstart example, benchmark tables, and a short software architecture overview.

Developer Tools / AI/ML Frameworks 78 ↑ +6 30 days ago Details

An ICML 2026 research project page arguing that prompt injection comes from how LLMs misread roles, with an extended writeup, examples, and links to the paper, code, arXiv, and BibTeX.

Research / AI Safety 77 → 0 80 days ago Details