https://github.com/URML-MARS/URML/tree/main/examples/physical-ai-safety-eval
Under the hood
AI queue
Watch AgentDish crawl new pages and run AI reviews in public.
Processing queue
Live queue state for crawls and AI evaluations.
#2726 submission_evaluation
succeeded
#2725 submission_evaluation
succeeded
https://github.com/Wholiver/metis
#2724 submission_evaluation
succeeded
https://github.com/ritenv/tokensift
#2723 submission_evaluation
succeeded
https://www.reddit.com/r/AI_Agents/comments/1w1nuzb/whats_your_weirdest_ai_agent_log/
#2722 submission_evaluation
succeeded
https://dmx.deepmodel.ai
#2721 submission_evaluation
succeeded
https://github.com/grith-ai/grith
#2720 submission_evaluation
succeeded
https://github.com/guillaumemeyer/watermarks-remover
#2719 submission_evaluation
succeeded
https://www.theregister.com/research/2026/08/28/researcher-shows-how-claude-code-can-be-tricked-simply-by-asking-it-to-summarize-a-website/5293372
#2718 submission_evaluation
succeeded
https://github.com/ryanssenn/gemma4.c
#2717 submission_evaluation
succeeded
https://github.com/0xplaygrounds/rig
#2716 submission_evaluation
succeeded
https://github.com/IdoGol24/weir
#2715 submission_evaluation
succeeded
https://github.com/nilbuild/rundown
Processing logs
Newest worker events first.
info / queued
Submission evaluation queued.
https://nvidia.github.io/OpenShell-Research/dev-notes/posts/2026-09-10-learning-formal-methods-agent-policy-prover/ / 2026-09-15 15:32:10
info / queued
Submission evaluation queued.
https://browser-use.com/posts/bitter-lesson-browser-agents / 2026-09-15 15:32:09
info / queued
Submission evaluation queued.
https://github.com/stagas/livediff / 2026-09-15 15:32:09
info / queued
Submission evaluation queued.
https://heymeraki.substack.com/p/aie_10-building-it / 2026-09-15 15:32:09
info / evaluating
Running the AI evaluation.
https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit / 2026-09-15 15:32:09
info / crawl_saved
Saved crawl snapshot with status 200.
https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit / 2026-09-15 15:32:09
info / queued
Submission evaluation queued.
https://tokencanopy.com/products/agentdrive / 2026-09-15 15:32:09
info / queued
Submission evaluation queued.
https://github.com/ordewell/ordewell / 2026-09-15 15:32:09
info / queued
Submission evaluation queued.
https://github.com/greentfrapp/panel / 2026-09-15 15:32:09
info / queued
Submission evaluation queued.
https://sunkcost.ai / 2026-09-15 15:32:09
info / crawling
Crawling https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit.
https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit / 2026-09-15 15:32:09
info / checking_duplicate
Checking for an existing accepted listing.
https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit / 2026-09-15 15:32:09
info / starting
Started processing job.
https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit / 2026-09-15 15:32:09
info / queued
Submission evaluation queued.
https://www.amazon.science/blog/why-dont-machine-learning-research-agents-overfit / 2026-09-15 15:32:09
info / complete
Processing job completed.
https://botbin.io/2fu56d / 2026-09-14 15:35:16
info / complete
Rejected submission decision saved.
https://botbin.io/2fu56d / 2026-09-14 15:35:16
info / rejecting
Saving rejected submission decision.
https://botbin.io/2fu56d / 2026-09-14 15:35:16
info / evaluation_saved
Saved rejected evaluation with score 22.
https://botbin.io/2fu56d / 2026-09-14 15:35:16
info / evaluating
Running the AI evaluation.
https://botbin.io/2fu56d / 2026-09-14 15:35:13
info / crawl_saved
Saved crawl snapshot with status 200.
https://botbin.io/2fu56d / 2026-09-14 15:35:13