AgentDish directory

Research

Accepted listings with this tag.

Listing Category Score Trend Checked

A research post from Prime Intellect comparing 153 autonomous runs across 18 frontier models on a nanoGPT optimizer speedrun. It presents the setup, harness, results, and discussion around autonomous AI research performance.

Research / AI Research Evaluation 82 ↓ -2 26 days ago Details
#1415 ↓ -2
BigTech AI News

Chrome extension that tracks major AI companies, pulls in AI news and research, and generates daily summaries with Gemini, including language-aware summaries and article deep dives.

Research / Knowledge Work 82 ↓ -2 93 days ago Details
#1421 ↓ -2
Wanderwhim

AI-native workspace for writers, bloggers, and lifelong learners. It combines source collection, an idea map, AI-assisted exploration, and a focused writing editor designed to support long-form thinking.

Writing / Copywriting 82 ↓ -2 100 days ago Details
#1455 ↑ +228
ShadowBrokers

AI-powered trade signal product for retail traders that turns financial news into ranked trade plans with entries, stops, targets, and tracked accuracy.

Research / Knowledge Work 82 ↑ +228 128 days ago Details

A research page comparing 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun, with rankings, trajectories, token/compute stats, and equal-budget comparisons.

Research / Knowledge Work 81 ↑ +2 19 days ago Details
#1512 ↑ +2
EuroMesh

A sourced model and short report exploring whether Europe could train a sovereign frontier AI model using public compute it already owns, with reproducible code, datasets, and a PDF report.

AI Research / Analysis / Reports 81 ↑ +2 88 days ago Details
#1513 ↑ +2
Foyer

Foyer is a local dashboard for watching AI coding agents work, with a narrated current-focus view and a research panel for reading sourced briefings while you wait.

Developer Tools / AI Developer Tools 81 ↑ +2 92 days ago Details

A research article from Plicara Labs analyzing 1.9 million GitHub agent skills and how often they include code, with breakdowns by language, mention-to-code gaps, and writing language.

Research / Data Analysis 80 ↓ -1 12 days ago Details
#1539 ↓ -1
Canvas Chat

Canvas Chat is a tree-based LLM chat interface for branching, comparing, and organizing conversations on an infinite canvas using your own API keys.

Research / Knowledge Work 80 ↓ -1 15 days ago Details
#1560 ↓ -1
Sharper AI

Sharper AI is an AI office assistant that works from your documents and connected apps to research, draft, summarize, schedule, and produce cited deliverables.

AI Productivity / Office Assistant 80 ↓ -1 52 days ago Details
#1567 ↓ -1
Lucid

Lucid is a browser-based tool for inspecting what a language model is representing before it answers, using Anthropic’s Jacobian lens. It supports small open models, lets users type prompts, view concepts across layers, and export shareable session slices.

AI Tool / Model interpretability 80 ↓ -1 63 days ago Details
#1573 ↓ -1
HEADING OS

A Claude Code-based operating system for running executive workflows across research, communications, CRM, content, and operations, with a split between a public engine repo and a private data repo.

AI Developer Tool / Agent workflow / private data orchestration 80 ↓ -1 72 days ago Details

A position paper arguing that AI alignment methods can be repurposed for censorship and manipulation, with examples across pre-training, post-training, and inference-time controls.

Research / Knowledge Work 79 ↑ +2 65 days ago Details
#1686 ↑ +6
Ramanujan

Ramanujan is a terminal-based multi-model agentic workbench for research in computational mathematics. It lets users chat normally, break down problems, survey literature, run parallel subagents, and produce a consolidated report with deterministic checking via SymPy and Z3.

Developer Tools / AI Agents 78 ↑ +6 5 days ago Details

A DeepMind technical report PDF about double-blind AI evaluations and confidentiality issues in AI safety research.

Research / AI Safety 78 ↑ +6 9 days ago Details

A research blog post describing an open-source small transformer trained from scratch to reach 44% on ARC-AGI-1 at very low cost, with architecture notes, ablations, and pointers to code.

Research / ML Experiment / Benchmark Result 78 ↑ +6 9 days ago Details
#1705 ↑ +6
HyperSAE

HyperSAE is a Python package for mechanistic interpretability that trains hyperbolic sparse autoencoders on LLM activations. The page includes installation steps, a quickstart example, benchmark tables, and a short software architecture overview.

Developer Tools / AI/ML Frameworks 78 ↑ +6 30 days ago Details

A research page on how censorship and behavior transfer during model distillation, with published models, data, evaluation code, and a benchmark called LineageEval.

Research / AI Safety / Model Behavior 78 ↑ +6 42 days ago Details
#1718 ↑ +6
BixRouter

A web app for branching AI conversations into a DAG-style chat tree, with model switching, response style controls, chat import/export, and example conversations.

Research / Knowledge Work 78 ↑ +6 45 days ago Details

Agora-1 is a multi-agent world model from Odyssey that simulates shared real-time environments for up to four participants, human or AI, with a focus on gaming, robotics, reinforcement learning, and foundation model research.

AI Research / World Models 78 ↑ +6 115 days ago Details

An ICML 2026 research project page arguing that prompt injection comes from how LLMs misread roles, with an extended writeup, examples, and links to the paper, code, arXiv, and BibTeX.

Research / AI Safety 77 → 0 80 days ago Details

PaperProfit explains an AI-assisted stock evaluation approach that combines fundamentals, technical signals, and qualitative analysis from transcripts and SEC filings into a weighted score.

Research / Knowledge Work 77 → 0 101 days ago Details

A research write-up on detecting AI agents through process differences in CAPTCHA and related cognitive tasks. It outlines the CogCAPTCHA30 approach, reports human-vs-model differences, and connects the findings to Roundtable’s Proof of Human product.

Research / Knowledge Work 77 → 0 104 days ago Details

arXiv paper on a self-speculative decoding framework for speeding up reasoning LLM inference on edge hardware, with hardware co-design and reported speedups.

Research / AI/ML Paper 77 → 0 105 days ago Details