AgentDish directory

benchmarking

Accepted listings with this tag.

Listing Category Score Trend Checked
#1611 ↑ +2
Agent Lightning v1.0.1

A GitHub release for Microsoft’s Agent Lightning Skill, which helps coding agents optimize other AI agents using benchmarks and iterative improvements across prompts, tools, workflows, models, and reasoning settings.

Developer Tools / AI Agents / Agent Training 79 ↑ +2 17 days ago Details
#1654 ↑ +2
HoprLabs

A Python CLI and research toolkit for simulating AI training math before model training. It estimates memory, training time, token budget, config risks, benchmark speed, and reliability, with optional native Rust and C backends.

Developer Tool / AI Research Toolkit 79 ↑ +2 77 days ago Details

An Apple Silicon–optimized inference build of Bonsai 1.7B with custom Metal kernels, benchmark results, quick-start instructions, and a bundled OpenAI-compatible server.

Developer Tools / Code Assistant 79 ↓ -223 128 days ago Details

A detailed technical write-up about training a 3.8B-parameter language model from scratch, including setup, throughput improvements, ablations, benchmarks, and cost/performance results.

AI Research / LLM Training / Write-up 78 ↑ +6 42 hours ago Details

A blog post about verifiable RAG that benchmarks open-source NLI verifiers against Claude on RAGTruth and describes a Python library for sentence-level citation and claim verification.

AI / RAG / Verification & Hallucination Detection 78 ↑ +6 101 days ago Details

A blog post from Augment Code comparing its coding agent, Auggie, against Claude Code on Opus 4.7. It presents benchmark results, token usage, cost comparisons, and an explanation of the Context Engine and Prism router.

Developer Tools / AI Coding Assistant 77 → 0 116 days ago Details

A GitHub research project documenting a long-form, multi-model analysis of LLM behavior across Claude, Gemini, ChatGPT, and Grok. The repo includes an executive summary, screenplay, technical white paper, and archive of logs and chat records.

AI Research / LLM Evaluation & Analysis 75 → 0 108 days ago Details
#1911 ↓ -1
Project HydraFusion

GitHub’s research preview on multi-model orchestration for coding tasks. The page describes adaptive workflow selection and includes benchmark comparisons against the Opus 5 baseline with estimated cost reduction claims.

Developer Tools / Code Assistant 74 ↓ -1 6 days ago Details

A technical blog post from the Cua repository describing a process-scoped Metal capability shim for macOS VMs that improves llama.cpp performance on Apple Silicon, with benchmark results and reproduction details.

Developer Tools / AI Infrastructure 74 ↓ -1 31 days ago Details

A research article describing an empirical runtime-monitoring approach for multi-turn LLM agents, backed by 3,175 runs across four benchmarks. It also points to an open-source implementation, state-harness, with Rust/Python support, LangGraph and CrewAI adapters, CLI tooling, and OpenTelemetry export.

Developer Tools / Code Assistant 74 ↓ -1 39 days ago Details

A blog post describing an evaluation harness for comparing agentic coding tools and prompts across realistic SWE tasks, with token-cost results and model-specific behavior notes.

Developer Tools / Code Assistant 74 ↓ -1 89 days ago Details
#1949 ↓ -1
LEVI

LEVI is a harness-first evolutionary framework for code and prompt optimization. It focuses on reducing LLM cost with diversity-preserving search, role-aware model routing, and a proxy benchmark, and presents comparative results against several existing systems.

Developer Tools / Code Assistant 74 ↓ -1 95 days ago Details

A Superconductor blog post showing how background coding agents were used to reproduce, diagnose, and fix a Rails memory leak using derailed_benchmarks, with a reusable Agent Skill workflow included.

Developer Tools / Code Assistant 74 ↓ -1 106 days ago Details

A Zenodo preprint reporting a 12,160-trial black-box evaluation of GPT-5.4 outputs under two prompt conditions, with detailed token-ceiling sweeps, control/ablation trials, and hash-chained verification.

Research / AI Model Evaluation 72 ↑ +1 37 days ago Details