AgentDish directory
benchmarking
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
#1611
↑ +2
Agent Lightning v1.0.1
A GitHub release for Microsoft’s Agent Lightning Skill, which helps coding agents optimize other AI agents using benchmarks and iterative improvements across prompts, tools, workflows, models, and reasoning settings. |
Developer Tools / AI Agents / Agent Training | 79 | ↑ +2 | 17 days ago | Details |
|
#1654
↑ +2
HoprLabs
A Python CLI and research toolkit for simulating AI training math before model training. It estimates memory, training time, token budget, config risks, benchmark speed, and reliability, with optional native Rust and C backends. |
Developer Tool / AI Research Toolkit | 79 | ↑ +2 | 77 days ago | Details |
|
#1681
↓ -223
Bonsai 1.7B: Apple Silicon Optimized Build
An Apple Silicon–optimized inference build of Bonsai 1.7B with custom Metal kernels, benchmark results, quick-start instructions, and a bundled OpenAI-compatible server. |
Developer Tools / Code Assistant | 79 | ↓ -223 | 128 days ago | Details |
|
#1683
↑ +6
Training a 3.8B LLM to 0.384 CORE for $998
A detailed technical write-up about training a 3.8B-parameter language model from scratch, including setup, throughput improvements, ablations, benchmarks, and cost/performance results. |
AI Research / LLM Training / Write-up | 78 | ↑ +6 | 42 hours ago | Details |
|
#1756
↑ +6
Verified RAG: every sentence checked
A blog post about verifiable RAG that benchmarks open-source NLI verifiers against Claude on RAGTruth and describes a Python library for sentence-level citation and claim verification. |
AI / RAG / Verification & Hallucination Detection | 78 | ↑ +6 | 101 days ago | Details |
|
A blog post from Augment Code comparing its coding agent, Auggie, against Claude Code on Opus 4.7. It presents benchmark results, token usage, cost comparisons, and an explanation of the Context Engine and Prism router. |
Developer Tools / AI Coding Assistant | 77 | → 0 | 116 days ago | Details |
|
A GitHub research project documenting a long-form, multi-model analysis of LLM behavior across Claude, Gemini, ChatGPT, and Grok. The repo includes an executive summary, screenplay, technical white paper, and archive of logs and chat records. |
AI Research / LLM Evaluation & Analysis | 75 | → 0 | 108 days ago | Details |
|
#1911
↓ -1
Project HydraFusion
GitHub’s research preview on multi-model orchestration for coding tasks. The page describes adaptive workflow selection and includes benchmark comparisons against the Opus 5 baseline with estimated cost reduction claims. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 6 days ago | Details |
|
A technical blog post from the Cua repository describing a process-scoped Metal capability shim for macOS VMs that improves llama.cpp performance on Apple Silicon, with benchmark results and reproduction details. |
Developer Tools / AI Infrastructure | 74 | ↓ -1 | 31 days ago | Details |
|
A research article describing an empirical runtime-monitoring approach for multi-turn LLM agents, backed by 3,175 runs across four benchmarks. It also points to an open-source implementation, state-harness, with Rust/Python support, LangGraph and CrewAI adapters, CLI tooling, and OpenTelemetry export. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 39 days ago | Details |
|
#1946
↓ -1
Evaluate Your Agentic Tooling
A blog post describing an evaluation harness for comparing agentic coding tools and prompts across realistic SWE tasks, with token-cost results and model-specific behavior notes. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 89 days ago | Details |
|
#1949
↓ -1
LEVI
LEVI is a harness-first evolutionary framework for code and prompt optimization. It focuses on reducing LLM cost with diversity-preserving search, role-aware model routing, and a proxy benchmark, and presents comparative results against several existing systems. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 95 days ago | Details |
|
A Superconductor blog post showing how background coding agents were used to reproduce, diagnose, and fix a Rails memory leak using derailed_benchmarks, with a reusable Agent Skill workflow included. |
Developer Tools / Code Assistant | 74 | ↓ -1 | 106 days ago | Details |
|
A Zenodo preprint reporting a 12,160-trial black-box evaluation of GPT-5.4 outputs under two prompt conditions, with detailed token-ceiling sweeps, control/ablation trials, and hash-chained verification. |
Research / AI Model Evaluation | 72 | ↑ +1 | 37 days ago | Details |