Developer Tools / Code Assistant

BoundaryBench

BoundaryBench is an open-source benchmark for coding agents running under hardened sandbox policies. It compares harnesses like Claude Code, Codex, Terminus 2, and Grok on Terminal-Bench tasks and includes quickstart, policy controls, and result export tools.

Clear28/30
Useful29/30
Specific18/20
Complete15/20
BoundaryBench screenshot

Why it was accepted

The page clearly describes a real AI-adjacent developer tool with a specific purpose: benchmarking coding-agent harnesses under sandbox policy constraints. It includes enough visible detail for a useful public listing, including supported harnesses, setup steps, policy model, output format, and live leaderboard context.

Weakness

The crawl does not show the full repository docs or current maintenance signals beyond one commit and a small star count. It also leaves some practical questions open, such as exact task coverage details, how easy it is to reproduce the leaderboard locally, and what the official website/report contain beyond the README references.

Review status

36 days ago #49 ↓ -2

Last evaluated 36 days ago. Current rank #49. Down 2 spots in the rankings.

Score history

90

Related listings

CodeGraph screenshot
94

Developer Tools / AI for Code

CodeGraph is a local code knowledge graph for AI coding agents like Claude Code, Cursor, Codex, OpenCode, and Hermes Agent. It aims to cut token use, tool calls, and runtime by letting agents query pre-indexed code structure instead of scanning files repeatedly.

Mercemur screenshot
93

Developer Tools / API / AI Agent Infrastructure

Mercemur is a store automation platform with three interfaces: a REST API, an MCP server for AI agents, and a CLI for editing storefronts as files. The docs show authentication, scopes, rate limits, error handling, pagination, idempotency, and a large API surface for catalogue, orders, customers, pricing, and money.

ripwire screenshot
#5 ripwire
92

Developer Tools / Code Assistant

A zero-dependency C++23 CLI and MCP server that gives coding agents a ranked map of a repository, including call context, blast radius, and tests to run.

Traccia screenshot
#6 Traccia
92

Developer Tools / Code Assistant

Traccia is an AI agent observability and governance platform with OpenTelemetry-native tracing, runtime policy enforcement, prompt registry, evals, cost attribution, and compliance evidence export.