Developer Tools / Testing

LLM-test-kit

An open-source CLI for testing LLM prompts across consistency, latency, cost, and behavior, with HTML reports and support for OpenAI and Anthropic.

Clear24/30
Useful27/30
Specific17/20
Complete20/20
LLM-test-kit screenshot

Why it was accepted

The page clearly presents a real AI developer tool with a specific use case: testing prompt behavior and reliability across multiple runs. It shows concrete commands, supported providers/models, installation steps, budget controls, and a report workflow, which is enough for a useful public listing.

Weakness

The snapshot does not show package publishing details, release history, or maintenance activity beyond the current repo view, so visitors cannot judge adoption or update cadence from this page alone.

Review status

128 days ago #347 ↓ -263

Last evaluated 128 days ago. Current rank #347. Down 263 spots in the rankings.

Score history

9088

Related listings

CodeGraph screenshot
94

Developer Tools / AI for Code

CodeGraph is a local code knowledge graph for AI coding agents like Claude Code, Cursor, Codex, OpenCode, and Hermes Agent. It aims to cut token use, tool calls, and runtime by letting agents query pre-indexed code structure instead of scanning files repeatedly.

Mercemur screenshot
93

Developer Tools / API / AI Agent Infrastructure

Mercemur is a store automation platform with three interfaces: a REST API, an MCP server for AI agents, and a CLI for editing storefronts as files. The docs show authentication, scopes, rate limits, error handling, pagination, idempotency, and a large API surface for catalogue, orders, customers, pricing, and money.

ripwire screenshot
#5 ripwire
92

Developer Tools / Code Assistant

A zero-dependency C++23 CLI and MCP server that gives coding agents a ranked map of a repository, including call context, blast radius, and tests to run.

Traccia screenshot
#6 Traccia
92

Developer Tools / Code Assistant

Traccia is an AI agent observability and governance platform with OpenTelemetry-native tracing, runtime policy enforcement, prompt registry, evals, cost attribution, and compliance evidence export.