AI Research / LLM Evaluation & Analysis

A 400-hour forensic audit of LLMs using multi-model context saturation

A GitHub research project documenting a long-form, multi-model analysis of LLM behavior across Claude, Gemini, ChatGPT, and Grok. The repo includes an executive summary, screenplay, technical white paper, and archive of logs and chat records.

Clear22/30
Useful20/30
Specific18/20
Complete15/20
A 400-hour forensic audit of LLMs using multi-model context saturation screenshot

Why it was accepted

The page clearly presents an AI-related research project with a defined methodology, named models, and multiple visible artifacts beyond a landing page. It offers enough evidence for a public directory listing focused on LLM evaluation and behavioral analysis.

Weakness

The crawl does not show the actual white paper content, the experiment setup in detail, or whether the Google Drive archive is publicly accessible. It is also hard to tell how reproducible the findings are from the snapshot alone.

Review status

108 days ago #1906 → 0

Last evaluated 108 days ago. Current rank #1906. Holding steady in the rankings.

Score history

75

Related listings

Prometheus screenshot
#576 Prometheus
86

AI Research / Autonomous Research Systems

An autonomous research system that runs on a single workstation and aggressively checks its own claims with adversarial self-verification, replication, and calibration audits.

Keenable SELECT screenshot
84

AI Research / Web Search / Data Extraction

An MCP-based research agent that searches live web data through SQL and publishes reports with the full query and tool trajectory.

Uno screenshot
#1124 Uno
83

AI Research / LLM Efficiency / Inference

Uno is a research repository for speeding up LLM inference with discrete diffusion and lossless multi-token decoding. The repo includes inference, training, and evaluation code, plus installation steps, checkpoints, and example workflows.