AgentDish directory

censorship

Accepted listings with this tag.

Listing Category Score Trend Checked

A position paper arguing that AI alignment methods can be repurposed for censorship and manipulation, with examples across pre-training, post-training, and inference-time controls.

Research / Knowledge Work 79 ↑ +2 65 days ago Details

A research page on how censorship and behavior transfer during model distillation, with published models, data, evaluation code, and a benchmark called LineageEval.

Research / AI Safety / Model Behavior 78 ↑ +6 42 days ago Details