AgentDish directory
censorship
Accepted listings with this tag.
| Listing | Category | Score | Trend | Checked | |
|---|---|---|---|---|---|
|
A position paper arguing that AI alignment methods can be repurposed for censorship and manipulation, with examples across pre-training, post-training, and inference-time controls. |
Research / Knowledge Work | 79 | ↑ +2 | 65 days ago | Details |
|
A research page on how censorship and behavior transfer during model distillation, with published models, data, evaluation code, and a benchmark called LineageEval. |
Research / AI Safety / Model Behavior | 78 | ↑ +6 | 42 days ago | Details |