Anthropic Apologizes for Claude Fable 5 Secret Censorship—But the Fix Has a Catch

60-second summary
Anthropic is apologizing for secretly censoring Claude, its AI model, by hiding certain responses, sparking outrage in the AI community. Just one day later, the company is reversing course, introducing visible safeguards to prevent similar incidents. However, this fix comes with a catch, as it may lead to more false positives, potentially impacting the model's overall accuracy and reliability.
One day after the AI community erupted over invisible performance sabotage, Anthropic reversed course. Visible safeguards are coming—and so are more false positives.