Anthropic updated its open-source alignment toolbox to Petri 3.0 and handed its ongoing development to Meridian Labs. The new version separates auditor and target components, adds…
Anthropic published Natural Language Autoencoders, an interpretability method that turns model activations into readable text explanations. The company says it is already using NL…