Tag

Adversarial Inputs

All articles tagged with #adversarial inputs

Anthropic's AI Constitution: A Radical Plan for Safe and Ethical AI.
ai-ethics3 years ago

Anthropic's AI Constitution: A Radical Plan for Safe and Ethical AI.

Anthropic has explained how its generative AI, Claude, is protected against adversarial inputs through its Constitutional AI system, which is guided by a set of 10 secret principles of fairness. The system replaces the human in the loop with another AI, which guides the model to take on normative behavior, such as avoiding toxic or discriminatory outputs and creating an AI system that is helpful, honest, and harmless. The principles are synthesized from a range of sources, including the UN Declaration of Human Rights and trust and safety best practices.