AI Glossary
Trust and Safety
What is Trust and Safety?
The field within AI development focused on identifying, measuring, and mitigating harmful model behaviors including bias, toxicity, misinformation generation, and misuse. Trust and safety teams at AI labs run red-teaming exercises, evaluate outputs for policy violations, and design guardrails. It is closely related to AI alignment and AI ethics.
Example in practice
An AI platform's trust and safety team that detects and removes prompt patterns enabling policy violations — and reports them back to the model team for retraining — is performing the operational side of responsible AI deployment.
Learn more
See Trust and Safety applied in a professional context through this free course.
AI Strategy for Leaders →