AI Glossary
Adversarial Attack
What is Adversarial Attack?
A technique that crafts inputs designed to fool an AI model into making incorrect predictions or producing harmful outputs. In text-based systems, adversarial attacks include jailbreaks, prompt injections, and inputs designed to trigger misbehavior. Understanding adversarial vulnerabilities is essential for anyone deploying AI in security-sensitive contexts.
Example in practice
A security researcher who adds invisible unicode characters to a malware sample's description to cause an AI classifier to label it as benign has mounted an adversarial attack against the model.
Learn more
See Adversarial Attack applied in a professional context through this free course.
AI Strategy for Leaders →