Skip to main content
Deliberate AcademyProfessional AI Education
AI Glossary

Adversarial Attack

What is Adversarial Attack?

A technique that crafts inputs designed to fool an AI model into making incorrect predictions or producing harmful outputs. In text-based systems, adversarial attacks include jailbreaks, prompt injections, and inputs designed to trigger misbehavior. Understanding adversarial vulnerabilities is essential for anyone deploying AI in security-sensitive contexts.

Example in practice

A security researcher who adds invisible unicode characters to a malware sample's description to cause an AI classifier to label it as benign has mounted an adversarial attack against the model.

Learn more

See Adversarial Attack applied in a professional context through this free course.

AI Strategy for Leaders