AI Glossary
Interpretability
What is Interpretability?
The degree to which the internal mechanisms of an AI model can be understood by humans. Modern large neural networks are largely 'black boxes' — their internal representations are difficult to interpret directly. Interpretability research aims to understand why models produce specific outputs, which is critical for debugging failures and building trustworthy AI systems.
Example in practice
A regulator requiring a bank to explain why its AI rejected a loan application is demanding interpretability — the ability to trace the model's decision back to specific input features rather than accepting a black-box output.
Learn more
See Interpretability applied in a professional context through this free course.
AI Strategy for Leaders →