AI Risk Assessment and Classification in Practice
Deliberate Academy Editorial Team
Reviewed for accuracy and professional relevance
You're 6 lessons in — don't lose your progress.
Sign up free to save where you are and earn a verified certificate when you pass.
- Explain why an internal, multi-factor risk classification methodology is necessary in addition to a regulatory tier such as the EU AI Act's
- Apply a five-factor scoring rubric — regulatory tier, data sensitivity, decision consequence, reversibility, and population vulnerability — to classify a described AI system
- Map internal risk tiers to proportionate control intensity, so classification drives a concrete governance decision rather than a label with no consequence
- Identify classification drift as a recurring failure mode and design a trigger-based re-classification process to prevent it
The EU AI Act's tier structure, covered earlier in this course, gives you a legally grounded classification for any system operating in scope of that regulation. It does not, by itself, give you everything a compliance function needs to run an internal governance program, for two reasons. First, most organizations operate in multiple jurisdictions, and a system that is minimal-risk under the Act may still carry significant internal risk — reputational, financial, operational — that a purely legal classification does not capture. Second, the Act's tiers are broad categories; within the high-risk tier alone, a credit-scoring tool touching 40,000 loan applications a year and a niche internal tool used by three underwriters on a specialized product line are both "high-risk," but they do not warrant identical governance intensity. An internal, multi-factor classification methodology sits alongside the regulatory tier and answers the question the Act's tier alone cannot: exactly how much governance rigor does this specific system need?
A Five-Factor Scoring Rubric
A practical internal classification methodology scores each AI system across five factors.
Regulatory tier. The system's classification under the EU AI Act (or an equivalent regime), established using the decision tree from earlier in this course. This is the baseline factor — a prohibited or high-risk regulatory tier sets a floor for internal risk that other factors can raise but should never be allowed to lower.
Data sensitivity. What categories of data does the system process — public information, internal business data, confidential customer data, or special-category personal data such as health, biometric, or financial information? Higher sensitivity raises the internal risk score.
Decision consequence. If the system produces an error, how severe is the resulting harm — a minor inconvenience, a moderate financial or operational cost, a significant harm to an individual's opportunities (employment, credit, housing), or a severe and potentially irreversible harm to health, safety, or fundamental rights?
Reversibility. Can an erroneous AI-influenced decision realistically be identified and corrected after the fact, or does the harm compound or become effectively permanent before anyone would notice — a declined loan application is more reversible, in principle, than an incorrect medical triage decision made under time pressure.
Population scale and vulnerability. How many people does the system affect, and does that population include groups with elevated vulnerability — children, people with disabilities, people in financial distress, or populations with limited practical ability to contest an automated decision?
A simple, transparent scoring approach works better in practice than an elaborate weighted formula that no one outside the compliance team can reproduce. Score each of the five factors on a 1-to-4 scale (minimal, low, moderate, severe), sum the scores, and map the total to an internal tier: roughly 5-8 as Low, 9-13 as Medium, 14-17 as High, and 18-20 as Critical is a reasonable starting calibration — adjust the thresholds to your organization's actual risk appetite once you have scored a representative sample of your AI inventory and can see where systems naturally cluster. What matters most is that the scoring is documented and repeatable, not that the formula is sophisticated.
Differentiating Governance Intensity Across a Portfolio — Regional Insurer
Context
An insurer's AI inventory included three systems that had all been informally labeled 'high-risk' by different business units simply because each involved automated decision-making: an AI claims-triage tool processing roughly 85,000 claims a year and directly influencing payout amounts; a niche AI tool used by four commercial underwriters to flag unusual policy language in specialty marine insurance contracts, covering fewer than 200 policies a year; and an AI-assisted internal tool that drafted first-pass responses to routine customer service emails, reviewed by an agent before sending.
Action
The Head of Model Risk applied the five-factor rubric to all three. The claims-triage tool scored high across data sensitivity, decision consequence, reversibility (claims decisions are difficult to unwind once paid or denied), and population scale, landing in the Critical tier. The marine underwriting tool scored moderate on regulatory tier and data sensitivity but low on population scale given its narrow use, landing in the Medium tier. The customer service drafting tool, despite superficially resembling the other two as an 'automated decision' tool, scored low across nearly every factor because a human always reviewed and could easily correct the output before anything left the organization, landing in the Low tier.
Outcome
The differentiated classification allowed the insurer to concentrate its limited compliance and model-risk capacity where it mattered most: the claims-triage tool received quarterly bias audits, a dedicated technical file, and monthly human-oversight sampling; the marine underwriting tool received an annual review and a lighter-weight documentation standard; and the customer service tool was moved to a self-certification track with periodic spot checks rather than continuous compliance attention. The Head of Model Risk's assessment was that treating all three as equally 'high-risk,' as the informal business-unit labeling had done, would have either over-invested compliance capacity in the lowest-risk system or under-invested it in the claims tool that actually carried the most exposure.
Two AI systems are both classified as high-risk under the EU AI Act because they both operate in the employment domain: one screens 15,000 job applications a year for a large retailer, and the other supports internal promotion decisions for a 12-person specialist team at a boutique firm. Using the five-factor rubric from this lesson, what is the most defensible reason these two systems might still warrant different levels of internal governance intensity, despite sharing the same regulatory tier?
Select one answer.
Exercise
Your Task
Apply the five-factor rubric to three AI systems from your organization or a plausible one. For each, assign a 1-to-4 score on regulatory tier, data sensitivity, decision consequence, reversibility, and population scale/vulnerability, sum the total, and assign an internal tier using the calibration in this lesson. Then write one sentence for each system stating what governance intensity that tier should trigger — documentation depth, review frequency, and who must sign off.
Success looks like
- Each of the five factors has an individually justified score, not just a total number pulled from intuition
- The three systems produce at least two different internal tiers, demonstrating the rubric actually differentiates rather than defaulting everything to the same result
- The governance-intensity sentence for each system names a concrete difference in review frequency or documentation depth, not just a restatement of the tier name
Watch out for
- Scoring every factor the same for every system out of convenience, which defeats the purpose of a multi-factor rubric
- Letting a system's political or budgetary importance influence the score instead of the five defined factors
- Assigning a low internal tier to a system that is genuinely high-risk under the EU AI Act — remember the regulatory tier acts as a floor other factors can raise but should not be used to justify lowering
Hint
Score the regulatory tier factor first using the decision tree from earlier in this course, then treat that as a floor — if the system is high-risk under the Act, its total score should not land in your Low internal tier no matter how the other four factors score.
Classification drift is one of the most common and most preventable compliance failures in AI governance. A system scored Low risk as an internal proof-of-concept, with a handful of users and no bearing on any real decision, can quietly evolve into a production system making consequential decisions about real people — without anyone re-running the classification. Build re-classification triggers into your process: a change in the system's user population, a change in what decision the system's output feeds into, a change in the data it processes, or a fixed review interval (annually, at minimum, for any Medium tier or above system) should each independently trigger a mandatory re-score, not wait for someone to notice the drift on their own initiative.
An AI tool was originally deployed as an internal pilot to help three analysts draft first-pass summaries of vendor contracts, and was scored Low risk at that time. Eighteen months later, the tool has been quietly adopted by the procurement team to generate risk flags that directly inform go/no-go vendor approval decisions across the organization, with no formal re-deployment process and no updated risk classification. What does this lesson identify as the correct governance response?
Select one answer.
- A regulatory tier such as the EU AI Act's is a necessary floor for AI risk classification, but a multi-factor internal methodology is needed to differentiate governance intensity within and alongside that floor.
- Score systems across five factors — regulatory tier, data sensitivity, decision consequence, reversibility, and population scale/vulnerability — using a simple, documented, repeatable scale rather than an opaque formula.
- Map internal risk tiers to concrete, differentiated governance intensity — documentation depth, review frequency, and sign-off authority — so classification drives an actual operational decision rather than producing a label with no consequence.
- The regulatory tier factor should act as a floor other factors can raise but never be used to lower — a system that is high-risk under the Act cannot be scored into a Low internal tier regardless of how the other factors look.
- Classification drift — a system's risk profile changing without triggering re-classification — is one of the most common and most preventable compliance failures; build concrete triggers (population change, decision-use change, data change, fixed review interval) rather than relying on someone noticing.