Risk Assessment and Analytical Procedures with AI
Deliberate Academy Editorial Team
Reviewed for accuracy and professional relevance
Enjoying the course?
Sign up free to track your progress and earn a verified certificate when you pass.
- Apply AI to risk identification under ISA 315 without allowing the tool output to substitute for the required understanding of the entity
- Explain why the precision of the expectation, not the sophistication of the model, determines the assurance a substantive analytical procedure provides under ISA 520
- Set and document a threshold for investigation that is derived from performance materiality rather than from the tool default
- Recognise when a model-generated expectation is circular because it was trained on the unaudited data it is being used to test
Analytical procedures are where AI is most immediately useful in an audit and most quietly misused. The usefulness is obvious: a model can build an expectation from many more variables than a two-period comparison, across every account, in minutes. The misuse is subtle, and it almost always comes down to one of two things — an expectation that is not precise enough to detect a material misstatement, or an expectation that was derived from the very data it is supposed to challenge.
ISA 315: Understanding the Entity Is Not Delegable
ISA 315 requires the auditor to obtain an understanding of the entity and its environment sufficient to identify and assess risks of material misstatement. AI can genuinely contribute here. It can process several years of financial data alongside industry benchmarks to surface unusual relationships. It can read board minutes, contracts, and correspondence far faster than a team can, and flag matters that warrant follow-up. It can compare the current year chart of accounts against prior year to surface new accounts that suggest new transaction types.
What it cannot do is hold the understanding. The standard requires the auditor to understand the entity, and an auditor who can only say "the tool identified inventory valuation as a high-risk area" has not met that requirement. The understanding has to include why: which products, whose judgment drives the estimate, what changed commercially this year, which control the client relies on and whether the person operating it changed roles in September.
The productive pattern is to use AI to widen the net at the start and then narrow it with human enquiry. The tool proposes candidate risks from the data; the team tests each candidate against what it knows about the business and discards the ones that are artefacts. What must not happen is the reverse — the risk assessment memo assembled from the tool output, with the team's understanding reverse-engineered from it afterwards.
A risk assessment that no member of the team can defend in a conversation with the audit committee has failed ISA 315 regardless of how comprehensive the underlying analysis was. If the only answer to "why is this a significant risk?" is "the model ranked it highest," the understanding requirement has not been met.
ISA 520 and the Precision of the Expectation
A substantive analytical procedure provides audit evidence by comparing a recorded amount against an independent expectation and investigating differences. ISA 520 requires the auditor to evaluate the reliability of the data used, develop an expectation precise enough to identify a material misstatement, and determine the amount of difference that is acceptable without investigation.
That middle requirement is where AI-assisted analytics most often falls down, and the failure is counter-intuitive: a more sophisticated model does not automatically produce a more precise expectation in the audit sense.
Consider a model that predicts monthly revenue from headcount, marketing spend, seasonality, and prior-year trend, and produces an expectation within plus or minus 9 percent. If performance materiality is 4 percent of revenue, that expectation cannot detect a material misstatement — a misstatement of 6 percent of revenue sits comfortably inside the model's own error band. The procedure produces a reassuring result and provides essentially no assurance. The model is not wrong; it is simply not precise enough for the purpose, and no amount of additional model complexity fixes that if the underlying relationships are genuinely that loose.
The disciplined question is therefore not "how good is the model?" but "is the expectation tighter than performance materiality, and can I demonstrate that?" If the answer is no, the procedure can still be used to direct attention — role two from lesson one — but it cannot carry substantive assurance, and the file must not claim that it does.
Disaggregation Beats Sophistication
The most reliable way to increase precision is not a better algorithm; it is disaggregation. An expectation built at the total revenue level is nearly always imprecise, because offsetting movements across products, regions, and channels cancel out. The same expectation built by product line by month is far tighter, because there is less room for compensating errors to hide.
This is where AI genuinely helps, and it helps by doing volume rather than by being clever. Building 240 disaggregated monthly expectations by hand is impractical; building them programmatically is routine. The precision gain comes from the disaggregation, and AI makes the disaggregation affordable.
The same logic applies to the data. A model that predicts payroll from headcount, grade mix, and the timing of the annual pay award will usually beat one that predicts payroll from last year's payroll, because it uses independent operational data rather than the accounting figure it is meant to be testing.
Circularity: The Failure That Looks Like Success
The most serious analytical failure specific to machine-learning approaches is circularity. If a model is trained on the current year's unaudited general ledger and then used to identify entries that deviate from expected patterns, the model has learned the misstatement as normal. A systematic error running through the whole year — a misapplied revenue recognition policy, a recurring incorrect accrual, a fraud scheme with consistent characteristics — becomes part of the baseline. The model will then confidently report that the population is unremarkable.
This is a genuinely dangerous failure mode because it produces a clean result with high apparent coverage. It is most likely to hide exactly the pervasive, policy-level misstatements that matter most.
Three defences, in order of strength:
- Train or calibrate on independently reliable data: audited prior periods, operational data from outside the finance system, or external benchmarks.
- Test the model against known errors: seed the population with realistic misstatements at performance materiality and confirm the approach detects them. This is the closest thing to a control test for the tool.
- Pair pattern-based approaches with criteria-based ones: rules derived from the risk (postings outside business hours, entries by unexpected users) do not learn from the data and therefore cannot be trained to accept a misstatement as normal.
An audit team builds a machine learning model on the client's current-year unaudited general ledger to predict expected account balances, then investigates accounts where the actual balance deviates materially from the prediction. Almost nothing is flagged. Why is this result weak evidence that the population is free from material misstatement?
Select one answer.
Setting the Investigation Threshold
ISA 520 requires the auditor to determine the amount of difference from the expectation that is acceptable without investigation, and that amount must be influenced by performance materiality and the desired level of assurance.
Analytics platforms ship with defaults — top 100 anomalies, scores above 0.8, deviations beyond two standard deviations. None of these is derived from your materiality. A threshold expressed in standard deviations is a statement about the distribution's shape, not about whether a difference matters to the financial statements. Two standard deviations on a tight population may be well below the clearly trivial threshold, generating pure noise; on a volatile population it may be several times performance materiality, silently accepting differences that should have been investigated.
Set the threshold in currency, derive it from performance materiality, document the derivation, and then configure the tool to match it. Where a tool only supports a score-based cut-off, establish the equivalent currency impact at the chosen score and document that mapping. The number of items the tool returns should be a consequence of your threshold, not the other way round.
A model precise to plus or minus 11 percent used to substantively test a balance with 5 percent performance materiality
Context
An engagement team used the firm's analytics platform to build a predictive expectation for a manufacturing client's cost of sales, incorporating production volumes, commodity price indices, and headcount. The model was well constructed and the team reduced substantive detail testing on cost of sales on the strength of it, documenting the analytical procedure as the primary source of assurance.
Action
During an internal quality review, the reviewer asked what the model's prediction interval was and how it compared with performance materiality. The model's historical accuracy was plus or minus 11 percent against a performance materiality set at 5 percent of the cost of sales balance. The recorded figure fell within the predicted range, but the range was wide enough to accommodate a misstatement of more than twice performance materiality without producing any signal at all.
Outcome
The team retained the model as a risk-direction tool but could no longer treat it as substantive assurance. They rebuilt the analysis disaggregated by product family and by month, which narrowed the interval to approximately 4 percent for three of the five product families — those three then supported substantive reliance, and detail testing was reinstated for the remaining two. The firm subsequently required that any analytical procedure relied on substantively record its prediction interval alongside performance materiality on the face of the working paper, so the comparison is visible to a reviewer without recalculation.
A model predicts cost of sales to within plus or minus 11 percent and the recorded figure falls inside that range. Performance materiality is 5 percent of the balance. What assurance has the procedure provided?
Select one answer.
Exercise
Your Task
Take a substantive analytical procedure from a recent file, ideally one supported by a tool or model. Establish three numbers and write them on the working paper: (1) performance materiality for that area, (2) the precision of the expectation, expressed as a currency range rather than a percentage accuracy claim, and (3) the investigation threshold actually applied, with its derivation. Then answer two questions in writing: could a misstatement equal to performance materiality have been detected by this procedure, and was the expectation derived from data independent of the recorded amounts being tested? If the answer to either is no, state what the procedure can legitimately be used for instead.
Success looks like
- Expectation precision is expressed as a currency range and compared directly against performance materiality
- The investigation threshold is traced back to performance materiality rather than to a platform default
- Any circularity between the expectation source and the tested data is identified explicitly
- Where the procedure cannot carry substantive assurance, the file is honest about downgrading it to risk direction
Watch out for
- Treating a high model accuracy percentage as evidence of sufficient precision without comparing it to performance materiality
- Accepting a standard-deviation or top-N default as the investigation threshold
- Building an expectation from the current-year unaudited ledger and describing it as independent
- ISA 315 requires the auditor to understand the entity. AI can widen the net for candidate risks, but a risk assessment the team cannot defend in conversation has not met the standard, however comprehensive the underlying analysis.
- Under ISA 520 the assurance from a substantive analytical procedure depends on whether the expectation is precise enough to detect a material misstatement. An expectation wider than performance materiality provides essentially no substantive assurance regardless of model sophistication.
- Precision comes from disaggregation and from independent operational data far more reliably than from algorithmic sophistication. AI helps mainly by making fine-grained disaggregation affordable at scale.
- Circularity is the critical machine-learning-specific failure: a model trained on the unaudited population learns pervasive misstatement as normal and returns a clean result. Defend with independent calibration data, seeded-error testing, and criteria-based rules.
- Investigation thresholds must be derived from performance materiality and documented, not inherited from a platform default expressed in standard deviations or top-N items.