Auditing Clients Who Use AI
Deliberate Academy Editorial Team
Reviewed for accuracy and professional relevance
You're 6 lessons in — don't lose your progress.
Sign up free to save where you are and earn a verified certificate when you pass.
- Identify where a client AI system sits in the flow of transactions and determine whether it affects an assertion you are required to audit
- Apply ISA 540 to an accounting estimate produced by a machine learning model, including evaluating the method, assumptions, and data
- Extend IT general controls thinking to cover model change, retraining, and data pipelines, which conventional ITGC scoping misses
- Recognise model drift as a period-specific risk that makes prior-year comfort unreliable
The previous lessons dealt with AI in your hands. This one deals with AI in the client's, which is the larger and less discussed problem. Clients are deploying models into revenue recognition, credit loss estimation, inventory obsolescence, and fraud screening, and those models are producing numbers that land in the financial statements you sign.
The first mistake is treating this as a technology topic. It is an ordinary audit problem — an estimate or a process affecting an assertion — with an unfamiliar mechanism. The second mistake is the opposite: treating it as entirely ordinary and missing the two things genuinely new about it, which are drift and opacity.
Scoping: Does the Model Touch an Assertion?
Clients use AI for many things that have no bearing on the financial statements. A marketing team using generative AI for campaign copy is not an audit matter. The scoping question is whether the system affects the initiation, recording, processing, or reporting of transactions, or produces an input to an amount or disclosure in the financial statements.
Systems that commonly do:
- Expected credit loss models under IFRS 9. The most consequential case in financial services and increasingly in corporates with significant receivables.
- Inventory obsolescence and net realisable value models driven by demand forecasting.
- Revenue recognition where an AI system determines contract terms, performance obligation satisfaction, or variable consideration estimates.
- Fair value estimation for assets without observable market prices.
- Automated contract review feeding lease classification or revenue treatment.
- Fraud and credit screening that determines which transactions are accepted, affecting the population of recorded transactions.
The mapping to make with the client is straightforward and worth doing formally: list the AI systems in use, and for each, identify the financial statement line items and assertions affected. Most entries on that list resolve to "none," and the exercise is worth it for the few that do not. Do it as part of understanding the entity under ISA 315, not as a separate technology review.
Clients frequently do not know which of their systems contain a model. A feature marketed as smart forecasting or automated matching inside an ERP module may be a model that materially affects an estimate, and the finance team may regard it simply as how the system works. Ask about functionality and outputs rather than asking whether they use AI.
ISA 540 Applied to a Model
Where a model produces an accounting estimate, ISA 540 (Revised) applies and provides more structure than most auditors expect. It requires evaluating the method, the assumptions, and the data — and those three map cleanly onto a machine learning model.
The method. What kind of model, and is it appropriate for the estimate? A model optimised for ranking is not necessarily appropriate for producing an unbiased expected value. The relevant question is whether the method suits the measurement objective in the applicable financial reporting framework, and management should be able to explain why this method was selected.
The assumptions. In a machine learning model the assumptions are less visible than in a discounted cash flow, but they exist and they are usually significant. The training period embeds an assumption that it is representative of the forecast period. The feature set embeds an assumption about what drives the outcome. The loss function embeds an assumption about the relative cost of over- and under-estimation, which directly affects whether the estimate is neutral or biased. Management bias is precisely what ISA 540 asks you to look for, and in a model it hides in these choices rather than in a rate that has been nudged.
The data. Training data completeness, accuracy, and relevance. A credit loss model trained exclusively on a benign period has not seen the conditions that make the estimate matter.
The standard also requires evaluating whether the estimate is reasonable in the context of the framework, and considering indicators of possible management bias. Two model-specific indicators are worth having in mind: a model retrained or reconfigured close to period end, and a post-model overlay or management adjustment applied to the output. Overlays are legitimate and common, particularly in expected credit loss, but they are also the easiest place to introduce bias, because they sit outside the model's discipline entirely. An overlay should receive the same scrutiny as any other significant management judgment, and often more.
Controls: Where Conventional ITGC Scoping Falls Short
Standard IT general controls cover access, change management, and operations. Applied to a model, that framework has gaps, because a model can change materially without any code changing.
The controls that matter:
- Model change control. Approval and testing for changes to model logic, features, or hyperparameters. Usually covered by existing change management if the model lives in a controlled codebase, and frequently not covered at all if it lives in a notebook or a spreadsheet.
- Retraining governance. This is the gap. If a model retrains automatically on a schedule, its behaviour changes without any change event that conventional change management would capture. Who approves a retrain, what validation runs before the new version goes live, and is there a record of which version produced the period's figures?
- Data pipeline controls. Completeness and accuracy of the data feeding the model, which is an ordinary ITGC concern but frequently unscoped because the pipeline is newer than the systems around it.
- Output review. A human control over the model's output. Its effectiveness depends entirely on whether the reviewer has the information and standing to challenge, which is the automation bias problem covered in the next lesson.
- Version records. Whether the client can tell you which model version produced the estimate in these financial statements. If they cannot, you have a documentation problem that limits what you can conclude.
Drift: Why Last Year Does Not Carry Forward
The genuinely novel risk is drift. A model's performance degrades as the world moves away from its training data, and this happens without any change to the model. Nothing breaks, no error is raised, and the outputs remain plausible. A demand forecasting model driving inventory provisioning becomes progressively less accurate as the product mix shifts, and the provision becomes progressively less reliable while the process appears entirely stable.
For audit, drift has a specific consequence: a model validated in a prior year provides limited comfort in the current year. The conclusion that the model was appropriate last year is a conclusion about a relationship between the model and last year's conditions.
What to look for each period: whether management monitors model performance against actual outcomes, what threshold triggers investigation or retraining, whether that monitoring was performed for the current period, and what it showed. Back-testing the model's prior-period predictions against actual outcomes is the most direct evidence available, and where management does not do it, the auditor often can, using the entity's own realised data. It is one of the more powerful procedures available in this area and it is underused.
A client uses a machine learning model to estimate inventory obsolescence. The model was validated by an external specialist two years ago, the code has not changed since, and the client's change management controls over the codebase were tested and found effective. What is the most significant remaining risk?
Select one answer.
A stable model, an unchanged codebase, and a provision that had drifted 40 percent adrift
Context
A lender used a machine learning model to calculate expected credit loss on an unsecured consumer portfolio. The model had been independently validated at implementation, change management controls were effective and tested, and the model had not been retrained in 26 months. The prior year audit had concluded the estimate was reasonable, and the current year team planned to rely on the same approach.
Action
Rather than repeating the prior-year approach, the director asked for back-testing of the model's predictions against realised losses for the preceding eight quarters. Management did not perform this monitoring, so the team performed it using the lender's own realised loss data. The model had predicted losses accurately for the first three quarters of the period examined and had then diverged steadily, under-predicting realised losses by an average of 40 percent in the most recent two quarters. The cause was a shift in the portfolio towards a customer segment barely represented in the training data, following a change in the lender's acquisition channels.
Outcome
The provision was understated materially and was increased before the accounts were signed. The absence of ongoing model performance monitoring was reported to those charged with governance as a significant deficiency, and management implemented quarterly back-testing with a defined retraining trigger. The director observed that every control the audit had previously relied on was operating effectively — the failure was that none of those controls was designed to detect a model becoming wrong while remaining unchanged.
In a discounted cash flow, management bias in an estimate typically shows up as a nudged rate. This lesson says that in a machine learning model it hides somewhere else. Where?
Select one answer.
Exercise
Your Task
For a current client, build the AI system inventory described in this lesson: list each system or system feature that involves a model, and for each, record the financial statement line items and assertions it affects. For any system that affects an assertion, establish four things and note which you cannot answer: (1) which model version produced the current period figures, (2) whether and when the model was last retrained or reconfigured, (3) whether management monitors predicted against actual outcomes and what the latest results showed, and (4) whether any post-model overlay or management adjustment was applied to the output, and on what basis.
Success looks like
- The inventory is built by asking about functionality and outputs rather than asking whether the client uses AI
- Each in-scope system is mapped to specific line items and assertions, with out-of-scope systems explicitly recorded as such
- Any post-model overlay is identified and treated as a significant management judgment
- Absence of management back-testing is recorded as a finding, with consideration of whether the team can back-test using realised data
Watch out for
- Relying on prior-year model validation as current-period evidence when drift produces no change event
- Scoping ITGCs around code change only, missing automated retraining that alters behaviour without a code change
- Scope by asking whether a system affects the initiation, recording, processing or reporting of transactions or produces an input to a reported amount. Ask clients about functionality and outputs, because finance teams often do not know a feature contains a model.
- ISA 540 (Revised) maps directly onto models: evaluate the method, the assumptions — training period, feature set, and loss function all embed significant assumptions — and the data. Management bias hides in those design choices rather than in an adjusted rate.
- Post-model overlays and management adjustments sit outside the model discipline entirely and are the easiest place to introduce bias, so they warrant scrutiny at least equal to any other significant judgment.
- Conventional ITGC scoping misses retraining governance. A scheduled automatic retrain changes model behaviour with no change event, so ask who approves retrains, what validation gates them, and which version produced the period figures.
- Drift is the novel risk: performance degrades with no code change and no error signal, so prior-year validation gives limited current-period comfort. Back-testing predictions against realised outcomes is the most direct evidence and can often be performed by the audit team where management does not do it.