Where AI Fits in the Audit Lifecycle
Deliberate Academy Editorial Team
Reviewed for accuracy and professional relevance
- Distinguish the three roles AI plays in an audit — preparing the file, directing attention, and producing evidence — and explain why the evidence role carries a materially higher justification burden
- Map an AI-assisted procedure to the specific financial statement assertion it provides assurance over, rather than describing it only as a tool that was used
- Explain why "the tool flagged it" is not a documented audit procedure and what has to be added to make it one
- Identify which parts of the audit lifecycle are least suitable for AI assistance and articulate the reason in standards terms
Most discussions of AI in audit collapse very different activities into one category. Drafting a client request list, ranking 400,000 journal entries by unusualness, and concluding that a revenue population is free from material misstatement are all "using AI in the audit," and they carry completely different obligations. Getting this distinction right is the foundation for everything else in this course, because it determines how much you have to be able to explain.
Three Roles, Three Burdens
Role one: preparing the file. AI drafts something a human would otherwise have typed — a lead schedule structure, a confirmation letter, a narrative for a standard file section, a summary of the prior year file. The output is not evidence. It is paperwork that sits alongside evidence. The obligation here is straightforward review for accuracy, which you already apply to work prepared by a junior. This is the territory covered in AI for Accountants.
Role two: directing attention. AI processes a population and tells you where to look — ranking entries by anomaly score, clustering contracts by unusual clause combinations, highlighting accounts whose movement diverges from a modelled expectation. The output is not evidence either. It is a scoping decision. But it is a scoping decision that quietly determines what you will and will not test, which means a flaw in the tool becomes a flaw in your coverage, invisibly.
Role three: producing evidence. AI performs the test itself and its output is what you rely on to conclude. The model reads 3,000 lease contracts and extracts the commencement dates you use to test the lease liability calculation. The tool reperforms a revenue cut-off test across the full population and reports the exceptions. Here the AI output is the audit evidence, and ISA 500 applies to it in full: you must be able to say why that evidence is sufficient and appropriate.
The burden rises sharply across those three. Firms get into trouble when a role-three procedure is documented as though it were role one.
The question that separates the roles is not how sophisticated the tool is. It is: if this output were wrong, would the audit conclusion be wrong? If yes, the output is evidence, and it needs the justification ISA 500 requires — not the lighter review you would apply to a drafted schedule.
Assertions, Not Tools
Audit procedures are not justified by what tool performed them. They are justified by which assertion they provide assurance over: existence, completeness, accuracy, valuation, rights and obligations, cut-off, classification, presentation.
This matters because AI tools are usually sold and described in tool language, and auditors then document them in tool language. "We used the platform's anomaly detection module across the journal entry population" describes an activity. It does not say what was tested. Compare it with: "To address the risk of management override, we tested the completeness of our journal entry population against the general ledger control totals, then applied criteria-based selection across all 412,000 entries to identify entries exhibiting characteristics associated with override risk — postings by users outside finance, entries to seldom-used accounts, round-sum amounts above materiality, and entries posted outside normal business hours."
The second version identifies the risk being addressed, the population and how its completeness was established, and the specific criteria applied. It happens to have been executed by software, which is unremarkable. The first version is unreviewable, because a reviewer cannot tell whether the procedure was appropriate to the risk.
A useful discipline: write the procedure description first, in assertion language, without naming the tool. If you cannot write it, you do not yet understand what the tool did — and that is a finding about your own file, not about the tool.
Population Completeness Comes First
Full-population analysis has a failure mode that sampling does not. When you select a sample of forty from a population, you have already established the population. When a tool analyses "the whole population," the population is whatever was extracted into the tool, and that is not the same thing.
Extractions silently drop data all the time: date-range filters that exclude the stub period, entity codes missed in a group with a newly acquired subsidiary, entries in a suspense account excluded by an account-range filter, a source system whose feed failed for two days. The analysis then runs perfectly across an incomplete population and returns a confident, clean result.
Reconciling the extraction back to the general ledger control totals — record count and debit and credit totals, by period and by entity — is the first procedure in any full-population approach and the one most often skipped. It is unglamorous and it is the whole foundation. A full-population test over an incomplete population provides less assurance than a properly selected sample over a complete one, while looking considerably more impressive in the file.
An audit team uses an AI tool to analyse the full accounts payable population and identify duplicate payments. The tool reports 14 potential duplicates, all of which are investigated and resolved. What is the most significant risk in this approach as described?
Select one answer.
Where AI Fits Poorly
Some parts of the audit resist AI assistance for reasons rooted in the standards rather than in technical capability.
Materiality and risk assessment judgments. Setting overall materiality, performance materiality, and the clearly trivial threshold requires judgment about the users of the financial statements and their information needs. A model can compute benchmarks. It cannot decide that this entity's users are lenders focused on covenant compliance and that a profit-based benchmark therefore understates what matters.
Evaluating the sufficiency of evidence. ISA 500 asks whether the evidence obtained is enough. That is a judgment about the aggregate of the audit work against the assessed risks, and it is the auditor's alone.
Concluding on going concern and significant estimates. These depend on evaluating management intent, the reasonableness of assumptions, and events after the reporting period — a synthesis of contradictory evidence, which is precisely where models perform worst and sound most confident.
Communicating with those charged with governance. Judgment about what matters enough to report, and how to frame it, is not a drafting problem.
A full-population revenue test that covered 83 percent of the population
Context
An audit team applied a full-population cut-off test to a client's revenue transactions using the firm's analytics platform, testing that revenue was recognised in the correct period against despatch documentation. The file documented full-population coverage across 96,000 transactions and the team reduced planned substantive sampling accordingly.
Action
During the manager review, the manager asked for the reconciliation between the transaction extract and the revenue figure in the trial balance. It had not been performed. When the team ran it, the extract totalled 83 percent of recognised revenue. The client had migrated to a new billing system nine weeks before year end, and the extract had been taken from the legacy system only. The 17 percent shortfall was the post-migration period — which included the year-end cut-off window the test was specifically designed to address.
Outcome
The team obtained the post-migration extract, reconciled both to the trial balance, and re-ran the test across the combined population. Two cut-off errors were identified in the post-migration data, both arising from a configuration difference in how the new system stamped despatch dates. Neither would have been found by the original procedure. The firm added a mandatory reconciliation step, evidenced by a signed-off control-total comparison, as a gate before any full-population procedure could be recorded in a file, and made it a standing question in the methodology review checklist.
Documenting the Role You Are In
The practical test at file level is whether a reviewer, and later an inspector, can tell which role the AI played. Three things make that visible:
- What the procedure tested, in assertion language, independent of the tool.
- How the population was established, including the reconciliation to control totals for any full-population work.
- What you did with the output, including how exceptions were resolved and how you satisfied yourself that items not flagged did not require attention.
The third point is the one that gets omitted. A file that documents nineteen investigated anomalies but says nothing about the basis for accepting the remaining 411,981 entries has documented the interesting part and left out the part that carries the assurance.
This lesson gives one question that decides which of the three roles an AI-assisted procedure occupies. What is that question?
Select one answer.
Exercise
Your Task
Take one AI-assisted or analytics-assisted procedure from a recent audit file. Classify it as role one (file preparation), role two (directing attention), or role three (producing evidence), and write down your reasoning. Then, without naming the tool anywhere, write a two- to four-sentence description of the procedure in assertion language: which assertion it addressed, how the population was established and reconciled, what criteria were applied, and what basis you have for the items that were not flagged. Note any of those four elements you cannot currently answer from the file as it stands.
Success looks like
- The role classification is justified by the consequence test — whether a wrong output would produce a wrong audit conclusion — not by how advanced the tool is
- The procedure description names a specific assertion and does not depend on the tool name to make sense
- Any missing element, particularly the basis for accepting unflagged items, is honestly identified rather than reasoned around
Watch out for
- Classifying a procedure as file preparation because the output was reviewed, when the audit conclusion actually rests on it
- Treating the tool vendor description of a feature as a description of the audit procedure performed
- AI plays three distinct roles in an audit — preparing the file, directing attention, and producing evidence — and only the third makes the output subject to the full sufficient-appropriate evidence test under ISA 500. The distinguishing question is whether a wrong output would produce a wrong audit conclusion.
- Procedures are justified by the assertion they address, not by the tool that executed them. If you cannot describe the procedure in assertion language without naming the tool, you do not yet understand what it did.
- Full-population work depends entirely on population completeness. Reconciling the extract to general ledger control totals by period and entity is the first procedure, and a full-population test over an incomplete population gives less assurance than a sample over a complete one.
- Materiality setting, evidence sufficiency, going concern, significant estimates, and governance communication resist AI assistance for reasons grounded in the standards, not in current technical limitations.
- Documentation must cover the basis for accepting items the tool did not flag. Recording only the investigated exceptions omits the part of the procedure that carries most of the assurance.