Audit Evidence and Documentation Standards for AI-Assisted Work
Deliberate Academy Editorial Team
Reviewed for accuracy and professional relevance
You're 5 lessons in — don't lose your progress.
Sign up free to save where you are and earn a verified certificate when you pass.
- Apply the ISA 500 sufficient-and-appropriate test to evidence produced or processed by an AI tool, including the reliability of the underlying data
- Meet the ISA 230 experienced-auditor standard for an AI-assisted procedure, so the work can be reperformed and understood without the original team
- Distinguish what must be documented about the procedure from what must be documented about the tool
- Retain the right artefacts so an AI-assisted conclusion remains supportable years later when the tool version has changed
Everything in the previous lessons converges here. A procedure can be well designed, run across a complete population, and produce genuine findings, and still fail at file level because the documentation does not let anyone else understand or reperform it. This is the most common inspection finding in AI-assisted audit work, and it is entirely avoidable.
ISA 500: Sufficient and Appropriate
ISA 500 requires the auditor to obtain sufficient appropriate audit evidence. Sufficiency is quantity; appropriateness is relevance and reliability. Where a tool produced or processed the evidence, the standard's requirement to consider the reliability of the information used as audit evidence does the heavy lifting, and it applies to three distinct things.
The source data. Where did it come from, and is it complete and accurate? This is the reconciliation from lesson one, plus consideration of whether the data passed through any transformation that could have altered it. Data exported from a client system, reshaped in a spreadsheet, and loaded into a tool has three points at which it could change.
The processing. Did the tool do what you believe it did? For a rules-based routine this is testable directly. For a model-based approach it usually cannot be verified by inspection, which is why seeded-error testing matters: constructing known misstatements at performance materiality and confirming the routine detects them is the practical evidence that the processing works.
The output. Is the result what it appears to be? An exception report is a claim about the population, and reperforming a subset by hand is the ordinary way to corroborate it. Taking twenty flagged items and twenty unflagged items and checking both by hand tests the routine in both directions — the unflagged sample is the more informative half and is usually omitted.
Reliability under ISA 500 is a property of the evidence in context, not a property of the vendor. A tool being widely used, well reviewed, or built by a large firm says nothing about whether the data extracted for this engagement was complete or the criteria appropriate to this entity's risks. Those remain engagement-level judgments every time.
ISA 230: The Experienced Auditor Test
ISA 230 requires documentation sufficient for an experienced auditor with no previous connection to the audit to understand the nature, timing and extent of the procedures performed, the results and evidence obtained, and the significant matters and conclusions reached.
Apply that literally to AI-assisted work and most files fall short. An experienced auditor reading "anomaly detection was run over the journal entry population and 19 items were investigated" cannot understand what was done. They cannot tell what the population was, what made an item anomalous, why 19 rather than 190, or what basis exists for the rest of the population.
The test is practical, not theoretical: could a competent auditor who has never met your team reproduce this procedure and get the same result? For AI-assisted work that requires the following in the file.
- The population and its reconciliation. Source system, extraction date and parameters, record count and control totals agreed to the ledger, by period and entity.
- The criteria or model configuration, in full. The actual rules, thresholds, and parameters — not a description of the category of rules. If a score cut-off was used, the cut-off and its derivation from performance materiality.
- The version of the routine. Which tool, which version, and which configuration, so a difference in later results can be attributed correctly.
- The complete output, or a reliable means of regenerating it. Not just the exceptions.
- The evaluation. How exceptions were resolved, how groups were formed and corroborated, and the basis for accepting unflagged items.
- The conclusion, tied to the assertion and the assessed risk.
Most of this is generated automatically by the tools and discarded because it is voluminous. Retaining it is cheaper than reconstructing it.
Procedure Documentation Versus Tool Documentation
A distinction that saves considerable effort: the file documents the procedure; the firm documents the tool.
Tool-level matters — how the vendor validates the model, the firm's evaluation before approving it for use, version history, known limitations, the results of firm-level testing — belong in firm methodology documentation, referenced from the file. Repeating them in every engagement file is waste and, worse, tends to crowd out the engagement-specific content that actually matters.
Engagement-level matters cannot be delegated to firm documentation, because they change every time: this population, this reconciliation, these criteria mapped to these assessed risks, this exception evaluation, this conclusion. A file that leans on the firm's tool approval as though it answered the engagement questions has confused the two, and it is a common way for a file to look well documented while saying nothing about the audit.
Explainability, Proportionate to Reliance
Auditors sometimes conclude that model-based tools cannot be used because their output cannot be fully explained. That is too strong. The requirement is that the auditor understands the procedure well enough to evaluate the evidence, and how much understanding that demands scales with how much the conclusion depends on it.
A model used to rank items for attention, where every ranked item is then examined by an auditor applying their own judgment, needs limited explanation — the evidence comes from the examination, not the ranking. What the file does need is the basis for the population that was not ranked highly enough to examine.
A model whose output is relied on directly, with no further human evaluation of individual items, needs substantially more: what it was trained on, what it optimises for, how it was validated, its performance against seeded errors, and its known failure modes. If that cannot be established, the output can direct attention but cannot carry substantive assurance. That is a legitimate and often correct conclusion, and the mistake is not reaching it — the mistake is relying on the output anyway while documenting it as though it had been established.
Retention and the Reperformance Problem
Audit files are retained for years and may be inspected long after the engagement. AI tools change on vendor timetables. A model retrained six months after your audit will not reproduce your results, and a tool version two releases on may not accept your configuration.
This makes retention an active decision rather than a default. At minimum, retain the input data as extracted, the full configuration, the complete output, and the tool and version identifiers. Where a cloud tool cannot guarantee that a historical version remains available, the exported output and inputs are the only durable record, and they must be preserved as files in the engagement archive rather than left in the platform.
A useful discipline: assume the tool will be unavailable when the file is inspected, and ask whether the retained artefacts alone would let an experienced auditor understand and evaluate what was done. If the answer depends on being able to re-run the tool, the file is not yet ISA 230 compliant.
Two years after an audit, an inspection team reviews a file documenting a full-population accounts receivable analysis. The file records the tool name, that anomaly detection was applied, and the twelve exceptions investigated. The tool has since been upgraded twice and cannot reproduce the original output. What is the principal documentation failure?
Select one answer.
Reperforming twenty unflagged items found the error the exception report never would have
Context
A team used a rules-based routine to test that every capitalised asset addition above a threshold had supporting approval documentation and correct useful-life assignment, across a full population of 3,100 additions. The routine returned 46 exceptions. Firm methodology required reperformance of a subset of both flagged and unflagged items.
Action
The senior reperformed twenty flagged items, all correctly identified. She then reperformed twenty unflagged items by hand. Three of the twenty had useful lives inconsistent with the client's own capitalisation policy but had not been flagged, because the routine compared useful life against the asset category recorded in the fixed asset register rather than against the category implied by the asset description. Assets miscategorised at entry were tested against the wrong expected life and therefore passed.
Outcome
The error affected an estimated 190 additions and produced a depreciation misstatement below performance materiality but material to the fixed asset disclosure. The routine was corrected to test the description-to-category mapping as a separate attribute and re-run. The team noted that reperforming only the flagged items would have confirmed the routine worked perfectly on the items it caught and revealed nothing at all about the population it passed. The firm made two-directional reperformance mandatory for any routine relied on substantively.
The senior reperformed twenty flagged items and then twenty unflagged ones. What does this lesson say the unflagged half establishes that the flagged half cannot?
Select one answer.
Exercise
Your Task
Select an AI-assisted or analytics-assisted procedure from a file that has been archived. Working only from what is in the file, attempt the experienced-auditor test yourself: can you determine the population and its reconciliation, the exact criteria or configuration applied, the complete output, and the basis for accepting unflagged items? List every element you cannot establish. Then, assuming the tool is unavailable, decide whether the retained artefacts alone would support the conclusion under inspection, and write a short retention checklist for the next engagement covering the gaps you found.
Success looks like
- The test is performed strictly from the file contents, without filling gaps from personal recollection
- Missing elements are listed specifically rather than summarised as documentation could be improved
- The retention checklist names concrete artefacts — extract, configuration, full output, version identifiers
- The basis for accepting unflagged items is explicitly checked, since it is the most commonly absent element
Watch out for
- Relying on firm-level tool approval documentation to answer engagement-level questions about population and criteria
- Assuming a cloud platform will retain a reproducible historical version of the routine and its output
- ISA 500 reliability applies separately to the source data, the processing, and the output. Seeded-error testing is the practical evidence that processing works, and two-directional reperformance — flagged and unflagged items — is how the output is corroborated.
- The ISA 230 experienced-auditor test is the operative standard: could a competent auditor with no connection to the engagement reproduce this and reach the same result from the file alone?
- The file documents the procedure; the firm documents the tool. Vendor validation and firm approval belong in methodology documentation, and they never answer the engagement-specific questions of population, criteria, and evaluation.
- Required explainability scales with reliance. A ranking tool whose every output is examined needs little; a tool relied on directly needs training basis, validation, and known failure modes — and where those cannot be established, the output can direct attention but not carry substantive assurance.
- Assume the tool will be unavailable at inspection. Retain the extract, full configuration, complete output, and version identifiers as archived files, because vendor upgrades routinely make original results irreproducible.