Skip to main content
Deliberate AcademyProfessional AI Education
~16 min left
Lesson 8 of 10
16 min read10 XP

Professional Scepticism in an AI-Assisted Audit

Deliberate Academy Editorial Team

Reviewed for accuracy and professional relevance

You're 8 lessons in — don't lose your progress.

Sign up free
What you'll learn
  • Explain how automation bias operates in audit judgment and why seniority and experience do not neutralise it
  • Identify the specific point in an AI-assisted procedure where scepticism is most likely to lapse — the clean result
  • Apply structured techniques that make scepticism a designed feature of the procedure rather than an individual disposition
  • Recognise when fluent AI-generated explanations are substituting for audit evidence

ISA 200 requires the auditor to plan and perform the audit with professional scepticism, recognising that circumstances may exist that cause the financial statements to be materially misstated. It is described as an attitude, which makes it sound like a matter of character. In an AI-assisted audit it is better understood as a design problem, because the specific ways AI erodes scepticism are predictable and can be engineered against.

Automation Bias

Automation bias is the documented tendency to over-trust automated output and to under-weight contradictory evidence from other sources. It has two forms. Omission errors: failing to notice a problem because the automated system did not flag it. Commission errors: acting on automated output that is wrong, in preference to one's own contrary assessment.

Both matter in audit and the first matters more, because an unflagged population is invisible. An auditor who reviews a clean exception report has no cue prompting further thought. There is nothing to look at. The absence of a signal reads as the presence of assurance, and no discipline is triggered.

Three features of AI-assisted audit work make this worse than for older automation.

Fluency. An AI system that explains its output in confident, well-formed professional language triggers the heuristics we use to assess human competence. Those heuristics are useless here: fluency is uncorrelated with accuracy in a language model, and a well-written wrong explanation is the normal failure mode rather than an unusual one.

Opacity. Where the reasoning cannot be inspected, there is nothing to disagree with. Scepticism needs a claim to interrogate, and a score of 0.87 offers none.

Apparent comprehensiveness. "All 400,000 entries were analysed" is a powerful reassurance signal that speaks to coverage while saying nothing about appropriateness — the exact conflation from lesson three.

Seniority does not protect against this. Experienced auditors are, if anything, more exposed, because they have more accumulated experience of automated tools being broadly right and greater confidence in their ability to spot problems. Automation bias is a feature of how attention works, not a deficiency of the individual.

Critical

The riskiest moment in an AI-assisted procedure is the clean result. An exception report with no exceptions produces no cue to think, no item to examine, and no natural point of challenge. Every design in this lesson exists to create the challenge the clean result does not prompt on its own.

Designing Scepticism Into the Procedure

Because scepticism cannot be relied on to appear spontaneously at the point of greatest risk, it must be built into the procedure as required steps.

Predict before you look. Before running a routine, record what you expect: roughly how many exceptions, of what type, concentrated where. Then compare. A material difference between expectation and result is informative in both directions — far fewer exceptions than expected is at least as interesting as far more, and without a recorded prediction the mind quietly adjusts its expectation to whatever appeared. This costs two minutes and is the single highest-value habit in this lesson.

Test the negative. Reperform a sample of unflagged items by hand, as in lesson five. This is the only reliable defence against omission errors, because it examines the population where automation bias makes you blind.

Seed known errors. Introduce realistic misstatements at performance materiality and confirm the routine catches them. This converts trust in the tool from an assumption into evidence.

Require the counter-explanation. For any significant AI-supported conclusion, have someone articulate how the output could be wrong: what would have to be true about the data, the criteria, or the population for this clean result to be misleading? Assigning this explicitly to a named person makes it a task rather than an aspiration.

Separate the finding from the explanation. Where an AI tool both identifies an item and explains it, the explanation is not evidence. Evaluate the item against source documentation before reading the tool's rationale, so the rationale cannot anchor the assessment.

Fluent Explanation Is Not Evidence

That last point generalises into the most practically important habit in an AI-assisted audit.

Generative tools increasingly explain their output: why an entry was flagged, why a balance movement is consistent with a business change, why a contract clause implies a particular treatment. These explanations are frequently correct and always plausible, because plausibility is what the model optimises for. Their persuasiveness is entirely independent of their accuracy.

The audit consequence is that an explanation from a tool has the status of an unverified assertion — the same status as an explanation from the client. Management telling you a receivables increase reflects a large contract signed in November is a starting point for testing, not a conclusion. A tool telling you the same thing is precisely as unverified, and rather more convincing in tone.

The habit that follows: when reading an AI-generated explanation, identify the specific evidence that would corroborate it, then obtain that evidence. If a movement is attributed to a November contract, look at the contract and trace the revenue. The explanation is useful because it directs you to the evidence quickly. It is not a substitute for it, and a file that records the explanation without the corroboration has documented a hypothesis as a finding.

Knowledge check

An AI tool analyses a client's gross margin decline and produces a well-written explanation attributing it to a documented change in supplier pricing and a shift in product mix. The explanation is consistent with what management has said and is internally coherent. What is the appropriate audit response?

Select one answer.

A clean full-population result accepted for three years, undone by a two-minute prediction habit

Audit Manager, mid-tier firm

Context

A team ran an automated completeness test over a client's revenue population annually, comparing recorded revenue against despatch records. The routine had returned zero exceptions in each of three consecutive years. Each year the result was documented and substantive testing reduced on the strength of it. No one had reperformed unflagged items, on the reasoning that there was nothing to reperform.

Action

A new manager introduced the predict-before-you-look step across the portfolio. Asked what she expected, the senior on the engagement said she would expect a handful of exceptions in any population of that size, given normal despatch timing differences around period end, and that three consecutive years of exactly zero was implausible rather than reassuring. On investigation, the matching logic joined the two datasets on a despatch reference field that had been populated for only part of the despatch population since a system change in the first of the three years. Unmatched records were being dropped from the comparison rather than reported as exceptions.

Outcome

The routine had been comparing a shrinking subset against itself and could not have produced an exception. Rebuilt to report unmatched records explicitly, it identified 63 exceptions in the current year, of which 9 were genuine cut-off errors aggregating just below performance materiality. The firm added two requirements: any routine returning zero exceptions must have a sample of unflagged items reperformed, and any join must report unmatched records as exceptions rather than dropping them. The manager noted the failure had survived three annual reviews because a clean result never prompted anyone to ask a question.

Team Dynamics and the Reviewer

Automation bias operates on reviewers too, and often more strongly, because the reviewer is further from the data.

A manager reviewing a working paper that records a sophisticated full-population procedure with a clean result faces a genuine disincentive to challenge. Challenging requires understanding the tool well enough to formulate a specific question, which takes time the review does not have and knowledge the reviewer may not possess. The path of least resistance is to accept the result, and the professional cost of doing so is invisible.

Two structural remedies help. Standard review questions for AI-assisted procedures, asked every time, so challenge does not depend on the reviewer generating it: what was the population and how was completeness established, what precision does this expectation have relative to performance materiality, and what is the basis for accepting unflagged items. Because they are standard, they are not read as scepticism about the individual.

Explicit permission to escalate uncertainty. Junior team members frequently notice that a result seems wrong and do not raise it, because the tool is the firm's approved platform and they cannot articulate a technical objection. A team culture where "this doesn't look right and I can't say why" is a legitimate contribution catches problems that no checklist will. That intuition is often the earliest available signal, and it is discarded by default in most teams.

Quick check

Three consecutive years of exactly zero exceptions from a revenue completeness routine went unchallenged until a new manager introduced the predict-before-you-look step. What did that step change?

Select one answer.

Exercise

~25 min

Your Task

Before the next AI-assisted or analytics procedure you run, write down three predictions: approximately how many exceptions you expect, what types they will be, and where in the population they will concentrate. Run the procedure and compare. Then, regardless of the result — and particularly if it is clean — reperform ten unflagged items by hand against source documentation. Finally, write two sentences answering: what would have to be true about the data, criteria, or population for this result to be misleading? Record all of it in the working paper.

Success looks like

  • Predictions are recorded before the procedure runs, not reconstructed afterwards
  • Unflagged items are reperformed even when — especially when — the result is clean
  • The counter-explanation identifies a specific mechanism by which the result could be wrong, not a generic acknowledgement of limitations
  • A result materially better than predicted is treated as a prompt to investigate rather than as reassurance

Watch out for

  • Treating a fluent AI-generated explanation as corroboration rather than as an unverified assertion to be tested
  • Skipping reperformance of unflagged items on the basis that a clean report contains nothing to reperform
Key takeaways
  • Automation bias produces omission errors — missing what the tool did not flag — and commission errors. Omission is the more dangerous in audit because an unflagged population generates no cue to think, and seniority does not protect against it.
  • Fluency, opacity, and apparent comprehensiveness make AI-assisted work more prone to over-trust than older automation. A well-written wrong explanation is the normal failure mode of a language model, not an unusual one.
  • Scepticism must be designed into the procedure: predict expected results before running, reperform unflagged items, seed known errors, require a named counter-explanation, and evaluate flagged items against source documents before reading the tool rationale.
  • An AI-generated explanation has the same status as a management representation — an unverified assertion. Use it to direct testing, then obtain the corroborating evidence; recording the explanation alone documents a hypothesis as a finding.
  • Reviewers are also subject to automation bias and face a real disincentive to challenge sophisticated clean results. Standard review questions and explicit permission for juniors to raise unarticulated doubts are the structural defences.