Skip to main content
Deliberate AcademyProfessional AI Education
~20 min left
Lesson 7 of 9
20 min read10 XP

AI Auditing — Documentation, Evidence, and Audit Trails

Deliberate Academy Editorial Team

Reviewed for accuracy and professional relevance

You're 7 lessons in — don't lose your progress.

Sign up free
What you'll learn
  • Distinguish a control that is documented from a control that is evidenced as operating, and explain why auditors specifically test for the second
  • Identify the core categories of audit evidence an AI governance program must maintain across classification, documentation, logging, and human oversight
  • Design a log or evidence record that would actually satisfy an auditor's sampling request, not just a narrative description of a process
  • Build a periodic internal audit sampling practice that surfaces gaps before an external audit or regulatory inquiry does

Every framework covered so far in this course — risk classification, ISO/IEC 42001's management system, an internal policy hierarchy — eventually gets tested the same way: someone asks for proof. An external ISO/IEC 42001 auditor asks for evidence a control operated, not just a policy describing it. A regulator conducting a supervisory review asks for the technical file and the human oversight logs for a specific high-risk system, not a general description of your governance philosophy. A plaintiff's counsel in litigation asks whether a specific automated decision about a specific person was actually reviewed by a human, and when. In every case, the question is the same: can you produce a specific, dated, attributable record — not can you describe your process convincingly.

Designed Versus Operating: The Distinction That Decides Most Audits

Auditors — internal, external certification, or regulatory — routinely draw a distinction between a control that is designed (a documented process exists, on paper, that should prevent or catch a specific problem) and a control that is operating effectively (there is evidence, across a meaningful sample of actual instances, that the process was actually followed). A beautifully written human oversight procedure is a design. A log showing that, across the last quarter's automated loan-decision outputs, a sample of borderline cases was actually reviewed by a named underwriter within a defined time window, with a record of whether the AI's recommendation was accepted or overridden, is operating evidence.

Most AI governance programs that fail an audit do not fail because their policies were badly written. They fail because the policies described a control that, when the auditor asked for evidence, turned out to exist only on paper. This is the single most important practical lesson in this course: write your controls assuming from day one that someone will ask for the evidence, and design the evidence-generation into the process itself rather than treating documentation as something to reconstruct after the fact.

Note

Specific evidence-retention periods and audit sampling requirements vary by regulation, sector, and jurisdiction, and continue to be clarified through regulatory guidance. This lesson describes evidence categories and practices structurally. Confirm your organization's specific retention obligations with legal counsel and any applicable sector regulator — do not treat a generic retention period as a substitute for your actual obligations.

The Core Categories of Audit Evidence

A reasonably complete AI governance evidence base covers several categories, each answering a specific auditor question.

Classification records. For every AI system in the inventory, a dated record of the risk classification decision — the decision-tree answers, the five-factor score if applicable, the resulting tier, and who approved it. Answers: how did you decide this system needed the level of governance it received?

Technical documentation. The technical file for each high-risk system — intended purpose, design and data overview, known limitations, and performance characteristics. Answers: what is this system, and what do you know about how well it works?

Data governance evidence. Records of the data quality and bias assessment conducted before deployment, and evidence of any subsequent monitoring for data or performance drift. Answers: did you check whether the data this system relies on was fit for purpose, and are you still checking?

Operational logs. Records generated by the system's actual operation — for a high-risk automated-decision system, this typically means a log entry per decision instance, sufficient to reconstruct what input the system received, what output it produced, and what happened next. Answers: can you reconstruct what this system actually did, for a specific case, months later?

Human oversight records. Evidence that the human review step required for a given risk tier actually occurred and was substantive — who reviewed, when, how long the review took, and whether the AI's recommendation was accepted, modified, or overridden, sampled at a defined frequency rather than logged for every instance if volume makes that impractical. Answers: was the human oversight control real, or nominal?

Incident records. A log of AI-related incidents — errors, near-misses, complaints, or unexpected behavior — with the investigation outcome and any corrective action taken. Answers: when something went wrong, did you catch it, and what did you do about it?

Training and competence records. Evidence that staff operating or overseeing AI systems received the training their role requires. Answers: did the people responsible for this control actually know how to execute it?

Vendor due diligence records. Evidence of the risk assessment conducted on any third-party AI vendor, covered in depth in the next lesson. Answers: did you assess the AI risk you inherited from a vendor, not just the AI risk you created yourself?

Human oversight log entry

Before

Human review conducted per policy.

This entry describes a policy, not an event. It cannot answer who reviewed, when, what was reviewed, or what the outcome was — it provides no auditable evidence at all.

After

2026-03-14, 09:41 UTC. Application ID 88213. Reviewer: J. Alavi, Senior Underwriter. AI recommendation: Decline. Review duration: 4 min 12 sec. Outcome: Recommendation upheld. Reviewer note: Income documentation consistent with AI's stated basis for decline; no override warranted.

This entry names a specific reviewer, a specific case, a review duration consistent with genuine review rather than a rubber stamp, and a substantive note — the kind of record that satisfies an auditor's sampling request.

A Stage-Two Certification Audit Finds a Documentation-Reality Gap — Healthcare Technology Vendor

Compliance Manager, health technology vendor providing an AI-assisted clinical triage support tool

Context

A health technology vendor pursuing ISO/IEC 42001 certification passed its stage-one documentation review without issue — the auditor confirmed the required policies, standards, and procedures existed and were internally consistent. At the stage-two audit, focused on operating evidence, the auditor requested a sample of 25 human oversight records from the preceding quarter for the company's clinical triage tool, which the company's human oversight procedure described as requiring a clinician to review every AI-flagged high-acuity case within 15 minutes.

Action

The compliance team could produce complete records for only 14 of the 25 sampled cases; the remaining 11 had been reviewed, according to clinical staff interviews, but the review had not been logged in the system designed to capture it, because clinicians had been using a workaround during a period when the primary logging tool had intermittent connectivity issues. The auditor raised this as a major nonconformity — not because the reviews likely had not happened, but because the organization could not produce evidence that they had. The compliance team's remediation plan added a redundant, offline-capable logging fallback and a daily reconciliation check comparing flagged cases against logged reviews, closing any gap within 24 hours rather than allowing it to accumulate silently for a quarter.

Outcome

The nonconformity was closed within the certification body's 90-day window through a documented corrective action and a follow-up evidence sample showing 100% logging completeness across the subsequent eight weeks, and certification was granted. The Compliance Manager's assessment afterward was that the underlying human oversight control had very likely been operating throughout — the failure was entirely in evidence capture, not in the substance of the review — but from an audit standpoint, an unevidenced control and a non-operating control are treated identically, which is the lesson the team took most seriously going forward.

Knowledge check

A company's AI incident response procedure states that all AI-related incidents are investigated and corrective action is tracked to closure. During an internal audit, the auditor asks to see the incident log for the past year and finds three entries, all closed the same day they were opened with no investigation notes. The compliance team explains that the process is followed 'informally' most of the time and not everything gets logged. What is the correct audit conclusion?

Select one answer.

Building an Audit-Ready Evidence Repository

Two practices meaningfully improve audit readiness beyond simply generating the individual records described above. First, maintain a control-to-evidence map — a single reference document listing each governance control your framework claims to have, and exactly which artifact, in which system, constitutes the evidence for it. When an audit request comes in, the compliance team should be able to answer "where is the evidence for X" in minutes, not days of searching across scattered systems and inboxes. Second, run your own internal audit sampling on a fixed cadence — quarterly for Critical and High tier systems, at minimum — pulling the same kind of sample an external auditor would, and treating any gap you find the same way the healthcare vendor in the case study treated theirs: as a real nonconformity requiring a documented corrective action, not an informal fix that goes unrecorded.

Tip

A useful internal test: pick one high-risk or Critical-tier AI system at random and ask your team to produce, within one business day and without advance notice, the full evidence set an external auditor would request for it — classification record, technical file, a sample of operational logs, and human oversight records for a specific two-week period. If your team cannot assemble this within a day, an external auditor or regulator will find the same gap, on their schedule rather than yours.

Exercise

~20 min

Your Task

Choose one control from your organization's AI governance framework (or a plausible one) — for example, human oversight for a specific system, or vendor due diligence. Design the evidence record that would satisfy an auditor's sampling request for that control: specify the exact fields the record must capture, the system or location where it will be stored, who is responsible for creating each entry, and how frequently your team will internally audit a sample of these records to confirm the control is operating before an external party asks.

Success looks like

  • The evidence record includes a timestamp, a named responsible party, and a specific, substantive outcome field — not a generic "compliant" checkbox
  • The record's storage location and creation responsibility are both named specifically, not left as "TBD" or "the relevant team"
  • An internal audit sampling cadence is specified with a frequency, not left open-ended

Watch out for

  • Designing an evidence record that captures only that an event occurred, without capturing enough detail to demonstrate the review was substantive (compare the before/after log entry example in this lesson)
  • Assuming a control is evidenced simply because a procedure describing it exists — that is exactly the designed-versus-operating gap this lesson addresses
  • Skipping the internal audit sampling step, which is what allows a compliance team to find and fix gaps on its own schedule rather than an external auditor's

Hint

Write the evidence record fields first, then ask: if I sampled ten of these six months from now, could I reconstruct exactly what happened in each case without needing to ask the person who created the record? If not, add fields until the answer is yes.

Quick check

Why does this lesson recommend running internal audit sampling on a fixed cadence, rather than only preparing evidence when an external audit or regulatory inquiry is announced?

Select one answer.

Key takeaways
  • Auditors distinguish between a control that is designed (documented on paper) and a control that is operating effectively (evidenced across a real sample) — most AI governance audit failures are evidence gaps, not design failures.
  • Maintain evidence across eight core categories: classification records, technical documentation, data governance evidence, operational logs, human oversight records, incident records, training records, and vendor due diligence records.
  • Design evidence capture into the process itself, with specific, substantive fields — a log entry that only confirms an event occurred, without capturing who, when, and what was actually reviewed, will not satisfy an auditor's sampling request.
  • Maintain a control-to-evidence map so any audit request can be answered in minutes, and run internal audit sampling on a fixed cadence — quarterly at minimum for your highest-tier systems — to find and correct gaps before an external party does.
  • Treat an unevidenced control the same way an external auditor will: as a nonconformity requiring a documented corrective action, regardless of how confident your team is that the underlying process was actually followed.