Journal Entry Testing and Fraud Risk
Deliberate Academy Editorial Team
Reviewed for accuracy and professional relevance
You're 4 lessons in — don't lose your progress.
Sign up free to save where you are and earn a verified certificate when you pass.
- Design journal entry testing that addresses the ISA 240 management override risk specifically, rather than generic unusualness
- Explain why criteria-based selection and pattern-based anomaly detection fail differently, and why fraud risk work should not rely on the latter alone
- Build selection criteria from the fraud risks actually identified for the entity rather than from a generic checklist or a tool default
- Handle the false-positive burden without tuning away the characteristics that define the risk
ISA 240 requires the auditor to test journal entries and other adjustments in every audit, because management override of controls is a risk present in every entity and journal entries are the primary mechanism through which it is effected. It is a mandatory procedure with no materiality exemption, and it is the single audit area where full-population analysis has changed practice most.
It is also the area where the difference between criteria-based selection and machine-learning anomaly detection matters most, and where using the wrong one produces work that looks rigorous and addresses the wrong risk.
Fraud Is Not Anomaly
The intuition behind anomaly detection is that fraudulent entries look different from normal ones. Sometimes they do. Frequently they do not, and the reason is structural rather than technical: an entry designed to avoid detection is designed to look ordinary.
A competent management override is posted by someone with legitimate posting rights, in an amount that does not stand out, to accounts that ordinarily receive entries, with a description matching the surrounding entries, during business hours. Every property that makes it fraudulent is contextual: it lacks proper authorisation, it lacks a business purpose, it was made to achieve a reporting outcome. None of those properties is visible in the entry's own attributes.
An unsupervised anomaly model ranks by statistical distance from the population norm. A well-disguised override has, by construction, minimal statistical distance. Meanwhile the top of the anomaly ranking fills with entries that are genuinely unusual and entirely legitimate: the annual impairment, the acquisition accounting, the restructuring provision, the year-end tax true-up. These are the largest and strangest entries in any ledger, they are unusual every year, and they are exactly what the model surfaces.
The result is a procedure that reliably finds the entries you already knew about and reliably misses the entries you are looking for.
ISA 240 journal entry testing addresses management override, a specific risk with specific characteristics. A generic anomaly ranking is not a substitute, because the entries that best conceal override are by design the least statistically anomalous. Criteria derived from how override is actually effected at this entity are the core of the procedure; anomaly scoring is a supplement to it.
Criteria-Based Selection
Criteria-based selection encodes what override looks like as explicit, defensible rules. Because the rules come from the risk rather than from the data, they cannot be trained to accept a misstatement as normal, which is the circularity problem from lesson two.
The standard starting set, each tied to a reason:
- Entries posted by users outside finance, or by users whose role does not ordinarily require posting rights. Override frequently requires someone senior to bypass the normal preparer.
- Entries posted outside normal business hours, at weekends, or on public holidays. Timing that avoids observation.
- Entries to seldom-used accounts, particularly suspense, intercompany, and manual reserve accounts, which absorb entries without drawing attention.
- Round-sum entries above a threshold. Estimates fabricated to achieve an outcome tend to be round; genuine transactional amounts rarely are.
- Entries with missing, generic, or duplicated descriptions, or descriptions inconsistent with the accounts involved.
- Entries posted in the closing days of the period or after period end but dated within it. The window where reporting-outcome adjustments cluster.
- Entries reversed shortly after posting without an obvious operational reason.
- Entries that credit revenue or debit an expense reversal against an unusual counterparty account.
The essential step, and the one most often skipped, is that this list is a starting point and not the procedure. ISA 240 requires the auditor to identify fraud risks for the entity. The criteria must then be extended to reflect those risks. If the identified risk is revenue recognition manipulation at period end, criteria targeting manual entries to revenue accounts in the final week are mandatory, and their absence is a gap regardless of how many generic criteria were run. If the risk relates to an earnout arrangement, criteria targeting the accounts affecting the earnout calculation belong in the set.
A file running the same eight generic criteria on every client has not performed an entity-specific fraud risk procedure. It has run a template.
Where Pattern Detection Genuinely Helps
Criteria-based selection has a real weakness: it only finds what you thought to specify. Pattern-based approaches complement it in three ways worth using deliberately.
Relationship anomalies. Account combinations that never otherwise co-occur are hard to enumerate in advance but easy to detect statistically. An entry debiting an account and crediting another where that pairing appears nowhere else in three years is genuinely interesting.
Behavioural change. A user whose posting behaviour changes sharply — volume, timing, accounts touched, typical amounts — is a signal no static rule captures. The comparison is the user against their own history, not against the population.
Text patterns. Description fields cluster in ways that reveal copied narratives, unusual phrasing, or entries whose stated rationale does not match the accounts.
Run these alongside criteria-based selection, never instead of it, and document them as supplementary.
The False-Positive Problem, Handled Honestly
Criteria-based selection on a large ledger generates substantial volume. A mid-sized entity may return several thousand entries against the standard criteria, most driven by legitimate routine processes: an automated interface posting outside business hours, a shared service centre in another time zone, a monthly accrual that is genuinely round.
The tempting response is to tighten criteria until volume is manageable. Applied to fraud risk work this is particularly damaging, because the characteristics being tuned away are the characteristics that define the risk. Raising the round-sum threshold from 10,000 to 100,000 removes noise and also removes the band where a fabricated accrual is most likely to sit.
The better approach is to explain volume rather than eliminate it, using the root-cause grouping from lesson three, with one addition specific to fraud work: exclusions must be justified by an understood and corroborated process, not by apparent normality.
If 1,400 out-of-hours entries come from a scheduled overnight interface, confirm the interface exists, confirm its schedule with IT, confirm the entries carry its system user ID, and test that the excluded population contains only entries matching that signature. Then exclude the group and document the basis. That is a defensible exclusion. "These looked routine" is not, and the difference is precisely what an inspector will probe.
Two rules protect the procedure. Never exclude a group without corroborating its cause independently of the finance team whose entries you are testing. And always examine a small number of items from every excluded group — if the exclusion logic is wrong, that is where it shows.
An audit team runs an unsupervised anomaly detection model over the full journal entry population and investigates the 25 highest-scoring entries, all of which prove to be legitimate large non-routine transactions. The team concludes that journal entry testing identified no indications of management override. Why is this conclusion poorly supported?
Select one answer.
An override found by a criterion the anomaly model ranked in the bottom half
Context
A services client with a profit-based management bonus had a fraud risk identified around period-end revenue recognition. The engagement team ran both an anomaly detection model across the full journal entry population of 96,000 entries and a criteria-based selection built around the identified risk, including manual entries to revenue accounts in the final ten days of the period.
Action
The anomaly model surfaced 30 entries, all explicable: two impairments, an acquisition adjustment, several large intercompany settlements. The criteria-based selection returned 214 entries, of which 31 were manual revenue postings in the closing window. Working through those, the team found four entries totalling slightly above performance materiality recognising revenue on contracts where the client's own signed acceptance documentation was dated after period end. The entries were individually unremarkable in size, posted by a finance manager with legitimate rights, during business hours, with descriptions consistent with surrounding entries.
Outcome
The anomaly model had ranked all four in the bottom half of the population. They were, by every statistical measure, ordinary entries. The revenue was restated before the accounts were signed, and the matter was reported to those charged with governance. The partner used the engagement to change firm methodology: anomaly scoring may supplement journal entry testing but may not be the primary selection method, and every engagement must document how its selection criteria map to the fraud risks identified for that specific entity.
A team attributes 1,400 out-of-hours journal entries to a scheduled overnight interface and excludes the group from individual investigation. What does this lesson require before that exclusion is defensible?
Select one answer.
Exercise
Your Task
Take the fraud risks identified on a current or recent engagement. For each identified risk, write the specific journal entry selection criteria that address it, and state the accounts, users, timing windows, or amount characteristics involved. Then compare that list against the criteria actually applied in the file. Identify any identified fraud risk with no corresponding criterion — that is a gap in the ISA 240 work. Finally, for any group of entries excluded from investigation, check whether the file records an independently corroborated cause or only an assertion that the entries appeared routine.
Success looks like
- Every identified fraud risk maps to at least one specific selection criterion
- Criteria name concrete accounts, users, windows, or amount patterns rather than restating generic categories
- Exclusions are supported by a corroborated process cause, verified independently of the finance team being tested
- Any gap between identified risks and applied criteria is recorded honestly
Watch out for
- Running the same generic criteria on every client and describing it as entity-specific fraud risk work
- Raising thresholds to reduce volume, removing the band where fabricated entries are most likely to sit
- Excluding a large group on the basis that it looks routine, without corroborating the underlying process
- ISA 240 journal entry testing is mandatory in every audit and addresses management override specifically. Override entries are designed to look ordinary, so they are structurally the least statistically anomalous items in the population.
- Criteria-based selection is the core of the procedure because the criteria derive from the risk rather than from the data, which also makes them immune to the circularity that undermines models trained on the unaudited ledger.
- Generic criteria are a starting point only. Criteria must be extended to address the fraud risks actually identified for the entity, and a file running the same list on every client has run a template rather than a risk-based procedure.
- Pattern-based methods add genuine value for account-combination anomalies, user behavioural change, and description text patterns — as a supplement to criteria-based selection, never as a replacement.
- Manage false positives by explaining volume through corroborated root causes, not by tightening thresholds. Tuning removes exactly the characteristics that define override risk, and every excluded group should still have a few items examined.