Skip to main content
Deliberate AcademyProfessional AI Education
~17 min left
Lesson 9 of 10
17 min read10 XP

Governance and Defensible Sourcing Decisions

Deliberate Academy Editorial Team

Reviewed for accuracy and professional relevance

You're 9 lessons in — don't lose your progress.

Sign up free
What you'll learn
  • Determine which regulatory obligations actually attach to AI-assisted supplier scoring, and avoid the common overstatement about the EU AI Act
  • Identify the structural biases that make supplier scoring systematically disadvantage smaller and foreign suppliers
  • Build a decision record that lets a rejected supplier be answered and a regulator be satisfied
  • Assign ownership for model and threshold changes so scoring logic cannot drift without approval

Everything in this course produces inputs to a decision: award, reject, escalate, monitor, exit. This lesson is about making those decisions defensible — to an internal auditor, to a rejected supplier who asks why, and to a regulator examining a due diligence claim.

What Actually Applies

It is worth being precise here, because supplier scoring attracts a lot of imprecise regulatory commentary.

The EU AI Act. General business-to-business supplier risk scoring does not typically fall within the Annex III high-risk categories. Those cover areas such as employment and worker management, creditworthiness assessment of natural persons, education, essential public services, law enforcement, and migration. A system scoring corporate suppliers on financial and delivery risk generally sits outside that list. Two qualifications matter: if your suppliers are individuals or sole traders, a creditworthiness assessment may bring it into scope, and general-purpose AI obligations and transparency provisions can still apply. Claiming your supplier scoring is high-risk when it is not is as much a governance failure as missing an obligation that does apply — it misdirects effort and misstates your position.

GDPR. This one frequently does apply and is frequently overlooked. Supplier due diligence routinely processes personal data: directors, beneficial owners, adverse media about named individuals, sanctions matches against people. You need a lawful basis, the data must be accurate, and individuals have rights including access and rectification. An adverse media hit about a named director held indefinitely, unverified, and acted upon is a data protection problem as well as a fairness one.

Public procurement rules. Where you are a contracting authority, duties of transparency, equal treatment, and proportionality apply to your evaluation process regardless of whether AI is involved. A scoring system whose logic cannot be explained is difficult to reconcile with an obligation to give reasons for an award decision.

Contract and tender terms. Your own tender documentation frequently commits you to an evaluation methodology. Introducing an AI-derived score not disclosed in that methodology is a straightforward breach of your own stated process, and it is the most likely basis for a supplier challenge.

Warning

The most common governance failure in this area is not a missed regulatory obligation. It is a scoring input that was never disclosed in the tender methodology. If your published evaluation criteria do not mention a supplier risk score, using one to eliminate bidders is a challengeable departure from your own stated process.

Structural Bias in Supplier Scoring

Supplier scoring models exhibit predictable biases, and they run in a consistent direction: against smaller, newer, and foreign suppliers. These are not exotic edge cases, they are the ordinary behaviour of models built on the available data.

Data availability bias. Larger suppliers file more, are covered more, and appear in more datasets. A model that penalises missing data penalises smallness rather than risk. A twelve-year-old family business with abbreviated accounts and no press coverage is not riskier than a listed company; it is quieter.

Jurisdictional bias. Suppliers in countries with strong registry disclosure produce richer, more current data. Where disclosure is limited, the model sees less and, again, often scores it worse. The effect penalises the jurisdiction rather than the supplier.

Adverse media asymmetry. Media coverage is far denser in some languages and countries than others. A supplier operating where local media is well indexed accumulates more hits, including trivial ones, than an equivalent supplier where coverage is thin — so measured adverse media partly reflects press density.

Historical performance bias. Models trained on your own past supplier performance learn from the suppliers you selected. Suppliers similar to those you have historically favoured score well, which reproduces past selection patterns and works against new entrants and diverse suppliers regardless of merit.

The mitigation is not to abandon scoring. It is to separate absence of data from evidence of risk — these should be distinct outcomes with different responses, not blended into one lower score — and to test the model's score distribution against supplier size, country, and tenure. If small suppliers systematically score worse, establish whether that reflects risk or reflects data availability. Where it is availability, the correct response is proportionate diligence, not exclusion.

The Decision Record

A supplier decision record has three audiences, and a record that satisfies all three is not much longer than one that satisfies none.

The rejected supplier. Can you explain why they were not selected, in terms of the criteria you published? A supplier told they scored poorly on a proprietary composite has been given a number, not a reason, and is likely to challenge.

The internal reviewer. Can they see what evidence was considered, what was screened, what was found, and what judgment was applied on top of the analytics?

The regulator or auditor. Can they see that the steps the applicable regime requires were actually taken, and taken before the decision?

The record that serves all three states: the criteria applied and their source; the analytics used, with the specific tool and what it covered; the screening performed and its results, including clear dispositions; the material gaps and what was done about them; the human judgment applied and by whom; and the date. Crucially it captures the data as it stood at decision time, for the same reason as lesson seven — scores, ownership data, and lists all change, and a reviewer needs to see what you saw.

The failure mode is a record that captures the conclusion and not the basis. "Supplier approved, risk score 84" is a conclusion. It cannot answer any of the three audiences.

Change Control on Scoring Logic

A supplier scoring model is a decision rule, and decision rules need change control.

In practice thresholds and weightings get adjusted informally: a category manager finds the model too conservative, a threshold is nudged, alert volume becomes manageable, and nobody records it. Six months later nobody can say what logic produced a decision made in the interim, which makes every decision in that period unreviewable.

Four controls, none expensive:

  • Named ownership of the scoring logic, distinct from the team using it day to day.
  • Approval and a record for any change to weightings, thresholds, or data sources, with the reason.
  • Versioning, so a decision can be tied to the logic that produced it.
  • Periodic revalidation, checking that scores still relate to outcomes — that suppliers scored high-risk actually experience more issues than those scored low-risk. Without that check a model can drift into producing well-formatted noise indefinitely, since the score never announces that it has stopped working.

That last control is the one almost nobody runs, and it is the only one that would detect a scoring model that has quietly become uninformative.

Knowledge check

A procurement team plans to eliminate bidders scoring below a threshold on an AI-derived supplier risk score. The published tender evaluation methodology lists price, technical capability, and delivery capability, and does not mention a risk score. What is the principal problem?

Select one answer.

A scoring model that penalised smallness, found by testing the score distribution

Procurement Director, public sector body

Context

A public body introduced an AI-assisted supplier risk score into its pre-qualification process, using it to flag suppliers for enhanced diligence. Over eighteen months the proportion of contracts awarded to SMEs fell noticeably, which conflicted with a published commitment to SME participation. No individual decision looked wrong on review.

Action

The procurement director asked for the score distribution to be analysed against supplier size, country of registration, and years trading, which had never been done. Small suppliers scored materially worse on average, and decomposing the score showed why: the model treated missing data as a negative indicator, and small suppliers filing abbreviated accounts with no press coverage and no credit history triggered several missing-data penalties at once. The score was measuring how much was known about a supplier, not how risky it was.

Outcome

The body separated the two outcomes. Missing data now produces a distinct data-insufficient flag routing to proportionate manual diligence — a short financial questionnaire and two references — rather than a lower risk score. Suppliers with genuine adverse findings continue to be flagged as risk. SME award share recovered over the following year and the enhanced diligence workload rose only slightly, because most data-insufficient cases cleared quickly. The director noted that every individual decision had been defensible and the aggregate pattern had not, which is only visible if someone looks at the distribution.

Quick check

Of the four controls this lesson puts on scoring logic, which is the only one that would catch a model that has quietly stopped being informative?

Select one answer.

Exercise

~30 min

Your Task

Take your supplier scoring or risk-rating process. First, check disclosure: compare the criteria you actually apply against the evaluation methodology published in your tender documentation, and identify anything applied but not disclosed. Second, test for structural bias: pull the score distribution against supplier size, country of registration, and years trading, and determine whether smaller or foreign suppliers score systematically worse. Third, establish whether missing data and adverse findings produce the same outcome in your process, and separate them if they do. Finally, document who owns the scoring logic, when it last changed, whether that change was approved and recorded, and when the model was last revalidated against actual supplier outcomes.

Success looks like

  • Applied criteria are compared against published methodology, with undisclosed criteria identified
  • Score distribution is tested by supplier size, jurisdiction, and tenure rather than assessed decision by decision
  • Absence of data and evidence of risk are established as distinct outcomes with different responses
  • Named ownership, change history, and last revalidation date for the scoring logic are all identified — or their absence recorded as a gap

Watch out for

  • Overstating regulatory scope by treating general B2B supplier scoring as EU AI Act high-risk, which misdirects effort
  • Overlooking GDPR obligations on personal data about directors and beneficial owners held in due diligence records
Key takeaways
  • General B2B supplier risk scoring typically falls outside EU AI Act Annex III high-risk categories, though sole traders and general-purpose AI provisions can change that. Overstating scope is a governance failure too. GDPR, by contrast, usually does apply, because due diligence processes personal data about directors and beneficial owners.
  • The most common practical exposure is procedural: applying an AI-derived criterion that the published tender methodology does not disclose. Either disclose it or use it for diligence rather than elimination.
  • Supplier scoring is structurally biased against small, new, and foreign suppliers through data availability, jurisdictional disclosure differences, adverse media density, and models trained on past selections. Separate absence of data from evidence of risk and route the former to proportionate diligence.
  • A decision record must serve the rejected supplier, the internal reviewer, and the regulator: criteria and source, analytics used and coverage, screening and dispositions, gaps and treatment, human judgment and who applied it, and the data as it stood at decision time.
  • Scoring logic needs named ownership, recorded and approved changes, versioning, and periodic revalidation against actual outcomes — the last of which is rarely done and is the only control that detects a model that has quietly stopped being informative.