What Strategic Sourcing Asks of AI That Tactical Buying Does Not
Deliberate Academy Editorial Team
Reviewed for accuracy and professional relevance
- Distinguish decisions where an AI error is recoverable from decisions where it is structural, and apply a proportionate evidence standard to each
- Interrogate a composite supplier risk score: what it contains, what it omits, and what weighting was applied
- Recognise the tier-one data horizon that limits almost every commercial supplier risk platform
- Explain why a supplier scoring decision now requires an evidence trail rather than an outcome
Tactical procurement decisions are mostly reversible. A poorly drafted RFP can be reissued. A weak negotiation position costs margin on one contract. If AI produced a mediocre supplier shortlist, you notice within a quarter and correct it.
Strategic sourcing decisions are not like that. Awarding a three-year category contract, consolidating from six suppliers to two, or approving a supplier into a regulated supply chain creates commitments that are expensive to unwind and, increasingly, exposures that are legally attributable to you. The same AI tool used with the same care produces very different consequences in the two contexts, and the evidence standard has to reflect that.
The Recoverability Test
Before relying on an AI-derived input, ask what happens if it is wrong.
Recoverable. A supplier research summary that misses a competitor, a spend classification that misallocates a small category, a draft market analysis with an out-of-date figure. Errors surface quickly through normal operation and cost is bounded. Ordinary review is proportionate.
Structural. A supplier scored as low-risk who is financially distressed and fails mid-contract. A category strategy built on a should-cost model with the wrong commodity input. A supplier approved into a regulated supply chain whose tier-two subcontractor triggers a forced labour finding. These do not surface through normal operation — they surface through failure, and by then the commitment has been made and, in the due diligence cases, the liability has already attached.
The distinction is not the sophistication of the analysis. It is whether the error announces itself. Structural decisions need documented inputs, understood limitations, and a record of what was checked, because the only opportunity to catch the error is before the decision.
In supply chain due diligence the decision and the liability are the same event. You cannot discover a forced labour problem in tier three and treat it as a supplier performance issue to be managed — under CSDDD, the Modern Slavery Act and UFLPA, what matters is what you did to look before you contracted, and whether you can evidence it.
Reading a Composite Risk Score
Supplier risk platforms present a single number. That number is a weighted blend of components, and the two questions that matter are almost never asked: which components, and weighted how?
A typical composite includes some subset of financial stability, delivery and quality performance, cyber posture, geographic risk, compliance flags, and ESG indicators. Vendors differ substantially in what they include, and a score of 82 from one platform is not comparable to 82 from another. More importantly, a high score on a composite dominated by financial and delivery data says nothing about labour practices — which is exactly the failure in the scenario that opens this course.
Three questions to put to any supplier risk score before relying on it:
What is in it? Get the component list. If the vendor will not disclose it, that is itself a finding, and it means the score cannot support a due diligence claim.
How is it weighted? A composite where financial data is 60 percent of the weight will rate a profitable supplier well regardless of its human rights profile. Weightings encode a view about what matters, and that view is the vendor's, not yours. Where the platform allows reweighting, align it to your actual category risks.
What is the data horizon? Covered next, and the most consequential of the three.
Where a category's principal risk is human rights or sanctions exposure, a general composite score should not be the primary instrument at all. Use the specific screening covered in lessons six and seven, and treat the composite as background.
The Tier-One Horizon
Almost every commercial supplier risk dataset is built on entities you can identify: registered companies, with filings, credit data, trade records, and news coverage. That works well for your direct suppliers. It degrades sharply below them.
Your tier-one supplier knows its own suppliers. It may not know theirs. By tier three, the relationships are commonly undocumented outside the parties themselves, and no external dataset reliably reconstructs them. Platforms that offer multi-tier visibility typically infer it — from shipping records, customs data, corporate ownership graphs, or industry-standard bills of materials — and inference is not the same as knowledge.
This matters because the regimes in lesson six do not stop at tier one. Obligations extend along the chain of activities, and a company that assessed only its direct suppliers has not met the standard, however well it assessed them.
The practical consequence is that tier-two and beyond requires a different method: contractual flow-down requiring suppliers to disclose their own suppliers for defined high-risk categories, supplier self-declaration, and targeted verification. AI is genuinely useful in processing the disclosures once you have them and in prioritising where to push. It cannot substitute for asking.
A sourcing team relies on a supplier risk platform that returns a composite score of 84 for a new supplier in a category with known forced labour risk. What is the most significant problem with treating this as adequate due diligence?
Select one answer.
Evidence, Not Outcome
The older procurement standard was outcome-based: choose a good supplier and the process is validated by the supplier performing well. Due diligence regimes have shifted the standard to process. What is assessed is whether you took reasonable and appropriate steps, identified risks, acted on what you found, and can demonstrate all three.
That shift has a direct implication for AI use. A decision supported by a tool needs a record of what the tool was asked, what it returned, what its known limitations were, and what the human did with it. "The platform rated them low risk" is an outcome. "We screened against sanctions and adverse media, we required tier-two disclosure for the three highest-risk components, we received disclosure for two of three and escalated the third, and we recorded the residual gap" is a process, and it is defensible even if something later goes wrong.
This is also the honest reason to keep the record: the process standard means you can do everything right and still have a supplier fail. What you cannot do is be unable to show what you did.
A green score, a missing component, and a supplier that had never been screened for what mattered
Context
A category manager sourcing electronic components consolidated to a supplier that had scored 86 on the company's third-party risk platform. The category was known to carry forced labour risk in upstream mineral processing, and the company had a Modern Slavery Act statement committing it to supply chain due diligence. The 86 was recorded in the approval file as the evidence of assessment.
Action
During an internal audit of the procurement function twelve months later, the auditor asked for the component breakdown behind the 86. The platform's composite was 45 percent financial stability, 30 percent delivery and quality performance, 15 percent cyber posture, and 10 percent geographic risk. There was no labour or human rights component in the score at all. The platform offered a separate human rights module that the company had not licensed, and the supplier had never been screened against it.
Outcome
Screening the supplier retrospectively produced no adverse findings at tier one, but the exercise established that the company had no visibility below tier one for the category and no contractual right to obtain it. The team added tier-two disclosure obligations to the standard terms for high-risk categories at next renewal, licensed the human rights module, and changed the approval template so that the risk evidence field requires the specific screening performed rather than a composite score. The auditor's finding was not that the supplier was bad — it was that the company could not demonstrate it had looked for the risk it had publicly committed to looking for.
This lesson's recoverability test separates a recoverable AI error from a structural one. What is the criterion it actually uses?
Select one answer.
Exercise
Your Task
Take a supplier risk score your organisation currently relies on. Obtain the component breakdown and the weightings from the vendor or the platform documentation, and write them down. Then answer three questions: (1) does the score contain any component measuring the principal risk in this category, (2) what is the data horizon — does the underlying dataset cover below tier one, and how is any multi-tier view derived, and (3) if you had to demonstrate to an auditor what you assessed for this supplier, what would you produce beyond the score itself? Record any question the vendor cannot answer, and treat that as a finding rather than a gap in your notes.
Success looks like
- The component breakdown and weightings are obtained as specifics rather than described in general terms
- The principal risk for the category is named first, then checked against the score components — not the other way round
- The data horizon question distinguishes observed tier-one data from inferred multi-tier data
- A vendor unable to disclose components is recorded as a finding affecting what the score can support
Watch out for
- Comparing scores across platforms as though the numbers are equivalent
- Accepting a vendor multi-tier visibility claim without establishing whether the relationships are observed or inferred
- Apply the recoverability test: an error is recoverable if it announces itself through normal operation, and structural if it only surfaces through failure. Structural decisions need documented inputs and limitations because the decision is the only chance to catch the error.
- A composite supplier risk score is a weighted blend, and the weighting encodes the vendor view of what matters. Establish the components and weights before relying on it, and never let a general composite stand in for screening the category actual principal risk.
- Commercial supplier datasets are built on identifiable registered entities and degrade sharply below tier one. Multi-tier visibility is usually inferred from shipping, customs, or ownership data, and inference is not knowledge.
- Due diligence regimes assess process, not outcome. A defensible record states what was screened, what was found, what was escalated, and what gaps remain — an outcome such as rated low risk is not evidence of having looked.
- In supply chain due diligence the sourcing decision and the legal exposure are the same event, so the evidence has to be assembled before the award rather than reconstructed after an allegation.