Skip to main content
Deliberate AcademyProfessional AI Education
~16 min left
Lesson 4 of 10
16 min read10 XP

Patient Data, Privacy, and GDPR Compliance in AI Use

Deliberate Academy Editorial Team

Reviewed for accuracy and professional relevance

You're 4 lessons in — don't lose your progress.

Sign up free
What you'll learn
  • Identify the legal basis required to process health data as special category data under UK GDPR and apply it to AI tool use scenarios
  • Distinguish between general-purpose AI tools and approved clinical AI platforms in terms of GDPR compliance for patient data processing
  • Evaluate whether a proposed data processing activity with an AI tool requires a Data Protection Impact Assessment
  • Apply the principle of data minimisation to reduce the data protection risk footprint of AI-assisted clinical workflows
  • Demonstrate the correct organizational process for approving an AI tool that processes patient data before clinical use

Patient data is the most sensitive category of personal data under UK and EU data protection law. It is also the data that healthcare professionals handle every day and that AI tools need to function. The tension between these two facts is not theoretical: it has already produced enforcement action against healthcare organizations and created professional conduct questions for individual clinicians who used AI tools in ways that compromised patient confidentiality. This lesson gives you the practical framework for knowing what data you can use with which tools and what that distinction requires of you professionally.

Health Data as Special Category Data Under GDPR

The UK GDPR, which retained the EU GDPR framework post-Brexit via the Data Protection Act 2018, creates a two-tier system of personal data protection. Most personal data requires a lawful basis for processing. Health data, which includes any information relating to the physical or mental health of an individual, is classified as special category data and requires both a lawful basis and a specific condition under Article 9 of the UK GDPR.

The conditions available for processing health data in healthcare settings are primarily: the data subject's explicit consent, processing that is necessary for the provision of health or social care, and processing for reasons of public interest in the area of public health. In routine clinical practice, healthcare treatment and care under the care of health professionals is the most common basis. The key phrase is "under the care of health professionals": care must be provided, and the data processing must be necessary for that provision of care.

This framework has a direct implication for AI use. Processing patient health data using an AI tool is not automatically covered by the healthcare treatment basis. The specific processing activity must be necessary for providing care and must be proportionate to the care being provided. Using a patient's identifiable health data to test an AI tool, experiment with a new application, or process through a general-purpose AI tool that will use the data for its own model training purposes cannot be justified under the healthcare treatment basis.

What Data You Can and Cannot Put Into General AI Tools

The practical question that most healthcare professionals are grappling with is which AI tools can be used with patient data and what restrictions apply. The answer requires distinguishing between categories of tool.

General-purpose consumer AI tools such as ChatGPT, Claude via the consumer app, Google Gemini, and Microsoft Copilot in its standard consumer configuration are not appropriate for processing identifiable patient data. These tools are not data processors under GDPR agreements with your organization. Their data handling terms do not provide the protections required for health data processing. Some may use inputs for model training. Even where a tool states it does not retain data, the absence of a formal Data Processing Agreement with your organization means the legal framework for health data processing is not in place.

This does not mean these tools cannot be used in healthcare workflows. A GP who uses ChatGPT to draft a template for patient letters, using no patient data, is using the tool appropriately. A junior doctor who uses an AI tool to look up general pharmacology information without entering any patient-specific detail is using the tool appropriately. The constraint is on identifiable patient data, not on general healthcare knowledge queries.

Approved clinical AI platforms that operate under Data Processing Agreements with NHS organizations or healthcare providers are a different category. Tools such as Nuance DAX, approved AI-assisted coding platforms, and AI tools integrated into EPR systems operate within a legal framework that includes the required data protection terms. Your organization's information governance team will have assessed these tools and the DPAs under which they operate.

Anonymised data is not personal data under GDPR if the anonymisation is genuine and irreversible. Truly anonymised patient data, processed through appropriate anonymisation pipelines, can be used with AI tools without the same restrictions. However, the standard for anonymisation is higher than many people assume. Removing a name and date of birth is not sufficient anonymisation if other data points in combination could re-identify the individual. Postcode, date of birth, diagnosis, and hospital number in combination are typically re-identifying. The ICO's guidance on anonymisation and pseudonymisation sets the relevant standard.

Warning

Pseudonymised data is still personal data under GDPR. Replacing a patient name with a reference number does not anonymise the data if the reference number can be linked back to the patient through any key held by any person with access to the system. Pseudonymisation reduces re-identification risk and can be a useful security measure, but it does not change the data's legal status. Do not assume that removing obvious identifiers makes data safe to process through general AI tools.

Knowledge check

A registrar is preparing a case presentation for a departmental teaching session. She wants to use an AI tool to generate a structured clinical summary from her notes. She removes the patient's name and date of birth before entering the information, but the input still includes the ward, admission date, diagnosis, specific comorbidities, and a reference to an unusual procedure performed during that admission. Which statement best describes the data protection position?

Select one answer.

NHS Data Security and Protection Toolkit

The NHS Data Security and Protection Toolkit (DSPT) is the assurance framework that NHS organizations and their suppliers use to demonstrate compliance with data security standards. For healthcare professionals, the practical implications are:

Your organization has a DSPT submission that covers approved data processing activities and tools. AI tools that process patient data should be assessed and approved through your organization's information governance process before clinical use. An individual clinician who adopts an AI tool that processes patient data without organizational IG approval is operating outside the framework, regardless of their good intentions.

The DSPT also establishes requirements for data sharing with suppliers, including the requirement for formal Data Processing Agreements. Any AI tool that processes personal patient data on behalf of your organization is a data processor, and the organization must have a DPA in place that covers the specific processing activities.

HIPAA Context for International Settings

For healthcare professionals working in or with US healthcare systems, the Health Insurance Portability and Accountability Act provides the relevant framework. HIPAA's Privacy Rule covers Protected Health Information, which is individually identifiable health information, and its Security Rule covers electronic PHI. AI tools that process PHI must either qualify as Business Associates and have Business Associate Agreements in place, or the PHI must be appropriately de-identified to the HIPAA standard before processing.

The HIPAA de-identification standard requires either expert determination that re-identification risk is very small or safe harbour de-identification that removes 18 specific identifiers. The 18-identifier safe harbour standard is more prescriptive than general anonymisation but shares the principle that individual data points in combination can be re-identifying.

HIPAA is also not the only privacy framework a US-facing AI healthcare workflow needs to satisfy. Several states layer additional obligations on top of it. California's Confidentiality of Medical Information Act (CMIA) protects medical information held by a broader range of entities than HIPAA's own definition of covered entities and business associates, and can reach some AI vendors and data flows that sit outside HIPAA's scope entirely. The California Consumer Privacy Act (CCPA), as amended by the California Privacy Rights Act (CPRA), also has provisions that interact with health data, though HIPAA-regulated PHI receives a partial exemption from it.

The CCPA/CPRA exemption for HIPAA-regulated data is narrower than it first appears: it exempts protected health information collected by a HIPAA-covered entity or business associate specifically, not "health data" in general. An AI vendor that is not itself a HIPAA-covered entity or business associate — for example, a general-purpose AI platform processing patient data without a signed Business Associate Agreement in place — falls outside that exemption and must independently satisfy CCPA/CPRA's own requirements. CMIA operates independently of HIPAA entirely: it applies to a broader range of entities than HIPAA's covered-entity/business-associate definitions, and HIPAA compliance does not, on its own, satisfy CMIA's separate obligations. In practice, a vendor's HIPAA compliance and BAA status should never be treated as proof of CCPA/CPRA or CMIA compliance — each framework's own applicability has to be checked independently for that specific vendor relationship.

For any AI tool processing patient data tied to a US care setting, and particularly one operating in California, treat HIPAA as the floor rather than the ceiling and confirm applicable state-level requirements before deployment.

Data Controller Obligations for Healthcare AI

If you manage a clinical team or are involved in selecting and deploying AI tools for your department or organization, you have data controller obligations to consider, not just data protection compliance obligations as an individual user.

Data controllers are responsible for: ensuring a lawful basis exists for the processing; completing Data Protection Impact Assessments (DPIAs) for high-risk processing; entering DPAs with data processors; maintaining records of processing activities; and ensuring that any AI tool deployed for processing health data meets the technical and organizational security requirements that special category data demands.

High-risk processing of special category data, which includes most AI-assisted clinical processing, requires a DPIA before deployment. A DPIA is not a compliance formality. It is an analytical process that identifies the risks to data subjects and the measures that will mitigate those risks. Deploying clinical AI without a DPIA where one is required is a GDPR breach.

Tip

If you are uncertain whether a specific AI tool can be used with patient data, the correct process is to raise it with your organization's information governance or data protection officer before using it, not after. The IG team can assess the tool, advise on lawful basis, and either approve its use or explain what data protection measures would need to be in place. Acting first and seeking approval later is the pattern that creates personal professional risk and organizational liability.

Pseudonymisation Misunderstanding — NHS Foundation Trust Research Team

Clinical Research Fellow, Respiratory Medicine

Context

A clinical research fellow was preparing data from a cohort of patients for an audit presentation and wanted to use an AI tool to help draft an analysis summary. Aware of data protection requirements, she removed patient names and NHS numbers from the dataset before uploading it. The remaining data included ward location, admission month and year, primary diagnosis, a rare comorbidity, and whether an unusual procedure had been performed. She considered the dataset anonymised and proceeded.

Action

The trust's information governance lead reviewed the presentation draft and identified the approach as pseudonymisation rather than anonymisation. The combination of ward, admission timeframe, rare comorbidity, and unusual procedure meant that individuals within the cohort could potentially be re-identified, particularly by colleagues who worked on the ward during the relevant period. The IG lead explained the difference between removing direct identifiers and achieving genuine anonymisation under the ICO standard, and flagged that the tool used had not been assessed for use with patient data and had no DPA with the trust. The draft was withdrawn and the analysis was redone using aggregate data with genuine anonymisation applied.

Outcome

No data was shared externally during the incident, which limited the regulatory exposure. The research fellow noted she had acted in good faith but had not understood that removing names and NHS numbers left a combination of indirect identifiers sufficient for re-identification in a small patient group. The trust used the incident in its next information governance training session for clinical research staff, focusing on the distinction between pseudonymisation and anonymisation and the requirement to raise AI tool use with the IG team before processing any patient-related data.

Practical Data Minimisation in AI Healthcare Workflows

A data minimisation approach to AI use in healthcare means using the minimum identifiable patient data necessary for the AI task, not the maximum available. Where an AI documentation task can be completed with initials and clinical identifiers rather than full names and NHS numbers, use the minimum. Where AI-generated drafts can be reviewed and approved before being associated with the patient record, do that rather than processing fully identified records through the AI tool.

Data minimisation is a GDPR principle, not just a good practice recommendation. It applies to all personal data processing, including AI-assisted processing. Building it into your AI workflow habits reduces the data protection risk footprint of AI use without reducing the clinical utility.

Quick check

A ward nurse wants to use a general-purpose AI chatbot to help draft a personalized discharge letter for a patient. She plans to enter the patient's name, diagnosis, medications, and discharge instructions into the AI tool to generate the letter. Why is this approach not compliant with UK data protection requirements?

Select one answer.

Exercise

~12 min

Your Task

Review the AI tools you or your team currently use — or are considering using — that interact with any patient information. For each tool, work through three questions: (1) Does it process identifiable patient data? (2) Is there a formal Data Processing Agreement between the tool's provider and your organization? (3) Has the tool been assessed through your organization's information governance process? If you cannot answer yes to questions 2 and 3 for any tool that touches identifiable patient data, document the gap and identify the correct person or team in your organization to raise it with.

Success looks like

  • You have identified at least one AI tool in your current practice that touches patient information and can answer all three questions for it
  • Where a gap exists — a tool without a confirmed DPA or IG approval — you have identified it clearly rather than assumed approval is in place
  • You can name the correct internal escalation route in your organization for IG approval of an AI tool (IG team, DPO, clinical governance lead, or equivalent)

Watch out for

  • Assuming that because a tool is widely used by colleagues it must have organizational IG approval — widespread use and formal approval are not the same thing
  • Treating a vendor's privacy policy or a no-retention statement as equivalent to a Data Processing Agreement — they are not

Hint

If you are unsure whether a DPA is in place, the most direct route is to ask your organization's information governance team or data protection officer. They maintain the register of approved data processors and can confirm the position quickly.

Key takeaways
  • Health data is special category data under UK GDPR and requires both a lawful basis and a specific Article 9 condition. Processing patient data through AI tools is not automatically covered by the healthcare treatment basis and requires case-by-case assessment.
  • General-purpose consumer AI tools are not appropriate for processing identifiable patient data because no Data Processing Agreement is in place with your organization and the data handling framework for special category health data is absent.
  • Pseudonymised data remains personal data under GDPR. Removing obvious identifiers does not anonymise data if any key exists that could link it back to an individual.
  • The NHS Data Security and Protection Toolkit framework requires that AI tools processing patient data be assessed and approved through organizational IG processes before clinical use, with formal DPAs in place.
  • HIPAA is a floor, not a ceiling, for US healthcare data: state laws including California's CMIA and the CCPA/CPRA can impose additional obligations on AI vendors and data flows, including in some cases reaching entities and data outside HIPAA's own scope.
  • A data minimisation approach to AI workflows reduces data protection risk without reducing clinical utility: use the minimum identifiable patient data that the AI task requires.