Skip to main content
Deliberate AcademyProfessional AI Education
~17 min left
Lesson 3 of 10
17 min read10 XP

Academic Integrity in the Age of AI: Assessment Redesign

Deliberate Academy Editorial Team

Reviewed for accuracy and professional relevance

You're 3 lessons in — don't lose your progress.

Sign up free
What you'll learn
  • Explain why unsupervised written assessment has reduced evidential value in an AI-accessible environment
  • Evaluate AI detection tools accurately, including their false positive rates and the risks of relying on them for misconduct decisions
  • Apply at least three assessment design strategies that maintain integrity without depending on AI prohibition
  • Design an assessment component — process portfolio, oral defense, or in-class task — that tests genuine learning regardless of AI access
  • Distinguish between what an AI-use disclosure policy requires and what it accomplishes for academic integrity culture

Academic integrity in education has always involved a certain amount of assumed trust. When a student submitted a written essay, educators generally assumed it represented the student's own thinking. AI tools have not made that assumption impossible to hold, but they have made it substantially less reliable for certain types of assessment. The response that serves students, educators, and the purposes of education best is not to restore the old trust assumption through prohibition, but to design assessment approaches that are genuinely informative about learning regardless of AI access.

How AI Has Changed What Student-Written Work Means

Before AI writing tools became widely accessible, a student-submitted essay was a reasonable proxy for what the student understood about a topic, how they could structure an argument, and how well they could communicate in writing. These proxies were imperfect — tutoring, ghost-writing, and collusion existed before AI — but they worked well enough as practical assessment tools at scale.

An AI-generated essay can now be produced that matches or exceeds the structural and stylistic quality of many student essays. It can present a coherent argument, cite relevant evidence, and address marking criteria points. A student who submits this essay may have engaged with the topic deeply and used AI as a drafting and refinement tool, or may have entered a prompt and submitted the output with minimal engagement. The written product does not distinguish between these two students, and the learning they have done is very different.

This changes the evidential value of written assessment, particularly unsupervised take-home essays. It does not eliminate the value of written work as a learning activity or as one component of a varied assessment portfolio. But it does require educators to think more carefully about what any given assessment is actually measuring, and whether the design of the assessment is still fit for purpose.

AI Detection Tools: Capabilities and Significant Limitations

The instinctive first response in many institutions to AI-generated student work has been to deploy AI detection tools. Turnitin, GPTZero, Copyleaks, and similar tools analyze student submissions for patterns associated with AI generation. These tools are widely marketed in education settings and have been adopted by many universities and some schools.

The problem is that the detection tools are substantially less reliable than their marketing suggests.

False positives. Multiple studies and real-world cases have documented that AI detection tools flag human-written work as AI-generated at significant rates, including the work of non-native English speakers, highly formal writing styles, and some students with autism who write in unusually direct and structured ways. Acting on an AI detection tool's positive result as a basis for academic misconduct proceedings without additional evidence is professionally and legally risky.

Arms race dynamics. AI models are developing rapidly, and AI-generated text is becoming harder to distinguish from human-written text. Detection tools that perform adequately today may perform poorly on text generated by more advanced models within months. Relying on detection as a long-term strategy is building on an eroding foundation.

Inconclusive results. Many detection tools provide a probability score rather than a binary determination. A "73% AI probability" determination tells an educator very little that they can act on with confidence. The score is influenced by writing style, topic familiarity, and tool version as much as by whether AI was used.

AI detection tools can be one piece of information in an academic integrity investigation, but they are not reliable enough to be the primary basis for a misconduct finding. The institutional and professional response to AI in assessment cannot rest primarily on detection.

Warning

Bringing academic misconduct proceedings against a student based primarily on an AI detection tool result, without corroborating evidence, creates significant risks for educators and institutions. Students have successfully challenged such findings, and institutions have faced reputational and legal consequences. The appropriate use of detection tools is as a prompt for further investigation, not as independent evidence of misconduct. This is a professional judgment issue as much as a technical one.

Redesigning Assessment to Work Without Detection

Lecturer in Business Studies, Further Education College

Context

A business studies lecturer at a further education college was marking a batch of BTEC Level 3 reports on marketing strategy when she noticed that several submissions had an unusually consistent structure and formal register, well above the typical standard for the cohort. She ran the submissions through an AI detection tool and received probability scores ranging from 55% to 79% for five students. She was aware that several of the flagged students were EAL learners whose writing was characteristically direct and formal.

Action

Rather than initiating misconduct proceedings on the basis of detection scores she knew to be unreliable, she scheduled brief individual conversations with each flagged student — 10 minutes each — asking them to walk her through one section of their report: what their argument was, why they had chosen that evidence, and what they would add if they had more space. She also reviewed their previous authenticated work from the start of the year. The detection scores were high for three students whose previous work was at a similar standard and who could discuss their reports clearly. Two students could not explain specific sections of their report at all.

Outcome

The two students who could not account for their report content were referred for formal investigation. The three whose detection scores were elevated but whose prior work and oral responses were consistent were not referred. The lecturer subsequently redesigned the next unit's assessment brief to include a mandatory five-minute verbal defense for all students, removing the need for detection tool screening entirely. She noted that the verbal component assessed the learning the written submission was supposed to represent, which detection could never do reliably.

Knowledge check

A student who is a non-native English speaker submits a history essay that scores 81% on an AI detection tool. The essay is well-structured and uses formal academic language consistently throughout. The student's previous work has been at a similar level. What does this scenario most clearly illustrate about AI detection tools?

Select one answer.

Assessment Approaches That Maintain Integrity in an AI-Accessible World

The more sustainable response to AI and academic integrity is assessment redesign: creating conditions under which the assessment demonstrates genuine student learning regardless of AI access outside the supervised context.

Process-based assessment. Assessing the process of producing work, not just the finished product, significantly reduces the value of submitting AI-generated work. A student who submits a portfolio that includes research notes, draft iterations, a reflective commentary on their process, and a final essay is much harder to replace with AI-generated output than a student who submits only the final essay. The research notes and drafts reveal whether genuine engagement with the material occurred.

Oral examination and viva voce. Requiring students to discuss and defend their submitted work in an oral examination directly tests whether they understand the work. A student who has submitted AI-generated work and has not engaged with the content will be exposed quickly by questions about specific arguments made, evidence cited, or analytical choices. Oral components do not need to replace written assessment entirely: a brief viva on a piece of submitted writing adds significant validity.

In-class and supervised tasks. Supervised assessments, whether handwritten examinations or in-class written tasks, are AI-inaccessible by definition. They place a different kind of pressure on students than take-home work, but they provide reliable evidence of what students can do independently. Balancing supervised and unsupervised assessment within a course gives a more complete and defensible picture of learning.

Authentic and contextualised assessment. Assessment tasks that require engagement with specific local, organizational, or experiential contexts are much harder to reproduce with AI. A case study analysis that requires applying concepts to a specific organization the student has visited, or a reflective account of a specific professional experience, or an analysis that references a specific dataset the class has worked with, demands the specific contextual knowledge that AI does not have.

Specification of AI use. Some educators are moving towards explicitly defining what AI use is permitted and requiring students to disclose and reflect on how they have used it. An assessment design that says "use AI to generate a first draft of this analysis, then revise it based on your research, and submit both the AI draft and your revised version with a commentary explaining what you changed and why" is actively teaching AI literacy while maintaining academic integrity. The submitted work demonstrates the student's capacity to evaluate, revise, and extend AI-generated content.

Tip

Redesigning assessments for an AI-accessible world does not always require creating new assessment types from scratch. Often the most effective adaptation is adding a process or discussion component to existing assessments. A reflective commentary on how a piece of work was produced, or a five-minute discussion of a submitted essay with the marker, adds validity without redesigning the entire assessment approach.

Institutional Policy Considerations

Assessment integrity in an AI-accessible world requires institutional policy as well as individual educator adaptation. The most effective institutional approaches share several characteristics.

Differentiated, not blanket, policy. A blanket prohibition on all AI use ignores the legitimate educational case for teaching students to use AI responsibly. A policy that differentiates between assessment types, year levels, and contexts, and that defines clearly what AI use is and is not permitted in different assessment situations, is more educationally coherent and more enforceable than a blanket ban.

Staff professional development. Policy without educator understanding is ineffective. Staff need to understand what AI tools can and cannot do, how to recognize the limitations of detection tools, how to have assessment integrity conversations with students, and how to redesign assessments where needed. This is a professional development investment, not a one-time communication.

Alignment across a program. Assessment integrity is weakest when different courses within a program take completely different approaches to AI. Students are confused by inconsistency, and the cumulative learning signal is reduced. Program-level consistency about the principles governing AI use in assessment is more useful than a patchwork of individual module policies.

Having Honest Conversations with Students About AI

The students who most benefit from honest educator conversations about AI are not those who are trying to cheat, but those who are genuinely uncertain about what the right approach is. Many students find themselves in a position where AI tools are accessible, where their peers are using them, where institutional policy is unclear or inconsistent, and where no one has explained the educational and ethical reasoning behind whatever policy their institution has adopted.

Educators who engage with AI openly, who explain the reasons behind their assessment design choices, who discuss what AI can and cannot demonstrate about learning, and who model a thoughtful professional relationship with AI are providing something more valuable than a set of rules. They are demonstrating the critical thinking about technology that students need throughout their professional lives.

Quick check

A university lecturer runs Turnitin's AI detection analysis on a student's essay and receives a result indicating 68 percent AI content. The student's previous work has been strong and their writing style in earlier submissions was similar. The lecturer wants to raise an academic misconduct concern. What is the most professionally sound approach?

Select one answer.

Exercise

~10 min

Your Task

Take one unsupervised written assessment you currently set — an essay, a report, a reflective piece, or a structured written response. Identify the specific learning outcomes it is designed to assess. Then redesign it by adding one process or demonstration component that would require students to show genuine engagement with the material, regardless of whether AI helped produce the final draft. Write the redesigned brief in full, including what students must submit and what the added component requires.

Success looks like

  • The redesigned brief names the learning outcomes explicitly and the added component tests them directly, not just the finished product
  • The added component — a process portfolio, a reflection, a short oral discussion, or a draft annotation — cannot be completed by submitting unrevised AI output
  • The brief is specific enough that a student reading it would know exactly what is permitted, what is required, and how the added component will be assessed

Watch out for

  • Adding a reflective commentary that asks 'how did you find the task?' rather than requiring specific disclosure of what AI did and what the student changed — the second is assessable, the first is not
  • Designing the added component as punitive catch-out rather than as a genuine learning demonstration — students disengage from assessment designs that feel like surveillance rather than professional development

Hint

The most effective add-on for most written assessments is a five-minute conversation requirement: 'be ready to discuss your three main arguments and what evidence you would add if you had more space.' A student who genuinely engaged with the material can do this. One who submitted AI output unmodified cannot.

Key takeaways
  • AI writing tools have reduced the evidential value of unsupervised written assessment as a reliable proxy for student learning, because AI-generated work can match or exceed the structural quality of student essays without the student having engaged with the material.
  • AI detection tools have significant false positive rates, are subject to rapid obsolescence, and produce probabilistic rather than binary results. They should be used as a prompt for further investigation, not as independent evidence in misconduct proceedings.
  • Assessment approaches that maintain integrity in an AI-accessible world include process-based assessment, oral examination components, supervised in-class tasks, authentic contextualised tasks, and explicitly designed AI-use frameworks where AI use is defined and disclosed.
  • The most sustainable institutional response to AI and academic integrity is differentiated policy supported by staff professional development, not blanket prohibition without educator understanding.
  • Honest educator conversations with students about AI, explaining the reasoning behind assessment design and the educational purposes that assessment serves, are more effective at building academic integrity culture than rules alone.