These are my first tentative thoughts on how assessment will have to change in response to AI. I have a short series of posts that I’m working on regarding this topic and I’d love to hear counter-arguments to help me test this idea.
First of all, we have to accept that soon (if not already) we won’t be able to confidently determine what content was created by generative AI.
Teachers and AI detectors can’t spot AI-generated text with the accuracy and reliability we need.
Generative AI can simulate student essay writing in a way that is undetectable for teachers, and AI-generated essays tend to be assessed more positively than student-written texts (Fleckenstein et al., 2024):
Our findings demonstrate that with relatively little prompting, current AI can generate texts that are not detectable for teachers, which poses a challenge to schools and universities in grading student essays.
And the solution is clearly not to be found in AI detection tools either (Weber-Wulff, et al., 2023):
The researchers conclude that the available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated text. Furthermore, content obfuscation techniques significantly worsen the performance of tools.
So now what?
I think we need to seriously consider the possibility that the only reasonable way forward is to allow students to use AI for their work, but we significantly increase the baseline of what we expect that work to look like. In my opinion, AI gives everyone a massive boost in capability in terms of what they are able to do. And so we should increase our expectations of what they submit for assessments.
For example, instead of asking students to:
- …design a building and create detailed blueprints, ask them to not only design the building but also create a fully immersive virtual reality walkthrough of the structure, complete with realistic textures, lighting, and interactive elements. They must also optimise the building’s energy efficiency, structural integrity, and environmental impact.
- …create a proposal for a public health initiative, ask them to launch an actual public health campaign that identifies at-risk populations, optimises resource allocation, and personalises outreach messages. In addition, the campaign should demonstrate a measurable impact on public health outcomes.
- …write a research paper on an environmental policy issue, ask them to develop an AI-powered decision support system for environmental policymakers. They should also create an interactive interface that allows policymakers to explore different policy options and understand their potential impacts.
If AI enables ‘superhuman’ performance, then our assessments should be modified to evaluate superhuman outputs.
- Fleckenstein, J., Meyer, J., Jansen, T., Keller, S. D., Köller, O., & Möller, J. (2024). Do teachers spot AI? Evaluating the detectability of AI-generated texts among student essays. Computers and Education: Artificial Intelligence, 6, 100209.
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19(1), 26.
Comments
8 responses to “How assessment needs to change in response to AI”
Hi Charmaine. Thanks for the question. I started writing a response here but it kept growing, so I published it as a separate post: https://www.mrowe.co.za/blog/2024/05/navigating-inequity-around-access-to-ai-in-higher-education/.
[…] (at least, not since the widespread availability of GPT-4 level models). But I do think we need to change assessment so that it looks very different to the current […]
Interesting ideas Michael, and I agree that we need to re-think assessment. You say, “AI gives everyone a massive boost…”, but how accessible is AI to all students (writing from a South African context). Will we disadvantage some students even more when assuming everyone has access to AI as a baseline?
[…] only way out of the dilemma – IMO – is to give students much more difficult assessment tasks. The kinds of tasks you can only complete with AI. The way we changed maths problems to be far more […]
Hi Ali. You’re right in that students need to put real effort into their assessment tasks. I’m coming from the perspective that AI augments their capabilities, and they will use AI, so they’ll be able to complete our existing tasks with no effort at all. Changing our expectations around what the task looks like means that they still need to put real effort into the task. And there’s no reason that you can’t include a reflective component (although AI can already write a better reflection than any student, so we’ll have no way of knowing if the student really wrote it).
How about going one step back to “pen and paper” exams and assessments? I mean I don’t mind students using AI for learning or preparations or so, at the end there will be an exam (as it used to be) and let the student reflect and write their learnings to some questions or problem solving scenarios in action. Maybe it is me but from the student perspective (as I am still), honestly, I should put real efforts in those kind of assessments rather any other type of tasks.
Agreed. And the role of lecturers needs to change as well. At the moment, I can’t imagine that individuals will have the skills or experience to grade these kinds of assessments.
Agreed! This does mean, however, that module outcomes and assessment criteria must be adapted accordingly – the entire curriculum must change to cater for cyborg students, cyborg teachers and cyborg citizens.