Leiter Reports: A Philosophy Blog

News and views about philosophy, the academic profession, academic freedom, intellectual culture, and other topics. The world’s most popular philosophy blog, since 2003.

  1. DH's avatar

    As an international student, I honestly don’t know how good an indicator the GRE is, at least for students whose…

  2. Cameron Buckner's avatar
  3. Herbert Hart's avatar

    This might filter out intelligent neurodivergent students who struggle with timed admissions tests.

  4. Mark J Engleson's avatar

    When jt comes to students’ ability to do advanced logic, the math section of the GRE is more valuable than…

  5. Brian Leiter's avatar
  6. Cameron Buckner's avatar
  7. Anonymous's avatar

A different approach to academic assessment in the age of AI

Philosopher Matthew Hammerton comments; an excerpt:

Prohibition simply doesn’t work when you can’t enforce it. Telling students not to use AI is like telling them not to think about pink elephants. The result is a two-tiered system where conscientious students follow the rules while the rest can cheat with impunity. This breeds cynicism and erodes trust in higher education.

Relying on technological safeguards for enforcement is also bound to fail. Once students discover workarounds to your safeguards—which they inevitably will—cheating will proliferate. The technological arms race between educators and students is not one that educators can win.

Retreating to in-class exams preserves the form of the essay but sacrifices its substance. The whole point of assigning essays is to give students time to think, research, synthesise and argue—cognitive processes that cannot and should not be rushed. An in-class essay is to real essay writing what speed dating is to marriage—it might look similar from a distance, but it’s missing everything that matters.

Abandoning essay writing might work for vocational courses where analytical writing was always peripheral. However, it’s not viable for most of the humanities and social sciences where the deep, higher-order, independent thinking involved in essay writing is essential to intellectual development….

In reality, those who excel at deep, independent thinking are precisely the ones best equipped to find creative, intelligent, and ethical ways of using whatever AI emerges. Students who focus on narrow, technical skills like ‘prompt engineering’ are the ones most likely to be outpaced as the technology evolves.

So, prohibition is naïve, abandonment is reckless, and retreating to in-class essay exams is shortsighted. Yet the essay is definitely worth saving. What is needed is a way of reconciling its pedagogical value with the reality of AI.

The mistake in current debates is the false choice between two extremes: all-out prohibition or uncritical embrace. Those who want prohibition don’t seem to understand that students can engage in deep, higher-order, independent thinking while using AI tools. For instance, a student might ask generative AI for critical feedback on their draft and judiciously use this feedback when making revisions. Other beneficial applications include copy editing, brainstorming when stuck, answering specific queries during the writing process, and converting detailed notes into fluent prose.

However, each of these can be misused in ways that undermine intellectual development. This then becomes the challenge: How do we enable students to use AI tools while preventing inappropriate use that would stunt their intellectual growth?

Many assume the solution involves detailed guidelines that meticulously catalogue all possible AI applications and specify when it is educationally appropriate to use them. But such an approach is doomed: it would be unwieldy, quickly outdated, and nearly impossible to enforce. Faculty would also disagree over what counts as ‘appropriate’, ensuring endless controversy.

I propose a different approach built on a simple, intuitive principle: Students must take intellectual responsibility for their work. This means that when you submit an essay, you should be able to explain and defend every significant choice you made—the thesis you advanced, the evidence you selected, the counterarguments you considered, the conclusions you drew. If you can’t explain why you structured your argument the way you did, or why you chose one interpretation over another, then you haven’t taken intellectual responsibility for your work.

This sounds to me like every essay has to be subjected to an oral exam. This will require a massive investment by universities in faculty to undertake such scrutiny.

Leave a Reply

Your email address will not be published. Required fields are marked *

15 responses to “A different approach to academic assessment in the age of AI”

  1. Let’s distinguish the practical question of how academics, particularly in humanities, should run their essay-form assessments in the short term, from the broader philosophical questions of how AI and related technologies will irreparably change an wide array of cultural practices going forward and the implications of those changes for the prospects of people knowing anything, being able to do anything, including live a good life. The problem is that answers to the philosophical questions must inform sound answers to the practical one. That is a problem because we don’t have answers to the philosophical ones (yet).

    These (otherwise admirable) comments may soon sound like the old geezer librarian in the late 1980s going on about the importance of diligently maintaining the Card Catalogue, just as hard drives and data systems were being built to forever replace them. Except in this case the changes go to the very structure and development of the human mind and the shape of the social world.

    I suspect that children who were born this morning will grow up with no acquaintance of what Hammerton refers to as “real essay writing”. They will never really know what that is, though we do, much like most of our ancestors knew something of what it really meant to “till the fields” (in agriculture) though a vanishingly small number of people know now. Something else replaced it. Hammerton’s choice of analogy is revealing: “An in-class essay is to real essay writing what speed dating is to marriage”. And yet, “speed dating” (on social media apps etc) has quickly become a fundamental part of how humans interact, whereas marriage has become far less common, for better or worse.

    Just as friendship was irreparably transformed by a series of technological advances (letter writing, telegram, telephone, rapid transport, and of course the internet), such that the relationships that we would reasonably characterize as friendships would be utterly alien to anyone living in (say) the 8th century, I suspect the use of these generative technologies will eventually reshape the concept of what it means for a person to know something, as well as (most germane to the assessment question) what it is for a person to understand something, what counts as a good explanation of something, and what it means for a person to be able to explain something.

    For that reason, appeals to “deep, higher-order, independent thinking involved in essay writing [that] is essential to intellectual development” may eventually sound like appeals to the virtues of the Card Catalogue: quaint, historically interesting, but ultimately practically and socially irrelevant. Children born today will also never know the unsettling and deep anxiety that we all feel in pondering that future, knowing the road we’ve travelled thus far.

    I tend to think that philosophers are quite possibly uniquely situated to sort much of this out, but I also suspect that the state of academic philosophy (and universities more generally) is such that the kind and extent of “reimagining” required to do it won’t be the sorts of things that yield high “impact” scores for tenure and promotion purposes and so will be discouraged, in direct inverse proportion to its importance for understanding ourselves — and our students.

    1. @Michael: do you really believe that? Consider the calculator, or all of its decendents including spreadsheets and accounting software and everything like that. If I ran a fortune 500 company and I was hiring an accountant, I would expect them, first and foremost, to be fluent in using all these tools. But I would also expect them to know how to add, to mutliply and to do long division. I would be highly suspicious of any accountant who couldn’t do those things. Wouldn’t you? Even if the future of writing philosophy papers is to chatgpt as accounting currently is to accounting software, I still don’t think anybody will be able to do it well if they cant write a cogent 3 page essay that’s been edited and revised without any assistance from a computer.

      And if I wrong about that, (and maybe I am!) I’m not sure what will be left to teach. If accounting software becomes so AI sophisticated that a person using it doesn’t have to actually learn any tax law or accounting principles, then I doubt well still be teaching accounting. And if a person interacting with chatgpt to make philosophy papers doesn’t have to even be able to write a three page paper unassisted, then I’m not sure what we’ll be teaching philosophy students other than “type ‘write something philosophically interesting’ into your chatbot”.

    2. Old geezer librarian here, one who happens to have overseen his city library’s pre-Web (c. ’91) deployment of Internet access for the library’s users, its migration to a new integrated library system (which encompasses functions like the online catalog, the registration of borrowers, etc.), and the decommissioning of its card catalog. Qua stereotype, your illustration suffices to make its point that people resist change and, in particular, technological change, which so often requires one to RTFM. But to the extent they existed, librarians who advocated for diligent maintenance of the card catalog were not so misguided, because their concern would have been for the integrity of the data and for the utility of the “syndetic structure,” the added-value cross-references that permit researchers to spring from a known object to relevant new ones. I’ll only mention that automated solutions, even these days, are not always elegant, due in part to–now I’m going to stereotype–developers’ perfect lack of interest in the purposes of a bibliographic catalog, either print or online. Librarians, on the other hand, have generally been early adopters of (at least experimenters with) new technologies, including AI.

    3. Matthew Hammerton

      Michael, you say “speed dating (on social media apps etc) has quickly become a fundamental part of how humans interact, whereas marriage has become far less common, for better or worse”. I think there are some counter-trends worth noting. Many young people are disillusioned with the superficiality and gamified nature of online dating, leading to more traditional forms of dating making a comeback – https://www.nytimes.com/2025/06/06/well/dating-irl-analog-online.html . And while total marriages are down, their quality has risen: divorce rates have fallen, marital satisfaction has increased, and those who marry today tend to be relatively privileged people whose lives are stable.

      This suggests that humans have a persistent need for genuinely meaningful experiences. The frustrations with online dating, and the high satisfaction among those who do marry, both reflect this.

      No doubt new technologies will transform our epistemic practices in countless ways that we don’t fully understand yet. However, I suspect that a desire for genuine understanding and intellectual autonomy will remain because they are central to the kind of meaningfulness that humans need.

      1. Again, the use of the example of marriage, as expressive of a “persistent need for genuinely meaningful experiences”, is question-begging in precisely the same way as your appeals to “real essay writing” and “intellectual development” and so on. As you may know, some people have claimed to have fallen in love with their chatbot: https://www.theguardian.com/tv-and-radio/2025/jul/12/i-felt-pure-unconditional-love-the-people-who-marry-their-ai-chatbots. There are many jarring phenomenological details here, but what is most striking to me is how soon this occurred after the introduction of this technology for public use. Given the exponential rate of development in these models, in 10 years time, the current LLMs that are the apparent objects of affection will seem extremely primitive. And these AI bots haven’t even been implanted seemlessly into humanoid robots, the sophistication of which is also accelerating. That is to say, one possible (and not too distant) future is expanding the concepts of both love and marriage to include relations with machines. I suspect it’s only a matter of time before activism begins in some jurisdiction (Japan?) to expand the legal right to marry to reflect this conceptual change.

        If the concepts, and practices, of love and marriage can be expanded in that way, why not the concepts around our epistemic practices as well? You seem to assume that such notions are incorrigible. You may even bite the bullet and claim the same of love and marriage. But that leaves increasingly common human experiences unexplained and not well understood, in a way that risks dismissing them out of hand. And I fear that appealing to (venerable and heretofore reliable) concepts in (particularly the Western) philosophical repertoire will function as roadblocks to understanding these implications of AI and other technologies.

        We tend to view “The Enlightenment” as a discreet historical period that has long been over, running from the late 17th to I suppose sometime in the early 19th century. And yet, it’s possible that in a few centuries, it will become clear that The Enlightenment actually ran to the end of the 20th century, and has just now ended. That transition will be marked by the philosophical concepts of the Enlightenment becoming increasingly inapt – culturally, politically, intellectually – in the face of the technological advances we are all witnessing unfold. That is a possible outcome, of course there are many others.

  2. As I’m sure you know, Hammerton discusses your objection at the end of the article. What do you make of his response?

  3. What’s proposed sounds like the arrangements I have been familiar with since my undergraduate days. The student writes a weekly essay, and that essay is the starting point of an hour’s discussion with the tutor (the word in Oxford) or supervisor (the word in Cambridge). This is great for learners: I used to love having a grown-up paid to listen to me expound my views. It is also great for teachers: it’s actually enjoyable to discuss a topic with youngsters who’ve made the effort to write an essay on its rudiments.

    1. Matthew Hammerton

      Yes, that is how I was thinking about it. Many responses to AI-written essays demand significant instructor time, but often in unpleasant ways—playing detective to identify and report suspected AI use, or combing through Google Docs histories to verify “human” writing. One reason why I favour oral exams is that they are a much better use of my time for both me and my students.

  4. I’ll likely provoke responses by suggesting that the answer isn’t oral exams, which are way too time-consuming to pay professionals to administer, but instead to have a chatbot administer the equivalent of your oral exam, performed in something like a “lockdown browser” preventing students from drawing upon outside resources. This is an economical way to give you most of the benefits of letting them use available tools to creatively explore a topic in a sustained way that enhances their own understanding, while also directly assessing how well they’d internalized that understanding, so they won’t be incentivized to take shortcuts to produce a finished paper.

  5. Why not set essay questions in the normal way (so that students have a chance to research at home), but instead of having them submit an essay, have them write it in an “open book” in-class examination? That would preserve most of the benefits of setting research essays, but avoid the possibility of cheating? (When I say “write it” in class I mean by hand, on paper – not on a computer).

    1. To Alex’s comment:
      Open-book in class essay-format exams don’t deal with many of the skills developed in standard essay writing.
      Take-home, untimed essays allow — and require — students to edit, revise, and sharpen their work in a way that’s not possible in an in-class, timed situation, open book or no.
      In terms of content, of students being able to state the arguments and material learned in class, there may not be much difference. But there’s a lot to be lost in developing students’ ability to *structure* their writing.
      We live with, even expect, some infelicities in the structure of in-class essay exams, because it’s inevitable. That doesn’t prepare students for any kind of real-world scenario, where these kinds of writing problems will not be tolerated.
      I don’t know what the answer is, but I do know that this isn’t going to cut it.

  6. Isn’t it ironic that AI just prompts the same old discussion: how to save the students from themselves? One would expect that they realize that it is in their interest to learn to think. But evidently, they won’t or wouldn’t.

    Otherwise, I also think oral exams and discussions are the only obvious solution, although I think it is strange that what I think is the major problem with them is not mentioned by anyone: they are massively prone to bias. The way the student looks, the way they speak and behave and so on. Endless sources for bias.

    Perhaps indeed chat bots are the solution.

    1. Matthew Hammerton

      To be fair to our students, I think we should acknowledge that there is also a significant number who recognize the value of learning to think for themselves, yet feel pressure to take the shortcuts their peers are taking lest they fall behind in the credentialing rat race (e.g., “if I can get an A for all my assignments without putting any work in, that will free up time to take another internship, which would be a big boost to my job prospects).

      You raise a fair concern about bias, but I would be cautious about overstating it. First, bias creeps into every assessment task we conduct. For example, when we mark essays, seeing a student’s name may influence how we assess their work, and even blind marking is not immune, since cultural references in the essay can reveal background. Yet we don’t take this as an objection against essay assignments.

      Second, although the risk of bias is real and worth mitigating, the evidence suggests its impact is often modest. Much of the science on implicit bias that was heavily promoted over the last few decades and not stood up to sustained critical scrutiny. For a competent, non-bigoted professor, bias seems to have only a minor impact on assessment. On that basis, I doubt it is a major problem for oral exams, though I agree we should remain alert to it.

  7. I’ve only had a handful of students cheating by using AI for essays. I’d like to think that is related to my incorporating the following in my teaching. 1. We extract modes of inference and other logical techniques from each of our readings. (E.g. how to find a counterexample to a conditional statement. Modus ponens. Etc.) 2. Then we learn how to create argument maps of selected longer arguments in the readings. (They have to turn these in once a week. 3. Once they get the hang of it and learn how to raise analyze arguments and critique them for validity and soundness, they gain confidence and seem to enjoy argument mapping. To break things up I sometimes ask students to draw a pictogram of an argument. Or hold an oral debate. 4. For essays, the prompts require that students use (generally) three additional sources (other readings we’ve done). The questions require that students think across all four sources. The questions are highly detailed and specific. (Never set a prompt that simply asks students to ‘evaluate the argument of paper X.’ Of course they will use AI to do that.) 5. Then scaffolding: they have to submit a draft of their thesis. Then an argument map of their paper. 6. In their first draft they are required to anticipate an objection to their own argument and respond to it. 7. I then comment fully on their first drafts. They are then required to submit a final draft which responds relevantly to my comments. The first draft is 4-5 pages. The final draft must be 5-7. It must also use full footnotes: each time they quote or paraphrase or otherwise rely on a source, they must cite the page/source. Only in rare cases do students use AI. When it happens, it is painfully obvious: there is no increase in length from first to final draft, there is no argument map submitted or the map doesn’t match the final product, there is no modification in argumentation between drafts or there is no footnoting or no original objection or relevant responses to my comments. In those cases, I have an office hour ‘oral exam’. But I can just fail the paper on the grounds these requirements were not met.

  8. AI helps me educate myself; it serves as a first reader for my humble tries at fiction and answers questions about the world and about philosophy and many domains. I do often have access to human experts and authorities and were I at a good college I’d avail myself of them.
    If you’re curious and have ideas of your own, AI is a fine intellectual companion- not world class authority but more like an intelligent Grad student; if you come to AI with a blank slate and empty mind it will do no more than put words in your mouth, words you’d often utterly fail to comprehend.
    Many students back in the day merely reflected back what the teacher lectured; perhaps today there are more students like that, or they have lost the ability for independent thought or for curisoity.

Designed with WordPress