Leiter Reports: A Philosophy Blog

News and views about philosophy, the academic profession, academic freedom, intellectual culture, and other topics. The world’s most popular philosophy blog, since 2003.

  1. David Wallace's avatar

    My experience is with fairly technical philosophy of physics papers, and with Opus 5. – copyediting: amazing, much better than…

  2. A B Carter's avatar
  3. sahpa's avatar
  4. Paul Canis's avatar
  5. Jason's avatar
  6. Michael Magoulias's avatar
  7. John S's avatar

An experience with feedback from an LLM

A colleague elsewhere, who subscribes to Astra, Open AI’s high end LLM model (at least $100/month), told me that he has found it useful for giving feedback on early drafts–he compares it to the feedback one gets from a good first-year graduate student. Not doing any historical work, he was curious how it would perform on that kind of paper, and kindly invited me to give him a Nietzsche paper I was working on, which I did. Astra generated several pages of feedback, that was impressive in a number of ways, but deficient in some others. On the positive side: it picked up incorrect citations to texts (e.g., GM II:23, not GM III:23 as I had it), filled in some missing citations (correctly), caught some logical ambiguities in formulations of points, and raise one or two interesting substantive objections. On the negative side: it has no sense of style, no sense of developing a dialectic (e.g., exploring a possible reading of a passage, then rejecting it in favor of another reading–it just wanted to go straight to the “preferred” reading), and it knows nothing substantive about Nietzsche. I recently discussed this same paper with a very good Nietzsche scholar, who gave me more meaningful feedback. But I think the comparison to a first-year graduate student is fair, with the caveat that many first-year graduate students know more about the primary text than Astra. What Astra could not do was challenge any aspect of my reading based on other texts of Nietzsche’s. What Astra could do is assess the internal logic of my argument (with the caveat that it gets confused about a developing dialectic). In only one case, did it misread the argument, and, as noted, n some cases, it read claims I was making carefully and revealed real ambiguities.

Curious to hear what other reader experiences have been with the better models. Please indicate what you work on before describing your experience.

,

Leave a Reply

Your email address will not be published. Required fields are marked *

5 responses to “An experience with feedback from an LLM”

  1. My suspicion is that you would get better feedback if you pointed the LLM at the additional relevant Nietzsche texts when asking it for criticism. It’s asking too much for the training weights to retain all possible interpretations and readings of all possible texts the LLM has been trained on and be able to comment on any given text from a cold start. (Even expert academics have to go back and consult classic texts because human memory is fallible.)

    In this case, I think it would be an interesting test to use an agentic AI which can access the Nietzsche texts on a computer alongside the paper. Why? Because the main limit AIs face is the context window they can maintain. The top Claude agents have a context window of 1 million tokens, but even that can be blown on a complicated task. The usual way around that is to break the task into subtasks and have independent agents read a single, separate chapter of Nietzsche, looking for the relevant parts for your paper. The independent agents write up detailed reports which are then sent to a single overseer who can integrate all the findings without having wasted time and context reading unnecessary parts of the text. This would then given a much more informed — and, I suspect — nuanced evaluation of your paper.

    Current Claude models can do this automatically: a single overseer will break the task down into multiple parts, allocate each task to a separate agent, collect the reports from all the individual agents (potentially validating their report using an adversarial agent which deliberately probes the report looking for mistakes and flaws), and then integrate all the separate findings.

    1. I also wonder if the paper was injected into the context window, in its entirety, at all. I don’t know Astra, but I know that many setups use selective retrieval on attachments (akin to interacting with the attachment through a search engine, with the LLM sending search queries to the engine, and the engine sending back relevant ‘chunks’ of the attachment). That can easily mean that Astra never ‘sees’ the whole paper at all. That would perhaps explain why it struggled to grasps the ‘developing dialectic’ of the paper, since that’s a more ‘meso-level’ feature of the text.

  2. I would readily give a passing grade on the process front, but I am not engaging these machines at the product front.

    We all have our process of assembling our thoughts and materials for a work. Over the course of a project I will have a large desk with piles of material, and I am essentially physically laying out for myself, the material I intend to build in the production of the work itself. We all have our ways at this level of “preparation” and “pre-production” so to speak.

    As a rough measure I would say my efficiency at this gathering-together stage is easily trebled by rustling AI’s ability to summon references and establish the basic “piles” on my work desk, so to speak. As a tool, the machine is responsive to this level of inputs at a delightful level. This use-case assumes the researcher is a good researcher and knows the subject well. As I say, an ideal undergraduate major that only exists in our dreams.

    On the production and post-production front, I am not sure I am interested in AI for that, one single bit.

    For that same reason, I am wary of being the one who says, “This is graduate level.” I’m with the Fields Medalists here. We risk destroying the discipline if we use it for our work. Graduate students, are not replicable.

  3. Brian, any chance of getting your colleague to specify how he used Astra to produce the feedback? Did he used standard chat, Canvas, Deep Research or Workspaces? What was the actual prompt? Did he use more than one? What was the total time spent by Astra to produce the final results?

  4. My experience is with fairly technical philosophy of physics papers, and with Opus 5.
    – copyediting: amazing, much better than any human I’ve worked with. (I accept ~80% of its copyediting advice.)
    – structure and argument: variable, but sometimes very good. It has a bit of a CS attitude (perhaps unsurprisingly) but it has certainly spotted structural infelicities. (I accept ~30% of its structural/argument advice).
    – style: it tries to homogenize my style towards its own, and has a bit of a tin ear for anything imaginative. (I accept ~0% of its style suggestions.)
    – literature and sourcing: it’s excellent on what’s out there, less good on the content of what’s out there (i.e. it provides good links but doesn’t always summarize them well). It’s extremely good at matching technical statements I’ve made to places in the literature where similar statements have been made already; I would anticipate it’s much better at doing this in math or theoretical physics than in philosophy.

    My default prompt includes “I am fairly thick-skinned; if I get something wrong, tell me.” I think this has helped a bit with sycophancy.

Designed with WordPress