Can AI Do Philosophy? An Experiment (Part Two)
MY SITE
Home
Academic Philosophy
Research
Teaching
Logic Textbooks
Autonomous Philosophy
Research Groups
Private Tutoring
Popular Philosophy
Talking In Circles
Absolute Irony (blog)
To Big Finities and Beyond
Selected Public Writing
Paintings
About Me
Absolute Irony
Philosophical musings on<br>whatever suits my fancy
A continuation of the old blog, found here
Can AI Do Philosophy? An Experiment (Part Two)
8/22/2026
0 Comments
This is my second post detailing my experiments in attempting to prompt current AI systems to autonomously generate publishable work in philosophy. In my first experiment, I did not do much work to optimize my prompting strategy. This was deliberate. I wanted to see how a frontier model would do if the basic instruction was just “write a philosophy paper.” The answer, unsurprisingly, was that it did not do great. How well could it do if it was prompted a bit more carefully? The answer, to my surprise, is that it could actually do quite well. Indeed, with just a bit of refinement of my prompting technique, I got ChatGPT-5.6-Sol to autonomously produce a philosophy paper in my area of expertise that I would likely conditionally accept if asked to review it for a good journal. This post documents how I got this outcome and what some of its implications might be.
The Process<br>Last week, I made a bet with my colleague that AI systems would soon be able to autonomously produce philosophy papers that would pass peer review in top journals. In the context of this bet, we stipulated some rules. First, we stipulated that, in order for the paper to count as “autonomously” produced by the AI system, the human prompter could not themself spend more than two hours actually engaging in the paper-writing process. Second, we stipulated that the human could not simply give the AI their own ideas and walk it through a paper that they already knew how to write themself. It wouldn’t be too surprising if, given such hand-holding, an AI could write a paper that was perhaps publishable but substantially worse than the one the human would write. No, in order to succeed in the challenge, the AI system had to really do it all autonomously.
The condition that the AI system autonomously generate the paper requires a different workflow from that developed by Simon Goldstein, which is designed for human/AI collaboration and involves quite a bit of human hand-holding. By contrast, for my approach, I needed to be entirely philosophically hands-off. To work within the confines of the rules of the challenge, I adopted a two-agent workflow, opening two different contexts in ChatGPT (using 5.6 Sol for both) and going back and forth between them. One agent was the author, who was actually working to write the paper, whereas the other agent was a consultant, who would give me instructions on how to prompt the first agent. This enabled me to provide substantially more detailed prompting than I did in my first experiment without having to spend time writing the prompts myself.
I started by having my consultant criticize my earlier approach. One clear issue with my first approach was that not enough time was spent on the idea-generation stage. Once the agent found an idea that it liked and that looked plausible enough to me, I had it go ahead and start writing the paper. But, in the end, the idea was really not that good, and so the quality of the resulting paper ended up being capped at a relatively low bar. Even if it would have been possible in principle for a good human writer to produce a publishable paper pursuing the basic idea, when the mediocre idea was coupled with the relatively poor writing skills of the AI agent, there was little hope of its producing a paper that was actually good.
Another related issue was that I let it search too broad a space of possible paper topics, and it ended up working on something that was outside my specific area of expertise. This made it impossible for me to evaluate at a glance whether certain ideas were promising. With these two issues in view, I had the consultant agent write up a starting prompt for the author, asking it to target the very specific subfield in which I work—inferentialist, proof-theoretic, and bilateralist approaches in philosophical logic and semantics. Rather than just telling the author to consult the literature on this topic, the consultant’s prompt provided a list of specific relevant papers in this field for the author to examine.
The prompt for stage 1 involved searching for a promising idea that engaged with this literature and preliminarily stress-testing it. When the author came back with an idea, I gave it to the consultant, who came back to me with the prompt for stage 2. This stage involved even more stress-testing, ultimately leading to a decision about whether to proceed with the idea or kill it. Though the author initially seemed...