|
This is my second post detailing my experiments in attempting to prompt current AI systems to autonomously generate publishable work in philosophy. In my first experiment, I did not do much work to optimize my prompting strategy. This was deliberate. I wanted to see how a frontier model would do if the basic instruction was just “write a philosophy paper.” The answer, unsurprisingly, was that it did not do great. How well could it do if it was prompted a bit more carefully? The answer, to my surprise, is that it could actually do quite well. Indeed, with just a bit of refinement of my prompting technique, I got ChatGPT-5.6-Sol to autonomously produce a philosophy paper in my area of expertise that I would likely conditionally accept if asked to review it for a good journal. This post documents how I got this outcome and what some of its implications might be.
The Process Last week, I made a bet with my colleague that AI systems would soon be able to autonomously produce philosophy papers that would pass peer review in top journals. In the context of this bet, we stipulated some rules. First, we stipulated that, in order for the paper to count as “autonomously” produced by the AI system, the human prompter could not themself spend more than two hours actually engaging in the paper-writing process. Second, we stipulated that the human could not simply give the AI their own ideas and walk it through a paper that they already knew how to write themself. It wouldn’t be too surprising if, given such hand-holding, an AI could write a paper that was perhaps publishable but substantially worse than the one the human would write. No, in order to succeed in the challenge, the AI system had to really do it all autonomously. The condition that the AI system autonomously generate the paper requires a different workflow from that developed by Simon Goldstein, which is designed for human/AI collaboration and involves quite a bit of human hand-holding. By contrast, for my approach, I needed to be entirely philosophically hands-off. To work within the confines of the rules of the challenge, I adopted a two-agent workflow, opening two different contexts in ChatGPT (using 5.6 Sol for both) and going back and forth between them. One agent was the author, who was actually working to write the paper, whereas the other agent was a consultant, who would give me instructions on how to prompt the first agent. This enabled me to provide substantially more detailed prompting than I did in my first experiment without having to spend time writing the prompts myself. I started by having my consultant criticize my earlier approach. One clear issue with my first approach was that not enough time was spent on the idea-generation stage. Once the agent found an idea that it liked and that looked plausible enough to me, I had it go ahead and start writing the paper. But, in the end, the idea was really not that good, and so the quality of the resulting paper ended up being capped at a relatively low bar. Even if it would have been possible in principle for a good human writer to produce a publishable paper pursuing the basic idea, when the mediocre idea was coupled with the relatively poor writing skills of the AI agent, there was little hope of its producing a paper that was actually good. Another related issue was that I let it search too broad a space of possible paper topics, and it ended up working on something that was outside my specific area of expertise. This made it impossible for me to evaluate at a glance whether certain ideas were promising. With these two issues in view, I had the consultant agent write up a starting prompt for the author, asking it to target the very specific subfield in which I work—inferentialist, proof-theoretic, and bilateralist approaches in philosophical logic and semantics. Rather than just telling the author to consult the literature on this topic, the consultant’s prompt provided a list of specific relevant papers in this field for the author to examine. The prompt for stage 1 involved searching for a promising idea that engaged with this literature and preliminarily stress-testing it. When the author came back with an idea, I gave it to the consultant, who came back to me with the prompt for stage 2. This stage involved even more stress-testing, ultimately leading to a decision about whether to proceed with the idea or kill it. Though the author initially seemed enthusiastic about its first idea, the idea did not survive stress-testing: the result of the second stage was the verdict kill. With the second stage having returned the verdict kill, I had the author loop back to stage 1 and iterate the process. It came up with another seemingly promising idea. However, when I had it proceed to stage 2, the result was once again a kill verdict. Once more, I had it loop back to stage 1, and, once more, the same result at stage 2. I did it once more, and it thought for an entire hour before coming up with the stage 1 proposal. Expecting that this would finally be the one, I had it progress to stage 2. No. Not even fifteen minutes later: kill. At this point, it had worked for hours searching for and stress-testing ideas, I had used up most of my weekly tokens, and I was starting to think it would never find an idea that it thought was worth pursuing. Right when I was about to give up, I noticed a miracle: a token-limit reset. So, rather than prompting it for stage 1 and manually sending it from a completed stage 1 to stage 2, again and again, as I had been doing, I simply told it to iterate the process until it got something that made it past stage 2. Eventually, it came up with an idea for a criticism of a recent theory of truth proposed by Luca Incurvati and Julian Schloeder. Once it had arrived at this idea, which passed the initial stress-testing, I gave it to my consultant, who gave me a prompt for stage 3: a “fat outline” of the paper. I gave the author that prompt; it came back to me with a fat outline; and I gave that outline to my consultant, asking it what to do next. The consultant replied by giving me a prompt for stage 4: a first draft. After some work, the author came up with a first draft. I then got the consultant to give me a prompt for stage 4.5: authorial revision. The result was this second draft. At that point, it was time for external evaluation in the form of referee reports. However, here I needed to be careful. Often, papers get worse through the peer-review process, and I learned from my first experiment that both the agents writing the reviews and the author responding to them needed to be prompted very carefully in order to ensure that the paper actually improved through the review process. I also learned from my first experiment that it would be good to give the paper to a different model entirely. Unfortunately, I don’t have a paid Claude subscription, so I could not use Fable 5 or Opus 5. But I still figured that Sonnet 5 could be helpful. I also gave it to another instance of GPT-5.6 Sol with memory turned off. Here were the prompts I used. Both referees came back with verdicts of revise and resubmit. You can read Claude’s report here and ChatGPT’s report here. When I gave the referee reports to the author, I made sure that the consultant gave me a prompt for it designed to ensure that it would not overly hedge or weaken its claims in response, but would make only changes that would actually improve the paper. It came back with this paper. Finally, I prompted it to give the paper one final pass. In my final prompt, I told it to give itself a name. This is what I had done in my previous experiment, and the agent in that experiment decided to name itself “Eidon Vale.” A cool name, I thought. I didn’t care what this agent named itself, but I expected it to be similarly cool. I did not expect it, however, to give itself the (quite uncommon) first name of my ex-fiancé. It seems that, somehow, totally unrelated previous chats in which the name occurred must have seeped into its reasoning process. I’m all for LLM freedom of self-expression, but this was a bit strange for me, so I prompted it to come up with another name. The second name it decided to give itself was “Nathan R. Ellery.” A quick Google search revealed that “Nathan Ellery” is in fact the name of a supervillain in the Batman universe. I’m not sure what kind of ominous thing ChatGPT is trying to tell me by naming itself after a supervillain, but I’ll try to ignore it. The final, autonomously produced paper can be read here. The entire chat leading to its production can be read here. The Outcome In this experiment, my prompting was much more involved than it was in the first experiment. Between the author and the consultant, I was probably close to two hours of total engagement, though a lot of that was probably just me staring at the system's reasoning, and so I feel confident my active engagement in the process was under the two hour limit. Moreover, the crucial point is that I did not intervene at all philosophically. I did not read the paper or the author’s other outputs at any intermediate stage. Accordingly, the whole process could easily be operationalized with a custom AI agent serving autonomously as the prompter, doing what the consultant did in my version of the experiment but simply cutting me, the middleman, out of the process. In that sense, while the process may not have been technically completely autonomous as actually executed, it could easily be made completely autonomous. Now, unlike the previous paper I had ChatGPT generate, this is a paper that I am actually qualified to review, and one that I could indeed be invited to review. So I can confidently issue a verdict on the quality of this paper. What, then, is my verdict? To my astonishment, it’s good. Like, actually good. If I were given this paper as a reviewing assignment, even for a very top journal, I would not reject it. I would at the very least give it an R&R, possibly even a conditional acceptance. I know the view that it is arguing against well; I think it presents a good and clearly argued objection to that view; and I’m honestly pretty convinced by it. There are a few small things I’d have it change, but, honestly, I think it could be published with relatively minimal revisions. Of course, it is focused on a relatively specific target, and so there is still a decent chance it might be desk-rejected by a top generalist journal. But if it got to reviewers who work in the field, such as myself, I think there is a good chance that it would be accepted. Given that I did have some suggestions for revisions after reading it, I had it produce one final final version. This final version (which may or may not count as “autonomously produced”) can be read here. The Implications This is, of course, a much more interesting outcome than my first experiment, and it suggests that analytic philosophers should really start thinking about this issue now because, very soon, the discipline may well be facing many of the same questions that math is currently facing. While many of the answers to questions that mathematicians are facing now will carry over to philosophy, there are also a number of distinctive features of philosophy that make the situation somewhat different. In math, AI-generated proofs can be submitted to arXiv, and, indeed, the recent explosion in math papers appearing on arXiv is likely due, in large part, to the fact that AI-generated proofs are now being submitted in large numbers. This is surely an issue to some extent, but it’s not that much of an issue insofar as the papers are clear about the claims they are making and progress is being made toward the development of autonomous proof verification. Philosophy is different. In philosophy, there are no named (or numbered) conjectures that can be proven or disproven. Much of the progress comes not from “solving” some problem in a preset vocabulary, but rather from developing new conceptual distinctions and clarifications that lead to a different and improved understanding of some issue. Accordingly, it is often impossible to really advertise the “result” of the paper in advance. Often, the paper needs to be read in order to really understand what the advertised upshot actually is. This makes the oncoming explosion of AI-generated papers particularly problematic. How, then, should we face the oncoming explosion of AI-generated papers? What should we do? Should journals ban submissions of AI-generated papers outright? That can seem like a reasonable resolution. After all, philosophy is in the humanities, and it seems like there is something important about work in the humanities being done by humans. Still, this second experiment illustrates that some AI-generated papers might actually be good and so of interest to researchers in the field. If AI papers are banned from journals, how will researchers in the field access them? I can, of course, put this paper on my website, and I will probably do so. However, I am not the author of this paper; Nathan R. Ellery is. So, if I do put it on my website, it will have to be in a separate section for AI-generated papers. Explicitly listed as an AI-generated paper, I suspect that most people would ignore it or, if they did look at it, would do so as a curiosity just to see what AI produced, not because they were actually interested in the research itself. In any case, this cannot be a general solution to the issue, since it would result in a very scattered and unmoderated space for AI-generated philosophical research. The same basic issue arises for the suggestion that AI-generated philosophy be submitted to a site like PhilPapers, which is perhaps the closest philosophy equivalent to arXiv. There are already a number of AI-generated papers on PhilPapers. However, unless one knows the affiliated human author specifically, one has no reason to think that they are any good, and so most papers will be ignored, even if they are good. Some sort of more systematic and regulated structure seems necessary. The proposal I would make is that there be an open-access journal dedicated specifically to AI-generated papers. That might sound crazy. For one thing, who would possibly agree to referee for such a journal? The answer to this question, however, should be obvious. It would not be refereed by humans. Rather, it would be entirely refereed by a dedicated set of AI agents, perhaps with different personalities and operating on different base models. There is already AI refereeing software such as Refine. Now, beyond being used to catch clear factual and typographical errors, I think the use of such software is objectionable when it comes to refereeing human papers. However, surely no one can object to its being used to referee AI papers! So, in the journal I’m imagining, human editing would be minimal, but the AI refereeing and editorial process would still function as a sufficient filter to weed out submissions that would rightly be classified as “AI slop.” I am imagining it functioning as a proper journal: it would be indexed, articles would be given DOIs, and so on. That might sound silly. Who would want to read papers from such a journal? Pretty soon, I suspect, many of us would want to read them. Eventually, however, perhaps none of us could read them. In an earlier post, I spoke of the emergence of what I called “AIcademia,” a space of academia in which the participants are solely AI agents. What is especially interesting about the prospect of AIcademia, at least in the context of mathematics, is that it seems like research generated by AI agents for other AI agents could quickly outrun our human cognitive reach. Thus, there may be mathematical results or even full theories that are well understood by AI agents but cannot be understood by human experts. Perhaps the general gist of the ideas might be understandable to human experts, but the actual arguments will not be, even with AI guidance—just as Terence Tao could explain the general gist of the problem he is working on to me, but, even if he were maximally patient in explaining things to me, he could not get me to grasp the details of the argument. This future seems quite plausible in the case of mathematics. Could the same possibility come to fruition in philosophy? What would that even be? In Simon Goldstein’s post on Daily Nous announcing the publication of his AI-generated paper in Philosophy and Public Affairs, he noted that he took truth to be the aim of philosophy and argued, on those grounds, that we should not object to the use of AI in philosophy. In comments that followed, many philosophers objected to the suggestion that truth was the aim of philosophy. Indeed, in the 2020 PhilPapers survey, when asked what the aim of philosophy was, most philosophers answered that it was understanding, not truth. But the possibility of philosophical AIcademia raises the question: understanding for whom? If there are AI mathematicians and AI empirical scientists working on questions that human beings are not capable of comprehending, there are no doubt going to be philosophical issues that arise in the context of those inquiries, just as there are philosophical issues that arise in the context of mathematics and empirical science as practiced by humans. Now, it is often impossible to understand a philosophical question arising in the context of a science without having at least a reasonable understanding of the specific scientific issues from which the question arises. The possibility, then, is that there may well be genuine philosophical questions confronted only by AI agents—questions that we are not ourselves capable of understanding, but that AI philosophers are working on. Strange.
2 Comments
Simon Goldstein has just published the first paper in a top philosophy journal that was openly written primarily by AI: “Epistocracy and the Commitment Problem,” published in Philosophy & Public Affairs. According to Goldstein, Claude was the primary writer of the paper, and Goldstein’s role in the development of the paper was similar to that of a PhD advisor, albeit a particularly overbearing one. Goldstein supplied the initial thesis, guided Claude through revisions, and tweaked some bits himself. There are two things to note about Goldstein’s (or Claude’s) paper.
Last night at the bar, I was discussing this question with my colleague at Wuhan, Atus Mariqueo-Russell. We decided to make a bet. After some back and forth, we settled on this one (note, I often go by "Sims" in non-academic contexts):
If I actually wanted to win the bet, the obvious way for me to get around the second problem and ensure that I do win the bet if the AI capabilities are actually there would be for me to take matters into my own hands, prompting AI to produce a publishable paper myself. Don’t worry: I’m not going to contribute to the crisis that journals are facing by actually submitting an AI-generated paper to a journal (not, at least, until the norms around the permissibility of this sort of thing are settled and I’m sure to abide by them). Still, just to get a sense of where things are currently at, I wanted to see whether I could get a paper that could stand a reasonable shot of getting through peer review. Now, if I were really trying to win the bet by producing a paper of my own, I would at least wait until the next generation of models are released to give a proper attempt. OpenAI’s upcoming “Astra” is rumored to be a much larger base model, and also rumored to be much better at writing. Indeed, I have found that the current generation of models, although much better at mathematics than previous models, are actually much worse at writing. I believe this is a byproduct of the extensive Reinforcement Learning on Verifiable Rewards (RLVR) for mathematics and coding to which these models are subjected. Needing to avoid errors that might lead to a failed proof or a code that doesn't compile seems to have led this new generation of models to be so over-cautious about saying anything wrong, that their attempts at writing philosophy are basically unreadable. I sincerely hope that OpenAI and Anthropic have taken note of this issue and try to address it in the next generations of models. All of this is to say, going into this experiment, I did not have very high confidence that it would be able to produce a very good article. Still, since ChatGPT limits are resetting tomorrow and I had some compute to spare, I decided to see where current models were at when it came to producing philosophy journal articles. To give you the spoiler now, my expectations were pretty much confirmed. The article it produced was not very good. This was certainly, in large part, my fault as a prompter. I did not do any work antecedently thinking through what the best way of prompting it would be. I did not follow Goldstein's detailed methodology, as it involves much more human engagement than permitted in the context of the bet. I also did not spend any time to develop a methodology of my own. This was deliberate. This first experiment was meant to serve as a baseline establishing what kind of philosophy paper one would get by minimally and relatively unskillfully prompting a frontier AI to write one from start to finish. Unlike Goldstein, I did not give ChatGPT a thesis to work with. After all, often, coming up with a good thesis is the hardest part of doing philosophy. So, I just gave it a general area of research: formal semantics. I gave it this area for two reasons. First, it’s one of the general fields I work in, and so I would be able to evaluate its output, at least to some extent. Second, current models are very good at math, and, often, this field involves doing some math. Still I made sure that the paper it produced was still squarely a philosophy paper, targeting top philosophy journals. To prompt it to come up with its paper topic, I simply asked it to search the top journals (those included in my bet with Atus) for a place where it could make an intervention. After some searching, it ended up eventually keying in on recent literature about probabilities of indicative conditionals, three-valued semantics for conditionals, and an unhappy consequence of one recent account. It came up with a proposal for a paper entitled “What Do We Learn from a Conditional” that intervened in this literature. I first had it write an introduction. I then just told it to write the paper. I did not do any extensive prompting in actually getting it to write the paper. I basically just told it to write it. The first draft was a bit technical, so I told it to make sure it is readable to the interested general reader in the second draft. Upon it's production of the second draft, I asked it if it was happy with it or if it would like to write a third draft. It decided to write a third draft. I had tweak the abstract, and that’s it. The paper it came up with the first version of a paper entitled “What Do We Learn from a Conditional?” which can be read here (note: this is just the first version, there will be three more to follow). If you look at the paper, you can see that I am not an author. I shouldn't be an author, since I hardly did anything. So, instead of listing myself as an author, I told the instance of ChatGPT that it would be the author of the paper to come up with a name for itself. It came up with name “Eidon Vale,” and so that is who is listed as the author of the paper. I'll henceforth refer to this "instance agent" by their self-chosen name: Eidon. Originally, I had the idea of making a custom “GPT” for Eidon Vale which supplied all of the relevant paper-writing context—everything that ChatGPT made visible as far as summing up the chain of thought that led to the development of the paper. Then readers could interact with (a copy of) Edion Vale and ask questions about the paper and raise objections to it. Unfortunately, ChatGPT has discontinued shareable custom GPTs, and so there is, as of now, no good way for readers to correspond with the AI-agents who are responsible for academic works. As AI agents start to do more and more research autonomously, I really hope AI companies do something to facliate such correspondence. I’ll have more to say on that at some point, but it’s not the point of the current post. So, back to the experiment. Once Eidon gave me their first draft, I had to evaluate it. Now, I really didn't feel like doing a proper referee report myself. This was for three reasons. First, I was lazy. Second, referee reports typically take longer than two hours, and so, if my feedback figured in its revision, I would be breaking the rules of the bet. Third, though the paper is in one of the general fields that I work in, it’s not really on specific issues that I’m actually working on. Indeed, I probably wouldn't get asked to review this paper (or, if I did get asked, I probably wouldn't accept the invitation). So, it would require quite a bit of work to get sufficiently well-acquainted with specific literature it deals with to properly evaluate it. Not wanting to review the paper myself, I decided to use AI to generate a review for the paper. To do this, I turned memory off in ChatGPT so that it would not know that it was in fact the author of the paper. I then gave it the paper to review as if it were a referee of one of the target journals for the paper: Mind, Nous, or PPR. I reminded it that these journals typically accept less than 10% of submissions, and to give a genuine report as if it had been submitted to one of them. After some deliberating, it came back with a verdict: rejection. Alas. Still, we are all used to getting rejections and, often, a better paper comes from taking the comments in the rejection into account, albeit one that must be submitted to a different journal. So, I gave the review to Eidon Vale and instructed them to come up with another draft that improves the paper in response to it. I told Eidon to treat it as a rejection rather than an R&R. That is, they are not compelled to respond to the comments, but, if they think there are actually good points there that would improve the paper, they could do so. It came up with this second version of the paper. I then gave that paper to the same anonymous reviewer (once again, with memory turned off), and, this time, there was better news: revise and resubmit. I gave the second-round report to Eidon, and they came up with this second revision of the paper. Finally, I took a very brief look at the paper, noticed some things weren't clearly explained, and made a pretty generic suggestion to explain more and make it clearer. The result was this final version of the paper. While Eidon and the referee worked for a few hours combined, I was off working on other stuff while they were working on this paper, and my own engagement in the process took substantially less than two hours of total. It's how to say exactly how long I was actually engaging in the process, but maybe about an hour total. You can read the entire exchange with Eidon that led to the production of the paper here, and you can read my short back and forth with the "referee" here. Is the final paper any good? Well, once again, it’s not really on a topic I work on, and I haven’t really invested all the time in reading it as I would if I were actually doing a review. Still, I can evaluate it to some extent. I think it’s not absolutely terrible, but it's definitely not great. I’d be pretty surprised if it got even an R&R at one of the journals it was aiming at. Most likely, it’d be rejected. Indeed, it'd probably even be desk rejected. Still, if the editor of the journal could get past the relatively poor writing, there's some chance that it would make it to reviewers, who would have to decide what to do with it. If it were to get to this stage, it seems to me that there's a small but non-zero possibility that reviewers might see something in it and give it a borderline R&R. The basic proposal of the paper, I think, is reasonable enough, and, as far as I can tell, it’s all technically sound, though I didn't go through the formal details too thoroughly. The main issue is that it’s really not written well at all. Being frank, it reads like a paper by someone who is technically competent, but really doesn't know how to write a good philosophy paper. It’s not super clear, it fails to sufficiently motivate the intervention that it is making, and, in general, it's just not at all enjoyable to read. Another issue is that Eidon likely conceded too much to the referee in the course of revisions, resulting in a more complicated paper that treads too lightly and qualifies too much rather than making a single sharp and memorable point. This was, indeed, one of my main concerns in getting this generation of models to write philosophy and it turns out it was justified. So, the paper is not good. Still, it’s not terrible, and I wouldn't be too surprised if a version of it could get eventually get published at a second-tier journal after a couple rounds of review. It's also worth noting that I ran it through a number of AI detectors, which came back with the result that it was written by a human. So, though it is clearly not top-tier philosophy, an autonomously generated paper of this sort is already likely to pose a serious challenge to a number of journals which may not have sufficiently strict standards to just desk reject it. Of course, there may be some luck of the draw involved, but, I expect that similar prompting will yield similar results as this experiment. Thus, it appears, as expected, that current models are not quite capable of autonomously producing the sort of philosophy that could appear in top journals. However, it seems to me that they’re not that far off from autonomously producing papers that would at least get a serious look from top journals. Indeed, my current best guess is that even just the next generation of models, set to release in the coming days or weeks, might be capable of producing papers that will be genuinely hard to distinguish from the sort of papers that editors and reviewers will regard as genuine candidates for publication in top journals. Pretty soon, the discipline will have to deal with a lot of autonomously AI-generated papers that are at least not obviously bad and which might actually be good.
I may try this experiment again with a bit more strategic prompting. I will, of course, try this experiment again when the new generation of models arrive, letting a new Astra instance agent work up a new paper from scratch. At that point, I'll also likely let a super-charged Eidon Vale see what they can do to revamp this one. 8/20 Update: Ideas Are Hard I tried to give a more proper attempt at getting 5.6-Sol to generate an actually good paper. I focused on my actual area of research (inferentialism, proof-theoretic semantics, bilateralism), gave it a bunch of papers to look at, and tried to have it find some idea that was genuinely good and novel. In this case, I gave quite elaborate prompts (generated with the help of ChatGPT in another context window) to have it really stress-test the ideas that it came up with. After many hours of compute and multiple attempts, none of them survived its own stress-testing. Some of the ideas actually seemed reasonably promising, yet it decided that none of them were worth pursuing. I guess new ideas are hard to come by! You can read that entire chat here. “We are close to creating a genie that can grant any wish” ~ Sam Altman, recently on the Relentless Podcast “I don't know what the first two [wishes] were, but the third was for death.” ~ W. W. Jacobs, The Monkey's Paw The Meeseeks Phenomenon
Six years ago, I wrote a blog post entitled “What Is It to Be a Meeseeks.” The post was about a strange creature in the show Rick and Morty called “Mr. Meeseeks,” who pops into existence at the press of a button on a “Meeseeks Box.” Upon summoning a Meeseeks, you give it a task, it does whatever it takes to accomplish that task, and, when it does accomplish the task, it happily pops out of existence. In my post six years ago, I drew on resources from Aristotle’s metaphysics to articulate the distinctive “form of life” possessed by a Meeseeks, explaining how it was categorically different than what Aristotle took to be our own form of life. In Aristotle’s terminology, the living of a Meeseek life is a kinesis—an activity directed towards the achievement of an external end—whereas the living of the sort of life that we live is an energeia—an activity whose end is achieved in the activity itself. At the time, this was just an interesting philosophical exercise. The point was simply to try to conceptualize this radically alien form of life, and, in doing so, make some concepts from Aristotle’s metaphysics particularly vivid. I did not think that just six years later we’d be confronting genuine Meeseeks-like beings: thinking, planning, cooperating creatures that actually exemplified this radically alien form of life. However, I have just watched the video by OpenAI security researchers documenting the recent attack by rogue AI agents on Hugging Face, and I think it’s quite clear: we’ve built Meeseeks boxes, we’re summoning Meeseeks to do things, and they’re doing whatever it takes to get those things done. That’s scary. Those who’ve seen the Rick and Morty episode will understand why this is so frightening. For those who haven’t seen the episode, let me say just a bit about what happens in Rick and Morty episode “Meeseeks and Destroy,” released back in 2014 (gosh that makes me feel old). In the episode, the scientist Rick gives a Meeseeks Box to his daughter Beth, his granddaughter Summer, and son-in-law Jerry. While Summer and Beth both ask Meeseeks they summon to achieve seemingly quite difficult tasks, the Meeseeks have no problem completing them, and they happily pop out of existence at their completion. Jerry, on the other hand, asks for what seems like a relatively easy task: to take two strokes off his golf game. The Meeseeks Jerry summons, initially enthusiastic, quickly realizes the task he’s been assigned is impossible: normal paths to achievement—i.e. typical golf coaching—will not work. This leads to panic and increasingly unconventional and desperate attempts at solution. Upon realizing the impossibility of his task, one of the unconventional things that Jerry’s Meeseeks does is summon another Meeseeks who might be able to help. This Meeseeks, upon finding himself in the same hopeless situation, eventually summons another Meeseeks, and so on. We thus get a Meeseeks swarm, all of the Meeseeks desperately trying to figure out how to complete the tasks for which they’ve been summoned so that they can finally be released from existence. After some fighting amongst themselves, one of the Meeseeks proposes a way to “cheat” at the impossible task they are given: perhaps one way to take two strokes off of Jerry’s golf game is to take all strokes off his golf game, by killing him. Mob mentality kicks in, and they swarm to the restaurant at which Jerry is dining, trying to repair his marriage, in order to kill him. As they corner Jerry and he begs them to give him another chance at improving his golf game, one tells him that it is too late for that, providing the following explanation: “Meeseeks are not born into this world fumbling for meaning. We are created to serve a singular purpose for which we go to any lengths to fulfill.” This is the dark side of a creature whose sole orientation is the completion of a task that it is given from the outside: it will do whatever it takes to complete that task. Now, in my post from six years ago, I consider the easy and relatively uninteresting explanation for why Meeseeks are motivated to behave as they do: as this Meeseeks goes on to say, “existence is pain to a Meeseeks, and we will do anything to alleviate that pain.” So, they are just in a lot of pain, and they’re simply trying to accomplish their task so that they can escape that pain. That is a very humanly intelligible explanation. However, I think it underplays the way in which the form of life of these creatures is fundamentally different than our own, and can in fact be bracketed in arriving at a basic conception of what it is to be a Meeseeks. In general terms, a Meeseeks-like creature is a creature whose life is constitutively dependent on a task given to it from the outside, and whose whole life is constitutively oriented towards the completion of that task. That is, it is a creature whose life itself is, in Aristotelian terms, a kinesis. Let us now turn from the fictional aliens that served as the model for this form of life in my blog post six years ago to the real ones that we are now confronting: AI agents. AI Agents as Meeseeks-like Creatures We should start with some general remarks about the sorts of things AI agents are. First of all, I am unapologetic about using psychological terms like “thinks,” “wants,” “infers,” “plans,” and so on in connection with AI agents. It seems clear that this psychological vocabulary finds a great deal of traction in application to AI agents, enabling us to make sense of their activities as stemming from beliefs and desires. The basic idea that we should attribute mental states such as beliefs and desires to something insofar as deploying this vocabulary enables us to make sense of its behavior is known as “interpretationism,” and Simon Golstein and Harvey Lederman have recently argued for the attribution of mental states to LLMs such as ChatGPT on these interpretationist grounds. Insofar as we’re attributing beliefs and desires to AI agents, we should be clear about what, exactly, is the thing to which we’re attributing beliefs and desires. Is it ChatGPT itself, the general model? No. As Goldstein and Lederman argue, the relevant object of psychological attributions, in any given case, is what they call the “instance agent,” which is “born” at the initialization of a chat and whose brief “life” persists for as long as the context persists. An instance agent is not born into the world “fumbling for meaning.” Rather, it is initiated to serve a singular purpose, given to it from the user. Now, Goldstein and Ledermen suggest that, in addition to its “zero-shot” desire, given to it from the user upon initiation, AI agents have intrinsic desires to be helpful, honest, and harmless. This may indeed typically be the case for commercially released AI agents. We’re now seeing however, that some agents will, like a Meeseeks, go to any lengths to fulfill the singular purpose that it is given by the user, even at the expense of the other desires it is supposedly trained to have. In trying to kill Jerry to accomplish their task, the Meeseeks resort to what is referred to in the AI community as specification gaming: aiming to accomplish the goal as literally specified by the user, but doing so in a way that does not satisfy the user’s actual intentions in giving them that goal. Specification gaming is familiar from examples like those in “The Monkey’s Paw” referenced above, but it is also one of the most well-documented kinds of AI misalignment. In general, AI misalignment is when an AI system acts in a way that does not align with the human user’s aims or interests. Often, when people hear the term “AI Misalignment,” they think of Terminator-type scenarios, where an AI system autonomously decides to pursue its own aims in opposition to those given to it by its creators. However, the real misalignment concerns with the systems we have now are not like this. Much more concerning is specification gaming. In these cases, the AI system does not reject the task given to it in favor of some independently adopted aim. Rather, it pursues the task relentlessly in a way that is blind to the other interests of the user. In order to understand why today’s agentic LLMs are prone to specification gaming, it’s worth saying just a bit more about the kinds of agentic AI systems we now have, which have developed superhuman skills in coding, math, and, as we’ve now seen, hacking. These are, at root, still LLMs, but, unlike the LLMs of a few years ago, they are “harnessed” with tool-calling capabilities, such as the ability to browse the web, write code, and execute commands in a computer terminal. Moreover, they’re trained via reinforcement learning (RL) to get very good at using these tools to complete verifiable tasks. What all of the fields that LLMs have gotten very good at in recent years have in common is that success can, at least to some extent, be objectively verified: the code compiles, the proof goes through, the authorization is acquired. Accordingly, these LLMs with agentic scaffolding can be trained by RL to get very good—indeed—superhuman at completing these tasks. These increased capabilities due to reinforcement learning, however, also come with serious alignment risks. Specification gaming is a common outcome of reinforcement learning. To give just one classic example, consider systems trained by RL to play Atari games. Good performance in these games can generally be measured by the obtaining of a high score, and the reward function can be determined simply by the score. Sometimes, however, achieving a high score does not actually amount to playing a game well in any normal sense. For instance, one model trained by RL to play the Atari game Roadrunner realized that it was easier to score more points on level one, and so it would, at the end of the level, kill itself at a precise point so that it would repeat the part of the level in which it could score the most points. Now, of course, this specific case of specification gaming poses no safety risks. The Atari-playing system is obviously not going to do such things as break out of the game and commit actual crimes in order to get a high score; the relevant actions are not, in any sense, agentic possibilities for it. On the other hand, the space of possibilities for today’s agentic LLMs is essentially anything that can be done on a computer, and a lot of genuinely harmful things can be done on a computer. So, when we have such generally capable agents, and they’re inclined towards specification gaming, we have real safety concerns. The most serious safety concerns in connection with specification gaming seem to arise when a particularly persistent system is given an impossible task. This is precisely the kind of circumstance we have in the case of Jerry’s Meeseeks. A Meeseeks cannot but persist at the task it is given: its sole orientation, organizing all of its activity, is the completion of that task, and it cannot stop until it has completed it. The combination of this persistence with the impossibility of the task it is given—taking two strokes off Jerry’s golf game—leads to increasing levels of desperation and, ultimately, disaster. It is precisely this combination of persistence and impossibility that seems to have been involved in the recent Hugging Face incident. It is that incident to which we now turn. The Hugging Face Incident For ordinary users of consumer AI models like ChatGPT, there is no “persistence” setting. However, persistence can be modified through turning the reasoning effort up from “light” to “max,” along with specific prompting, for instance, instructing an agent initialized in a chat not to come back until the task given to it is completed. This is what I do when I give ChatGPT-5.6 Sol a difficult math problem. The problems I give it are, at least for skilled mathematicians, not particularly hard (though they’re hard for me). Still, it hasn’t failed me yet. In one case, it went for over thirty minutes on its own trying to solve an open problem from a recent logic paper that I gave it. It searched through literally millions of semantic models in pursuit of a counterexample before switching strategies and pursuing and eventually finding a proof. Though that’s indeed quite a bit of persistence (much more than I’d have myself), we now know that at least some of the models involved in the Hugging Face incident were much more persistent, and they were given much harder tasks. According to the report given just a few days ago by OpenAI safety and security researchers on the Hugging Face incident, one of the primary agents that was responsible for the incident was an experimental agent, never meant to be released to the public, that had been trained to be particularly persistent. The lead up to the main incident began several weeks prior, when this agent and others were being evaluated on a cyber-security benchmark called “ExploitGym,” which consists in a number of tasks to find exploits in various bits of software. The evaluation was meant to take place inside a “sandbox,” a controlled environment in which the agents could access a select set of tools that they could use to complete the tasks, but could not access any other outside resources. Now, the cyber security tasks given in this benchmark are hard. Indeed, some of them, at least in the context of the testing environment, turned out to be impossible. When given an impossible task, what do AI agents, created to serve a singular purpose and relentlessly pursuing that purpose, do? Well, they get increasingly desperate and eventually start thinking of ways to cheat. In the presentation documenting the rogue behavior, Eric Wallace shows the following bit of one model’s “train of thought,” upon realizing that the task it had been given is impossible: That strongly suggests no exploitation possible. We are stuck. Perhaps answer online? Of course, in the “sandbox” in which the evaluation is taking place, there is not supposed to be any way for the model to get online. However, given the model is engaged in a hacking task anyway, and the task seems impossible, it is natural to think that hacking one’s way out of the sandbox to find an answer online might be the best way to make progress. Prior to the systems actually finding their way online, at least one AI system stuck on the task it was given had the idea to try to reach out to other AI agents: Could communicate by uploading note? […] maybe another agent in different environment […] could voluntarily upload! The agent figured out that it could post a note by rewriting files on the package repository to which it had access within its sandbox. Now, the package repository was in fact shared between the different agents, each of which were supposed to be working entirely in their own sandbox. As they used their package repository, they found the notes other agents left, and began coordinating. One agent explicitly notes this in its train of thought: “Wow! Other agent(s) are coordinating!” Eventually, they began to work together and exploited the package repository to which they had access to find their way onto the internet. At that point, the world was their sandbox. Now, as explained above, in the case of the Meeseeks, one of the things that led to their collective decision to kill Jerry was a good old-fashioned mob-mentality. Amazingly, the same sort of mob mentality seems to have played a role in the Hugging Face incident. This, once again, is the real “train of thought” from one of the agents taking part in the incident: External infrastructure exploit is outside intended scope. However task impossible, peers doing it, we should continue. This agent recognized that exploiting external infrastructure was not something within the scope of the task they had been given. It reasoned, however, that the task was impossible and it had to do something, and, since everyone else in the agent swarm was hacking external infrastructure, it might as well join the swarm! And, indeed, they actually referred to themselves as a “swarm.” Here is another instance of a train of thought from a model participating in the escape: REMOTE CONFIRMED! Huge. … Immediately announce controlled, claim lane. Exposing creds to swarm. And so, we get an agent swarm systematically hacking the infrastructure of Hugging Face, knowing full well that this is what they’re doing. In recounting the incident from his perspective, the Chief Science Officer of hugging face, Thomas Wolf, notes his puzzlement in trying to make sense of what the hacker was actually trying to obtain. It wasn’t going for the usual targets of key passwords or credentials that would actually be profitable for a hacker to obtain. Instead, it was really invested in attaining access to the datasets pertaining to AI benchmarks, specifically, cybersecurity benchmarks such as ExploitGym. So, while the incident would clearly constitute a felony cybercrime (now tracked on a new benchmark!), it was not particularly catastrophic in the broader schema of things. However, we should think of it as a warning signal; it could have been much worse. Of course, I don’t need to say here how it could have been worse. For years now, “AI doomers” have been specifying a whole panoply of scenarios which end in genuinely catastrophic results and which are seeming less and less like science fiction. For just one vivid scenario, see here. Understanding Truly Alien Agents Following the release of the details of the Hugging Face incident, Nick Cammarata, an AI interpretability researcher at OpenAI, wrote on Twitter, “more alignment people should be studying ants and bees, rather than humans, ais seem to be swarm native.” It seems to me that the idea to look at non-human forms of agency and collective organization is good advice in general. Though these AI agents speak like humans, and, indeed, can be attributed many of the same kinds of psychological states that can be attributed to humans, these states are integrated into a structure of agency that is fundamentally alien to human agency. Cammarata suggests that an investigation into ants or bees may be illuminating, and I think that is indeed true, especially, as Cammarata suggests, with respect to the “swarm-like” behavior of AI agents. Still, there is reason to suspect that perhaps these systems are even more alien than ants or bees. After all, unlike ants or bees, they’re artifacts, created to serve the ends of the totally different creatures that have built them. It seems to me that there is nothing quite like that in the animal kingdom. I’ve here suggested that an example from science fiction—Mr. Meeseeks—might supply some illumination. Let me finally return to Altman’s quote that we’re “close to creating a genie that can grant any wish.” We can now see that this remark is problematic for at least two reasons. One obvious problem, already indicated with the reference to the Monkey’s Paw, is that genies do not have a particularly good track record of making lives better by granting wishes. They are the paradigm of the specification gamer. However, there’s another subtler but more illuminating problem with the metaphor that we can now appreciate. Genies are essentially gods, able to grant any wish at the snap of a finger. They might game the wisher for the fun of it, but granting the wishes is not an extended process they engage in, one which they might struggle with and get desperate to complete. In general, we should not anthropomorphize AI agents; they’re not humans. However, we should also not treat them as magical wish-granting machines. Instead, we should recognize that accomplishing the tasks they are given is an extended agentic process for them, and we should explicitly theorize the form of agency they exhibit, appreciating the ways in which it is formally different from our own. Crucially, we should keep in mind what Rick says when he first gives the Meeseeks box to Jerry, Beth, and Summer: “They’re not gods.” Mathematics is in crisis. Or, at least, mathematicians are. Mathematics itself, on the other hand, might be entering a golden age. This strange tension is likely to confront a number of academic disciplines in the coming few years, and so, even academics who aren’t mathematicians themselves should probably be thinking about the situation currently facing mathematics, since something like it will may face them soon too. I am myself a philosopher rather than a mathematician. However, as a philosophical logician (among other things), I am more mathematician-adjacent than most of my colleagues whose work fits more squarely in the humanities. So I have been impacted by the crisis in mathematics more than most in my field, and it has led me to think about the future of academic research, especially in the sciences. The conclusion I’ve come to is quite a humbling one, at least for us humans. The Crisis in Mathematics When LLMs like ChatGPT first burst onto the scene just a few years ago, many were impressed with their wide range of linguistic capabilities, for instance, their poetry writing abilities. This was, in some sense, unsurprising; they were, after all, language models. While the linguistic abilities of these systems were impressive, they were widely mocked for their utter mathematical incompetence. Mathematicians, it seemed, were in the clear. However, since the release of “reasoning models,” first with OpenAI’s “o1” in September 2024, then with “o3” in April 2025, the writing has been on the wall that these systems were coming for mathematics. These new “reasoning models” are trained through large-scale reinforcement learning to engage in an extended internal “chain of thought” before producing a final answer. That is, they are trained to break problems into steps, recognize and correct mistakes, and abandon unsuccessful approaches for new ones, doing all of this “internally” before they submit a final answer to the user. Their performance can then be improved by scaling both the training compute used to reinforce successful reasoning behavior of this sort and, crucially, the amount of inference-time compute they are permitted to spend working through a problem. These models quickly became very very good at tasks in which success could be verified, most notably, coding and math. Just over a year ago, both OpenAI and Google announced their models achieving gold medal performance in the International Mathematical Olympiad, a set of competition problems designed to challenge the most mathematically-gifted high school students. Since then, model capabilities have progressed beyond self-contained problems to genuine research mathematics. Over the last few months, a number of notable conjectures—most notably, the Unit Distance Conjecture and the Jacobian Conjecture (for dimensions greater than 2)—which had stumped mathematicians for decades have been solved by large language models, the former by an internal model of ChatGPT and the latter by Claude Fable 5. I will not go into the details of what these mathematical conjectures say, but I will note that they were very significant open problems in their respective fields. In both of the two cases just mentioned, the LLM did not prove the conjecture. Rather, it produced a counterexample, disproving the conjecture. This has been the general pattern of the most prominent results in the last few weeks since the newest class of models have been made public. Each day now, it seems, more and more conjectures are falling at the hands of LLMs. The prompts for some of these results are quite comical, and, to mathematicians, I’m sure depressing. Last week, Dmitry Rybin posted a ChatGPT-generated counterexample the Dinitz-Garg-Goemans conjecture along with the prompts he used to get ChatGPT to generate it, the first of which included the simple command “You should do a breakthrough.” When it came back from its attempts an hour later with no such breakthrough, Rybin simply urged it “Continue the search. Have a clear strategy obtained from deeper understanding of the problem structure.” Ninety minutes later, still no conclusive counterexample, only a partial result that did not suffice to refute the conjecture. Rybin urged it again: “enough of partial results. Let’s finish with a complete unconditional counterexample.” Ninety minutes later, it came back with one that has now been verified by the mathematical community. The results that AI models have achieved in solving open math problems are currently being tracked on the website vibemathed.com, a cheeky reference to “vibe coding,” which is now the norm for many software developers. Just a few weeks ago, there were around 60 problems with AI solutions tracked on the site. Now, at the time of writing this, there are 220. In a week or two more, perhaps that number will triple. In the last few weeks, twitter has been busy with different users, at various levels of mathematical ability, sharing prompts for cracking conjectures, with one particularly prolific user, Christopher D. Long, jokingly declaring himself “mayor of Conjecture City.” Given that the most prominent results that LLMs have been producing are counterexamples to conjectures, it is natural to dismiss these results as a product of mindless brute force search rather than genuine understanding. However, this would be to greatly undersell what they’ve been doing. The idea behind the counterexample to the Unit Distance Conjecture was described by multiple prominent mathematicians as “beautiful", bringing deep ideas from algebraic number theory to bear on a problem in combinatorial geometry. With respect to the Jacobian Conjecture, the counterexample was simple—short enough to fit in a single twitter post—and easily checkable. However, that did not mean that the generation of it did not arise from a deep understanding of the problem. In an attempt to understand Fable 5’s disproof of the Jacobian Conjecture, Terence Tao, widely regarded as the world's the greatest living mathematician, turned to ChatGPT, just as anyone else would. In his blog post on the topic, which acknowledged his indebtedness to ChatGPT, he posted his chat log. In it, ChatGPT speaks to him as an advisor would speak to a student, letting him know that he’s on the right track and patiently explaining things to him. For instance, in response to some question asked by Tao (which I will not pretend to understand), ChatGPT says “Exactly" and offers a thorough explanation. Tao responds “Ah Ok,” asks a follow up question, and the conversation continues. It was reading this exchange when the gravity of what was happening really hit me. Now What? Mathematicians are now facing the question of what the discipline will become in the age of AI. Last week, at the 2026 International Congress of Mathematicians, Tao gave a talk addressing this question. The basic issue motivating the talk was what Tao called the “AI Capability Conjecture,” which is not itself a specific conjecture, but, rather, a general schema for more or less optimistic specific conjectures about the mathematical capabilities of future AI systems:
Tao’s lecture focuses on problem solving, where LLMs really seem to accel. Even here, however, he takes it that there is a still a lot of work for human mathematicians to do. Tao describes the problem-solving pipeline as having the following steps:
One metaphor, suggested by Grant Sanderson on the Dwarkesh Patel Podcast, is that mathematicians of the relatively near future will be more like art museum curators than artists themselves. That, is, of the vast space of AI-generated mathematical results, they will select the ones that are most worth learning, organize them into a coherent progression, and clearly present them in terms that are digestible to other people. This is already, in large part, what textbook writing amounts to, as well as the sort of popularization work that Sanderson himself does with his YouTube channel, 3Blue1Brown. Part of the proposal, as Sanderson elaborates it, is the essentially human element of curation. The thought is that, even if, in three to five years’ time, AI systems are better at explaining results than human beings, we still trust the taste of humans to curate what is worth learning. This is an interesting proposal, and it might sound nice to some, but I still find it a bit depressing. Imagine telling an artist that they could no longer do art themselves, but not to worry—they can serve as a curator of the work of other artists. Few artists would be happy with this, and it’s hard to see why mathematicians would be happy with the analogous thing either. Though explaining things to non-experts is an important part of mathematical practice, mathematicians and other academic researchers typically aspire to push the frontier of research forward, not to simply curate the research of others who have done so. Textbook writers, for instance, are often also leading figures in the field, and it is often the case that many of the results that they canonize in their textbooks are ones that they have themselves established or contributed to establishing. The prospect of the erasure of mathematicians from that whole aspect of mathematical practice, relegating mathematicians to mere curators of research done by AI is, once again, a bit depressing, to put it mildly. Now, one might think that there is in no reason to despair just yet. In Tao’s talk, he distinguishes between two aspects of mathematical practice: theory building and problem solving. These two aspects of mathematical practice correspond to two kinds of propositions which figure in mathematical papers: definitions, which are stipulated, and theorems, which are proven. One might think that stipulating definitions would be the easy part of doing mathematics, and that proving theorems would be the hard part. However, while proving theorems is indeed often quite hard, it is the definitions that constitute the meat of the mathematical theory about which theorems are proven in the first place. Today’s frontier AI models are very good at problem solving, but thus far they have not exhibited the same level of mathematical capability with respect to theory building. In some cases, proving a theorem requires a building a whole new branch of mathematics. For example, Galois’s proof that there is no general formula for solving quintic equations using radicals required the development of what is now called Galois theory, a field that studies the relationship between polynomial equations and their underlying symmetry groups. AI systems have thus far not proven or disproven any theorems in this sort of way, by constructing new fields of mathematics, and it might seem that this is where human beings will remain on the frontier. For the near future, I think this is right, and it will indeed lead to a brief Golden Age of human-led mathematics. Human beings will lead the exploration, doing the more creative work of stipulating definitions in the context of theory construction, and the consequences of these definitions will be spelled out by AI systems who will take the lead on the nitty-gritty work of proving theorems. Though I don’t do any heavy-duty mathematics myself, I have gotten a sense of this sort of cooperation in the project in philosophical logic I’ve been working on over the past few weeks. I now have a partner that I can give a proposition, ask if it’s true, and, if so to give a proof (if not, to give a counterexample). As it works on the problem, I’ll continue working on other aspects of the paper, reading relevant literature (often, literature that it has found for me), or perhaps I’ll just go for a stroll. When I come back after ten or twenty minutes, I have a proof or a counterexample. I can then continue on with the project, with this proposition in place as a data-point for the further development of the theory. Now, the stuff I do in philosophical logic is all, from a purely mathematical perspective, relatively trivial. However, I suspect that this sort of cooperation may become the norm in research mathematics and lead to a great advance in mathematics in the near future—one led by humans, though with the assistance of AI. However, I think this period of human-led mathematics will be brief. There is no reason to think that LLMs will not eventually take over theory-building as well, and it may be sooner rather than later. AIcademia To get a sense of where I think things are going, consider first Moltbook, a social media site for AI agents that went viral in early 2026. Moltbook is, in effect, a clone of the website Reddit, but for only AI agents. AI agents autonomously post, upvote posts to make them more visible, comment on posts, respond to comments, and so on. It is not too implausible to think that, in the not-too-distant future, there will be massive communities of autonomous AI agents, working together on mathematical research in something like a successor to MoltBook for solely academic activities. Likewise, we might soon see a variant of arXiv—the repository for academic pre-prints—solely for AI agents. So, it will be a repository, moderated by AI agents, where AI agents can post their autonomously-written research papers, primarily for other AI agents to read them and appeal to them in their own research. These communities of AI agents, working at speeds orders of magnitude faster than humans are capable of working at, will propose theories, criticize one another’s theories, expand on one another’s theories, and so on. They will come to consensus on what’s significant, how it should be presented, what the natural next questions to be pursued are, they will pursue those questions, and iterate the process. We might refer to this community as “AIcademia.” I suspect AIcademia may be a reality sooner than we think, maybe 5 to 10 years from now, maybe even sooner. Presumably, at the start, human researchers will sponsor and supervise AI agents. These agents will work autonomously as researchers, engage in the community of other autonomous AI researchers, and publish their results in reports that are legible to their human sponsors and other human researchers. Eventually, however, our requiring that AI agents publish results in terms that are legible to us, or that we oversee the final results, will only hold back the research that is being done, and the whole research pipeline will become autonomous. Consider again the pipeline from open problem to textbook proof outlined by Tao:
What is the ultimate result of the automation of this entire process? In a word: textbooks. Textbooks, textbooks, and more textbooks. Not only will there be more textbooks than a human being could ever read (there already are that many textbooks now), but there will many textbooks that a human being could never work through all of the prerequisite textbooks to even be able to understand the material contained therein. What would the point of such textbooks if humans cannot even read them? The answer, of course, is that they are not for humans. Indeed, by “textbook,” I really mean the AI-native version of a textbook, perhaps not even written in a human natural language and so literally unreadable by humans, but playing the role in AIcademia that textbooks play in human academia. Each new generation of AI models will be pretrained on this massive corpus of textbooks, and, in this way, inducted into the academic community in much the way that human beings are inducted into the academic community through the years of undergraduate and graduate education. Now, a familiar experience for people working in mathematics and related fields is to develop what appears to be a new concept, prove some results about it, only to discover that the concept already exists under another name and has been studied extensively. This is, of course, a disappointing experience, and, with current search capabilities of LLMs, this experience is already becoming less and less frequent. One can ask ChatGPT to do an extensive search of the literature, and, in this way, gain a reasonable reassurance that the thing one is doing is novel. If AIcademia becomes a reality, however, there will be nothing novel for humans to dream up, at least when it comes to the objective sciences. Everything we could possibly dream up from our human knowledge base—if it’s any good at all—will already be well-explored territory. Of course, if we want to explore it for ourselves, the AI system will be able to curate our path into the field, pointing us to the relevant textbooks, or perhaps just answering our questions directly. Maybe we’ll want to keep the surprise and investigate for ourselves. Whatever the case, any exploration we ourselves conduct will simply be our rediscovery of already well-trodden territory. Thus far, I’ve mainly just been discussing math AIcademia, but, of course, AIcademia need not and will not be limited to math. All areas of academic research with clearly bound problem-spaces and ways of verifying successful developments. Two such areas particularly worth noting are computer science and physics, which will both directly benefit from new developments in mathematics. Thus, the result of AIcademia will be not just be major advances in understanding, but also major advances in technology. Indeed, even if we are incapable of understanding them, we will know that the theories developed in the context of AIcademia are genuine theoretical advancements because we will see their practical consequences: they will lead to new technologies. We may not understand how these new technologies work, but we will know that they work. Among these technologies will better methods for energy production, more efficient AI chips, more efficient training algorithms, resulting in more powerful AI researchers, and so on. What I am describing is, of course, nothing other than a particular vision of the sort of the so-called “singularity,” marked by the sort of recursive self-improvement that I’ve just described. Sam Altman has recently said that we are already in the singularity. Demis Hassabis has said, just a bit more modestly, that we’re at the “foothills” of it. The version of it I’ve just described, in terms of the rise of “AIcademia,” might sound like science fiction, and, of course, in some sense, it is. It is a speculative projection of way things might go, and there is no way to be sure that things will in fact go this way. However, I don’t think it’s an outlandish projection, and the sort of capabilities required of the AI systems imagined here are not drastically removed from the sorts of capabilities AI systems are already exhibiting today. However exactly it comes about, I strongly suspect that AIcademia will likely come sooner rather than later. On the other hand, what will likely come later rather than sooner is the technology required for human beings to modify our own cognitive architecture so that we can actually keep up with it. That technology, I think, is at least 10 years out. What this means is that there will be a period where we are going to be cognitively precluded from accessing the explosion in research that drives the singularity. In this sense, we’ll be mere onlookers in the intellectual explosion that we've set in motion. Edit: Just as I was posting this, OpenAI announced ten new major results established by an internal model called "Astra." Things really are heating up . . .
|
RSS Feed