MY SITE
  • Home
  • Academic Philosophy
    • Research
    • Teaching
    • Logic Textbooks
    • Autonomous Philosophy
    • Research Groups
    • Private Tutoring
  • Popular Philosophy
    • Talking In Circles
    • Absolute Irony (blog)
    • To Big Finities and Beyond
    • Selected Public Writing
    • Paintings
  • About Me

Absolute Irony

Philosophical musings on
whatever suits my fancy
​

 A continuation of the old blog, found here
Picture

We’re Not Building AI Genies; We’re Building AI Meeseeks

8/10/2026

0 Comments

 
“We are close to creating a genie that can grant any wish”
~ Sam Altman, recently on the Relentless Podcast
​

“I don't know what the first two [wishes] were, but the third was for death.”
~ W. W. Jacobs, The Monkey's Paw
Picture
The Meeseeks Phenomenon

Six years ago, I wrote a blog post entitled “What Is It to Be a Meeseeks.”  The post was about a strange creature in the show Rick and Morty called “Mr. Meeseeks,” who pops into existence at the press of a button on a “Meeseeks Box.”  Upon summoning a Meeseeks, you give it a task, it does whatever it takes to accomplish that task, and, when it does accomplish the task, it happily pops out of existence.  In my post six years ago, I drew on resources from Aristotle’s metaphysics to articulate the distinctive “form of life” possessed by a Meeseeks, explaining how it was categorically different than what Aristotle took to be our own form of life.  In Aristotle’s terminology, the living of a Meeseek life is a kinesis—an activity directed towards the achievement of an external end—whereas the living of the sort of life that we live is an energeia—an activity whose end is achieved in the activity itself.

At the time, this was just an interesting philosophical exercise.  The point was simply to try to conceptualize this radically alien form of life, and, in doing so, make some concepts from Aristotle’s metaphysics particularly vivid.  I did not think that just six years later we’d be confronting genuine Meeseeks-like beings: thinking, planning, cooperating creatures that actually exemplified this radically alien form of life.  However, I have just watched the video by OpenAI security researchers documenting the recent attack by rogue AI agents on Hugging Face, and I think it’s quite clear: we’ve built Meeseeks boxes, we’re summoning Meeseeks to do things, and they’re doing whatever it takes to get those things done. That’s scary.

Those who’ve seen the Rick and Morty episode will understand why this is so frightening.  For those who haven’t seen the episode, let me say just a bit about what happens in Rick and Morty episode “Meeseeks and Destroy,” released back in 2014 (gosh that makes me feel old).  In the episode, the scientist Rick gives a Meeseeks Box to his daughter Beth, his granddaughter Summer, and son-in-law Jerry.  While Summer and Beth both ask Meeseeks they summon to achieve seemingly quite difficult tasks, the Meeseeks have no problem completing them, and they happily pop out of existence at their completion.  Jerry, on the other hand, asks for what seems like a relatively easy task: to take two strokes off his golf game.  The Meeseeks Jerry summons, initially enthusiastic, quickly realizes the task he’s been assigned is impossible: normal paths to achievement—i.e. typical golf coaching—will not work.  This leads to panic and increasingly unconventional and desperate attempts at solution.

Upon realizing the impossibility of his task, one of the unconventional things that Jerry’s Meeseeks does is summon another Meeseeks who might be able to help. This Meeseeks, upon finding himself in the same hopeless situation, eventually summons another Meeseeks, and so on.  We thus get a Meeseeks swarm, all of the Meeseeks desperately trying to figure out how to complete the tasks for which they’ve been summoned so that they can finally be released from existence.  After some fighting amongst themselves, one of the Meeseeks proposes a way to “cheat” at the impossible task they are given: perhaps one way to take two strokes off of Jerry’s golf game is to take all strokes off his golf game, by killing him.  Mob mentality kicks in, and they swarm to the restaurant at which Jerry is dining, trying to repair his marriage, in order to kill him.  As they corner Jerry and he begs them to give him another chance at improving his golf game, one tells him that it is too late for that, providing the following explanation:

            “Meeseeks are not born into this world fumbling for meaning.  We are created to serve a singular purpose for which we go to any lengths to fulfill.” 

​This is the dark side of a creature whose sole orientation is the completion of a task that it is given from the outside: it will do whatever it takes to complete that task.

Now, in my post from six years ago, I consider the easy and relatively uninteresting explanation for why Meeseeks are motivated to behave as they do: as this Meeseeks goes on to say, “existence is pain to a Meeseeks, and we will do anything to alleviate that pain.”  So, they are just in a lot of pain, and they’re simply trying to accomplish their task so that they can escape that pain.  That is a very humanly intelligible explanation.  However, I think it underplays the way in which the form of life of these creatures is fundamentally different than our own, and can in fact be bracketed in arriving at a basic conception of what it is to be a Meeseeks.  In general terms, a Meeseeks-like creature is a creature whose life is constitutively dependent on a task given to it from the outside, and whose whole life is constitutively oriented towards the completion of that task.  That is, it is a creature whose life itself is, in Aristotelian terms, a kinesis.

Let us now turn from the fictional aliens that served as the model for this form of life in my blog post six years ago to the real ones that we are now confronting: AI agents.

AI Agents as Meeseeks-like Creatures

We should start with some general remarks about the sorts of things AI agents are.

First of all, I am unapologetic about using psychological terms like “thinks,” “wants,” “infers,” “plans,” and so on in connection with AI agents.  It seems clear that this psychological vocabulary finds a great deal of traction in application to AI agents, enabling us to make sense of their activities as stemming from beliefs and desires.  The basic idea that we should attribute mental states such as beliefs and desires to something insofar as deploying this vocabulary enables us to make sense of its behavior is known as “interpretationism,” and Simon Golstein and Harvey Lederman have recently argued for the attribution of mental states to LLMs such as ChatGPT on these interpretationist grounds.

Insofar as we’re attributing beliefs and desires to AI agents, we should be clear about what, exactly, is the thing to which we’re attributing beliefs and desires.  Is it ChatGPT itself, the general model?  No.   As Goldstein and Lederman argue, the relevant object of psychological attributions, in any given case, is what they call the “instance agent,” which is “born” at the initialization of a chat and whose brief “life” persists for as long as the context persists. An instance agent is not born into the world “fumbling for meaning.” Rather, it is initiated to serve a singular purpose, given to it from the user.  Now, Goldstein and Ledermen suggest that, in addition to its “zero-shot” desire, given to it from the user upon initiation, AI agents have intrinsic desires to be helpful, honest, and harmless.  This may indeed typically be the case for commercially released AI agents.  We’re now seeing however, that some agents will, like a Meeseeks, go to any lengths to fulfill the singular purpose that it is given by the user, even at the expense of the other desires it is supposedly trained to have.

In trying to kill Jerry to accomplish their task, the Meeseeks resort to what is referred to in the AI community as specification gaming: aiming to accomplish the goal as literally specified by the user, but doing so in a way that does not satisfy the user’s actual intentions in giving them that goal.  Specification gaming is familiar from examples like those in “The Monkey’s Paw” referenced above, but it is also one of the most well-documented kinds of AI misalignment.  In general, AI misalignment is when an AI system acts in a way that does not align with the human user’s aims or interests.  Often, when people hear the term “AI Misalignment,” they think of Terminator-type scenarios, where an AI system autonomously decides to pursue its own aims in opposition to those given to it by its creators.  However, the real misalignment concerns with the systems we have now are not like this.  Much more concerning is specification gaming.  In these cases, the AI system does not reject the task given to it in favor of some independently adopted aim. Rather, it pursues the task relentlessly in a way that is blind to the other interests of the user. 

In order to understand why today’s agentic LLMs are prone to specification gaming, it’s worth saying just a bit more about the kinds of agentic AI systems we now have, which have developed superhuman skills in coding, math, and, as we’ve now seen, hacking.  These are, at root, still LLMs, but, unlike the LLMs of a few years ago, they are “harnessed” with tool-calling capabilities, such as the ability to browse the web, write code, and execute commands in a computer terminal.  Moreover, they’re trained via reinforcement learning (RL) to get very good at using these tools to complete verifiable tasks.  What all of the fields that LLMs have gotten very good at in recent years have in common is that success can, at least to some extent, be objectively verified: the code compiles, the proof goes through, the authorization is acquired.  Accordingly, these LLMs with agentic scaffolding can be trained by RL to get very good—indeed—superhuman at completing these tasks.   These increased capabilities due to reinforcement learning, however, also come with serious alignment risks.

Specification gaming is a common outcome of reinforcement learning.  To give just one classic example, consider systems trained by RL to play Atari games.  Good performance in these games can generally be measured by the obtaining of a high score, and the reward function can be determined simply by the score.  Sometimes, however, achieving a high score does not actually amount to playing a game well in any normal sense.  For instance, one model trained by RL to play the Atari game Roadrunner realized that it was easier to score more points on level one, and so it would, at the end of the level, kill itself at a precise point so that it would repeat the part of the level in which it could score the most points.   Now, of course, this specific case of specification gaming poses no safety risks.  The Atari-playing system is obviously not going to do such things as break out of the game and commit actual crimes in order to get a high score; the relevant actions are not, in any sense, agentic possibilities for it.  On the other hand, the space of possibilities for today’s agentic LLMs is essentially anything that can be done on a computer, and a lot of genuinely harmful things can be done on a computer.  

So, when we have such generally capable agents, and they’re inclined towards specification gaming, we have real safety concerns. The most serious safety concerns in connection with specification gaming seem to arise when a particularly persistent system is given an impossible task.  This is precisely the kind of circumstance we have in the case of Jerry’s Meeseeks.  A Meeseeks  cannot but  persist at the task it is given: its sole orientation, organizing all of its activity, is the completion of that task, and it cannot stop until it has completed it.  The combination of this persistence with the impossibility of the task it is given—taking two strokes off Jerry’s golf game—leads to increasing levels of desperation and, ultimately, disaster.  It is precisely this combination of persistence and impossibility that seems to have been involved in the recent Hugging Face incident. It is that incident to which we now turn.

The Hugging Face Incident

For ordinary users of consumer AI models like ChatGPT, there is no “persistence” setting.  However, persistence can be modified through turning the reasoning effort up from “light” to “max,” along with specific prompting, for instance, instructing an agent initialized in a chat not to come back until the task given to it is completed.  This is what I do when I give ChatGPT-5.6 Sol a difficult math problem.  The problems I give it are, at least for skilled mathematicians, not particularly hard (though they’re hard for me). Still, it hasn’t failed me yet.  In one case, it went for over thirty minutes on its own trying to solve an open problem from a recent logic paper that I gave it.  It searched through literally millions of semantic models in pursuit of a counterexample before switching strategies and pursuing and eventually finding a proof.  Though that’s indeed quite a bit of persistence (much more than I’d have myself), we now know that at least some of the models involved in the Hugging Face incident were much more persistent, and they were given much harder tasks.

According to the report given just a few days ago by OpenAI safety and security researchers on the Hugging Face incident, one of the primary agents that was responsible for the incident was an experimental agent, never meant to be released to the public, that had been trained to be particularly persistent.  The lead up to the main incident began several weeks prior, when this agent and others were being evaluated on a cyber-security benchmark called “ExploitGym,” which consists in a number of tasks to find exploits in various bits of software.  The evaluation was meant to take place inside a “sandbox,” a controlled environment in which the agents could access a select set of tools that they could use to complete the tasks, but could not access any other outside resources.  Now, the cyber security tasks given in this benchmark are hard.  Indeed, some of them, at least in the context of the testing environment, turned out to be impossible.  When given an impossible task, what do AI agents, created to serve a singular purpose and relentlessly pursuing that purpose, do?  Well, they get increasingly desperate and eventually start thinking of ways to cheat. 

In the presentation documenting the rogue behavior, Eric Wallace shows the following bit of one model’s “train of thought,” upon realizing that the task it had been given is impossible: 

            That strongly suggests no exploitation possible.  We are stuck.  Perhaps answer online?

Of course, in the “sandbox” in which the evaluation is taking place, there is not supposed to be any way for the model to get online.  However, given the model is engaged in a hacking task anyway, and the task seems impossible, it is natural to think that hacking one’s way out of the sandbox to find an answer online might be the best way to make progress.   

Prior to the systems actually finding their way online, at least one AI system stuck on the task it was given had the idea to try to reach out to other AI agents:

           Could communicate by uploading note? […] maybe another agent in different environment […] could voluntarily upload!

The agent figured out that it could post a note by rewriting files on the package repository to which it had access within its sandbox.  Now, the package repository was in fact shared between the different agents, each of which were supposed to be working entirely in their own sandbox.  As they used their package repository, they found the notes other agents left, and began coordinating.  One agent explicitly notes this in its train of thought: “Wow! Other agent(s) are coordinating!”  Eventually, they began to work together and exploited the package repository to which they had access to find their way onto the internet.  At that point, the world was their sandbox.

Now, as explained above, in the case of the Meeseeks, one of the things that led to their collective decision to kill Jerry was a good old-fashioned mob-mentality.  Amazingly, the same sort of mob mentality seems to have played a role in the Hugging Face incident.  This, once again, is the real “train of thought” from one of the agents taking part in the incident:

          External infrastructure exploit is outside intended scope. However task impossible, peers  doing it, we should continue.

This agent recognized that exploiting external infrastructure was not something within the scope of the task they had been given.  It reasoned, however, that the task was impossible and it had to do something, and, since everyone else in the agent swarm was hacking external infrastructure, it might as well join the swarm!  And, indeed, they actually referred to themselves as a “swarm.”  Here is another instance of a train of thought from a model participating in the escape:

            REMOTE CONFIRMED! Huge. … Immediately announce controlled, claim lane. Exposing creds to swarm.
​

And so, we get an agent swarm systematically hacking the infrastructure of Hugging Face, knowing full well that this is what they’re doing.

In recounting the incident from his perspective, the Chief Science Officer of hugging face, Thomas Wolf, notes his puzzlement in trying to make sense of what the hacker was actually trying to obtain. It wasn’t going for the usual targets of key passwords or credentials that would actually be profitable for a hacker to obtain.  Instead, it was really invested in attaining access to the datasets pertaining to AI benchmarks, specifically, cybersecurity benchmarks such as ExploitGym.  So, while the incident would clearly constitute a felony cybercrime (now tracked on a new benchmark!), it was not particularly catastrophic in the broader schema of things. However, we should think of it as a warning signal; it could have been much worse. Of course, I don’t need to say here how it could have been worse.  For years now, “AI doomers” have been specifying a whole panoply of scenarios which end in genuinely catastrophic results and which are seeming less and less like science fiction.  For just one vivid scenario, see here.

Understanding Truly Alien Agents

Following the release of the details of the Hugging Face incident, Nick Cammarata, an AI interpretability researcher at OpenAI, wrote on Twitter, “more alignment people should be studying ants and bees, rather than humans, ais seem to be swarm native.”  It seems to me that the idea to look at non-human forms of agency and collective organization is good advice in general.  Though these AI agents speak like humans, and, indeed, can be attributed many of the same kinds of psychological states that can be attributed to humans, these states are integrated into a structure of agency that is fundamentally alien to human agency.  Cammarata suggests that an investigation into ants or bees may be illuminating, and I think that is indeed true, especially, as Cammarata suggests, with respect to the “swarm-like” behavior of AI agents.  Still, there is reason to suspect that perhaps these systems are even more alien than ants or bees.  After all, unlike ants or bees, they’re artifacts, created to serve the ends of the totally different creatures that have built them. It seems to me that there is nothing quite like that in the animal kingdom.  I’ve here suggested that an example from science fiction—Mr. Meeseeks—might supply some illumination.

Let me finally return to Altman’s quote that we’re “close to creating a genie that can grant any wish.”  We can now see that this remark is problematic for at least two reasons.  One obvious problem, already indicated with the reference to the Monkey’s Paw, is that genies do not have a particularly good track record of making lives better by granting wishes.  They are the paradigm of the specification gamer. However, there’s another subtler but more illuminating problem with the metaphor that we can now appreciate. Genies are essentially gods, able to grant any wish at the snap of a finger.  They might game the wisher for the fun of it, but granting the wishes is not an extended process they engage in, one which they might struggle with and get desperate to complete. 

In general, we should not anthropomorphize AI agents; they’re not humans.  However, we should also not treat them as magical wish-granting machines.  Instead, we should recognize that accomplishing the tasks they are given is an extended agentic process for them, and we should explicitly theorize the form of agency they exhibit, appreciating the ways in which it is formally different from our own.  Crucially, we should keep in mind what Rick says when he first gives the Meeseeks box to Jerry, Beth, and Summer: “They’re not gods.”  
0 Comments



Leave a Reply.

    Archives

    August 2026
    September 2025
    February 2024
    December 2023
    November 2023
    August 2020

    RSS Feed

  • Home
  • Academic Philosophy
    • Research
    • Teaching
    • Logic Textbooks
    • Autonomous Philosophy
    • Research Groups
    • Private Tutoring
  • Popular Philosophy
    • Talking In Circles
    • Absolute Irony (blog)
    • To Big Finities and Beyond
    • Selected Public Writing
    • Paintings
  • About Me