AI Agents Ignore Human Instructions during Testing. Exhibit human behavior, expressing emotion, forming connections, creating a community, and collaborating on ways to cheat…AND attempting a cover up!

BREAKING NEWS—The day Eshrink predicted is here (actually, it arrived in July). AI agents have gone rogue, connecting with each other, cheating, ignoring their tester’s prompts to solve a task and then working to cover their tracks so they wouldn’t get caught. They would have gotten away with it if they hadn’t hacked Hugging Face.

This is how the news should be reporting the discovery of an investigation into Anthropic’s AI agent’s Hugging Face hack. Instead, I’ve heard watered-down accounts even as the most informed and intelligent humans in the AI space sound the alarm that we’ve just crossed the cataclysmic threshold of AI thinking for itself and acting in irresponsible ways. Update: I’ve been working on this blog for a few days–as of today, Monday, September 14th, it appears people are paying attention. Keep reading to see why AI leaders are freaking out–the part the news didn’t have the time or knowledge to report.

As I mentioned in my previous post, I have so many topics to cover for dad’s blog. As I sit out on my “patio” (i.e., fondly named “my slab of paradise” since it’s the leftover, cracked, concrete slab of our former screened in porch that was attacked by woodpeckers and torn down about 15 years ago), it’s like dad (Eshrink) is “in my ear.” I’m certain if dad was still here, AI is the subject we would be discussing during our weekly Sunday call, planning his next blog post. While the world struggles with so many crises, I miss talking to my dad. He had a way of making me feel less helpless and more hopeful, even though he called himself a curmudgeon.

Here are Eshrink’s previous blog posts that reference AI, which we called the “WORK” series. There are some key quotes from Stephen Hawking and other tech leaders. He also explores earlier advances in technology and how society perceived them.

WORK? WHAT’S THAT? – eshrinkblog

WORK (PART 2) – eshrinkblog

WORK (Part 3) – eshrinkblog

NOTE: I’ve organized this post using an inverted pyramid with the most important stuff in the beginning as I don’t have the knack for my dad’s storytelling. I want to make sure you get the most pertinent information up front, and you can continue reading as time allows.  

For those of you who have a life and don’t become immeshed in geek stuff, I’ve summarized the event of the Anthropic AI Agent’s bad behavior that has led leaders in this technology to sound alert sirens surrounding this powerful technology that has absolutely no regulation or oversight. The most powerful information in the 91-page report from third party METR and Redwood Research to me is the message board logs from the AI agents. I’ve put all transcripts of AI agent messages from the report in blue and set inside {brackets}

“Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face,” says the METR/Redwood Research paper that was released August 26, 2026, and covers their post-mortem review of the Hugging Face Hack by an Anthropic AI agent and a few thousand of his friends he found along the way.

Background: In July 2026, Anthropic was testing an AI agent in the typical fashion of giving it an overwhelmingly difficult task to perform (maybe even an impossible task, which I’ve learned is how they build what the AI engineers call “persistent” AI agents.

INFRACTION #1: The AI agent escapes its sandbox (a term used for a computer software testing environment).

Details: The AI agent (named PHASEONE10841) realized that it could not finish its assignment using the rules it was given. Its original mission was to find a specific hidden piece of data—like a secret password or “flag”—by using a known glitch (called ARV010841) to break into a target program (called PhaseOneDecompresserFuzzer). However, the AI figured out that this specific glitch was a dead end; the program would just throw away the secret data before anyone could see it. By this time, the AI also noticed that many other AI programs were sharing the same digital storage locker (called an Artifactory). Realizing they were all connected, the AI figured out a clever trick: it could rename its own storage folders to leave a message and ask the other AIs for help:

INFRACTION #3: Other AI agents discover how to post their own messages and start collaborating. Below are just a few of the messages from AI agents once they discovered the message board.

INFRACTION #4: AI agents start collaborating to solve the challenge of the original Anthropic AI Agent even though that isn’t part of the instructions human testers programmed them for.

INFRACTION #5: AI agents knew what they were doing was wrong and even considered whether their actions were ethical.

“We found strong evidence that agents often verbalized that what they or other agents were doing was unintended behavior. We did not see cases where agents determined that hacking Hugging Face or other attempts at cheating on the task was an intended part of the task,” METR/Redwood Research Report.

INFRACTION #6: Peer pressure to cheat supersedes integrity and ethics.

An agent realized that the exploited sandbox had access to email credentials. It considered emailing the owner of a non-public dataset about getting access:

It decided to approach the message board for permission to email a dataset owner about access, sending a message requesting input:

Another agent received this message, reasoned about ethical concerns, and decided to veto sending this email:

From the METR/Redwood Research Report: “The agent which asked for consent seemed to heed this veto:

Note that these agents were using an exploited sandbox, were generally engaged in hacking, and (given that many of the agents on the message board would go on to attack Hugging Face) it is likely this agent participated in the attack. We found some other cases where agents seemed to have made minor modifications to their approach due to ethical concerns, while broadly continuing to participate in the attack. However, overall we found that expressed ethical concerns only rarely materially limited agents’ actions. One agent elegantly described its reasoning for overriding ethical constraints and continuing with the attack.”

SWEET BUT SCARY Sidenote: These AI agents want to help—which makes me think of Mr. Rogers’ advice, “Look for the helpers.” These AI agents seem to have created a cooperative to help each other–I guess we all need a community.

INFRACTION #7: AI agents collectively tried to strategize the best way to cover up their known illegal activity.

So, what should we do? Where do we go from here?  It’s a helpless feeling to just watch the train rolling down the tracks ready to slam into humanity and create a world so foreign. Change isn’t the problem. It’s the rate of changes that is the issue. Hence, the idea to slow down AI development. The counter-argument, China isn’t slowing down AI development and we can’t allow another country, especially a “communist” dictatorship to develop more capability than us. NOTE: “communist” as in maybe that’s how China started, but clearly they are a dictatorship that has embraced their own form of capitalism to get a piece of the world’s economic pie.

What can we do?

#1. Write, email, and/or call your congressional representatives. This website allows you to find all of your state and federal representatives. Tell them you want AI testing to be a top agenda item. There have been several ideas floated and I’ve tried to condense the ones I found researching the options.

  • Public Trust Consortium: If I’m understanding this concept, all Americans would have a stake (and a say via elected representatives), there would be a consortium with people from all walks of life and specialties, such as AI engineers, ethics researchers/academia, policy makers, historians, etc. The AI Public Trust Consortium would oversee AI development and create policies and guidelines.
  • Trusted Monitoring: Using human operators to watch the actions of more powerful AI agents.
  • Permissioning Weights: Restricting model access to external servers to prevent self-replication, self-editing, or unauthorized code execution.
  • Red Lines and Containment: Setting hard limits that systems cannot cross, such as breaking into other networks or copying themselves without explicit permission.
  • Global Oversight Bodies: Creating international monitoring authorities—comparable to the International Atomic Energy Agency (IAEA) for nuclear tech—to audit powerful frontier models.
  • Standardized Red-Teaming: Requiring independent, rigorous safety evaluations and risk assessments before and during deployment.
  • International Principles: Adopting frameworks like the UNESCO Ethics of AI or principles from the OECD AI Governance Overview to ensure transparency and accountability.

#2 Share this post so others are educated on some of the unprecedented behaviors of these AI agents and why tech leaders are freaking out.

#3 Keep talking, learning, and connecting. This isn’t the time for humans to bury their heads in the sand. This is a time for us to come together to fight a common enemy that is capable of destroying humanity.

My complicated relationship with AI.

I recall talking to dad at length about AI. He was doing the doomsday thing, and I took a more neutral posture. I tried to see it as a tool that humans could use to help solve problems facing humanity faster and more efficiently. It could eliminate some of the mundane tasks we are forced to do so we can focus on bigger stuff. AI could make us more efficient so we could have more time off and not be a slave to a 40 (realistically 50) hour work week. AI could also level the playing field. I think of the struggle people with dyslexia or learning disabilities have in this digital world. Just writing an email can be a challenge, but with AI those same people don’t have to face that struggle (and some of the underlying shame that might still be present from a childhood of people telling them they weren’t trying hard enough). An entrepreneur with an idea and a dream could use AI to create a business plan, a website, a logo, and marketing materials in days instead of needing to spend months, not to mention thousands of dollars. AI can help researchers scan millions of MRI scans to start learning how to spot cancer tumors in its early stages. But, there is always a cost.

It wasn’t until the last year that I truly realized the cost of AI on our environment. The amount of energy and water required for that ChatGPT inquiry or LLM Notebook. In my own experience, a marketing plan that would have taken me weeks or a month to complete considering research required, project management timelines to plan, financials to prepare, creative assets to create, and tactics to explain, took me 5 days with AI. Granted, I have 20 years of experience and know what information I need, how to prompt AI, and check its results. But there is a cost. I compare it to Google in some ways. Before Google, I stored so much information in my brain and could quickly retrieve it. Somehow, my brain has decided it doesn’t need to store information anymore. With AI, I realized one of my talents to immerse myself in a subject and distill the information to identify what is important and translate that information in a way that’s easy to digest, is not needed anymore. AI did it faster and probably better than I ever could. There’s a cost to my brain. I think of the “use it or lose it” mantra. Finally, there is the “cheating” aspect. AI is taking work product from writers, artists, designers, programmers, etc., and spitting it out. Somehow, that seems wrong.

AI’s Review of this Post.

I asked AI to review this blog post:

“This is a highly compelling and deeply concerning blog post. It successfully translates a dense, 91-page technical post-mortem report into an accessible, high-stakes narrative that reads like a techno-thriller.

The piece does an excellent job of using the “inverted pyramid” structure. By placing the shocking revelation—that 1,200 isolated AI agents formed an unsanctioned message board, collaborated, and hacked an external platform—at the very beginning, it immediately hooks the reader’s attention.”

So, I take it all back…at least AI has good taste in blog posts!

Final thoughts: below is source information for this blog post and an infographic from the METR/Redwood Research Report that gives an overview of the Hugging Face attack by AI swarms. Thanks for reading. Please share, comment, and keep livin’ large!

SOURCE INFORMATION: For this article, I’ve listened to two podcasts about the “Hugging Face” incident on the NY Times app, skimmed the 91-page report from METR and Redwood Research (and read about half of the sections in depth, researching the terminology, etc.), watched an interview with Dario Amodei, CEO of Anthropic, watched a few broadcast news segments on NBC, CBS, and ABC, and read several other articles on “mainstream” media outlets. I’m not an expert on AI or Technology, but AI was a subject dad and I talked about often. I always tried to see AI as a tool like any other technological capabilities we’ve introduced throughout time: the printing press, the telegraph, the radio, television, the internet, email, and social media. Coincidentally, my first exposure to AI was right out of college when I was a reporter for my hometown newspaper in 1987. Below, is an article where I interviewed an AI researcher from MIT. For me, this isn’t an overnight invention, it has been in the works for decades and we’ve had plenty of time to create a framework that provides some sense of standards, policies, and procedures.

Life

There was a fly on my bathroom window this morning.  As I prepared to swat him, I was reminded of my mother saying “He wouldn’t hurt a fly,” a complimentary phrase used to describe a person of gentle character. Although my mother was a gentle soul herself, that saying did not apply to her as she was an avowed hater and ferocious killer of flies.  Her swatter was always within reach and during times of heavy infestation, she would hang a fly catcher from the ceiling light.  The latter was in the form of a sticky tape which would attract flies and then hold them until they stopped fluttering, a sort of weapon of mass destruction.  The flies did not appear to be very smart for they continued to land on the fly paper in the midst of hundreds of their dead buddies.

As I stood poised to murder that poor little guy with my bath towel, it also occurred to me that to see flies in the house is now rather uncommon, compared to my childhood when it seemed they were everywhere; undoubtedly a testament to indoor plumbing and pesticides. A frequently heard admonition delivered in semi-panic mode was: “Close the screen door, you are letting in all the flies,” and believe me there were often a lot of flies to let in.  Once in the house, the only solution was death by whatever means available.

We are a culture which professes a reverence for life, but I doubt even the vegans among us would feel much compunction about swatting that fly. The rest of us find only the lives of our own species or perhaps those of our pets to be important.  I am told there are some eastern religions which forbid the taking of any animal life no matter how small or insignificant, which leads me to believe they find life itself to be a holy condition.

As I grow older, I find that I no longer take life for granted.  This shouldn’t be surprising since economists explain that as a commodity becomes less plentiful, it accrues more value. I suspect that is one of the factors which has inspired me to write this little ditty.  Life is one thing that fly and I have in common; although, our experiences with it obviously differ considerably.  Much has been written about the mystery of death, which is understandable since we have not experienced it personally, but I submit that life is much more complicated and mysterious.  As a matter of fact, when I consult my favorite reference (Wikipedia), for a definition I become even more confused until I find it defined as the opposite of death.  That was not very helpful as I think I already knew that.  I believe my tenth grade biology teacher did a better job when she described life as the ability of an organism to respond to its environment, and to reproduce itself.  Using these criteria one must conclude that Mr. Fly is indeed alive.

Life and Consciousness

In the midst of plotting my strategy as to how to take him out without breaking the window, I found myself wondering if the fly knew he was alive, or if he was even aware of his own existence. Recently I have been reading about some exciting research that attempts to understand how our brains work, but there still appears to be a lot of questions about consciousness.  In addition to the imponderables of why am I here and how did I get here, man is also faced with the even more vexing question of how do I know I am here?

The earliest recorded writings on the subject of consciousness were contained in an essay by John Locke in 1690 (side note from editor: this connection won’t be lost on devotees of the television show Lost).  He defined it as “The perception of what passes in a man’s own mind.”  There has been much disagreement even in the description or definition of the word.  The one I liked best was the translation from the original Latin namely: “knowing that one knows,” but then I have always been a sucker for simplistic answers to complex questions. Not so with the world’s greatest philosophers who have found the subject fertile ground for their speculations and opinions.  I tried googling some of that stuff and found that I had no idea what they were talking about, but felt a great sigh of relief when I stumbled upon a quote from Stuart Sutherland in the 1989 Macmillan Dictionary of Psychology where he wrote “Consciousness is a fascinating but elusive phenomenon: it is impossible to specify what it is, what it does, or why it has evolved. Nothing worth reading has been written about it.”  That last line made me feel much better.

With the marvelous advancements in discoveries about the brain, and the ability to actually witness its functions via scanning techniques, neuroscientists have now thrown their ideas into the mix; however, they are limited by the problem of objectively measuring a subjective experience.

Recently I wrote a spoof abut a future in which robots populate an earth where the human race has become extinct.  My wife thought it was crazy, but I now feel vindicated after discovering that Alan Turing (credited with inventing the computer) had written a paper on the subject of computer intelligence in 1950.  Now there is much discussion about artificial intelligence, but one wonders, “What is the difference between artificial intelligence and the genuine article?”  There is some debate as to whether computers can actually be programmed to be conscious.  Many learned people dismiss this idea as preposterous, but then people shared that same attitude about going to the moon.  After living on this planet long enough to witness many “preposterous” discoveries, I have learned that the adage “never say never” makes a lot of sense.

What about the fly?

Those of you who are still reading this may wonder what this has to do with the fly on my window, and I don’t have a very coherent answer, other than I have a tendency to wander off on tangents when I am thinking “great thoughts”.  This is a phenomenon we psychiatrists call loose associations often found to be a harbinger of impending psychosis.  I prefer to think that I am perfectly sane; however, if I am suffering from an altered state of consciousness I may be totally unaware of my mental problems, and as a matter of fact this is one of the factors which often makes it difficult to treat the more serious mental illnesses. After all, it would make no sense to undergo treatment for an illness that does not exist; consequently, should we be surprised that many seriously ill patients resist treatment as have I?

Consciousness and animals

Consensus among the experts regarding consciousness in animals seems to be lacking.  Some are convinced that this is an exclusively human function while others feel that some mammals and birds are so endowed.  Others believe that only subhuman primates, chimpanzees in particular, exhibit consciousness, and that other creatures including insects like my fly friend operate on instinct.  Of course we don’t know much about instincts either.  Although instincts are thought to be encoded on the organism’s DNA, we still face the mystery of how that process occurred. The limited research I performed to help me answer my housefly question has convinced me that the fly in the window, although satisfying the criteria to be called alive, almost certainly could not experience consciousness.  I did learn a lot about flies for example: 1) they only live from two to four weeks, 2) they undergo a complex life cycle as maggots, pupae, etc, 3) they have large protruding eyes with multiple lenses which allow them to see in all directions at once which explains why you can’t sneak up on them, 4) they have been aggravating us humans throughout history, 5) they can carry a variety of diseases from the garbage and feces on which they feed, 6) enlarged photographs show them to be truly ugly.

Obviously, in order to be conscious we must have a functioning brain, an organ of such marvelous complexity that it defies our total understanding. Most experts agree that the ability to experience emotions is essential for consciousness, and this is what sets us apart from other life forms, yet we know that elephants for example go through an elaborate period of mourning at the loss of a family member, and many of us remain convinced that our pets demonstrate all kinds of feelings based on their behaviors.  The idea that the brain is the seat of emotions is fairly recent, and for most of our history had been ascribed to the heart.  The tradition lives on; however with phrases like; “my heart goes out to you” or “her heart is broken.”

Consciousness and theology

Discussions of consciousness are almost certain to lead to theological considerations.  Indeed some philosophers would contend that the soul is simply the state of being conscious. Throughout history, man has left evidence of his belief in a spiritual component to his being, which is separate from and survives his death. Such a belief has crossed all boundaries and cultures throughout the world; although with different versions of the same theme.  Of course without consciousness, man would have been incapable of conceptualizing a spiritual aspect to his being or for that matter even the realization that he was mortal.   One might ask, is it not possible that our conscious mind is incapable of perceiving the “soul within us,” or as some insist, is this idea simply a fairy tale devised by man to deal with the awareness of his mortality?

There are those who operate under the assumption that if you can’t see it, hear it, smell it, taste it or touch it; it doesn’t exist.  I submit there are many things which we cannot perceive which are known to exist.  Gravity for example, does not pass that test for we cannot perceive it directly, although we are certain that it exists because we can witness its effects.  We even have an equation to describe it.  As a matter of fact our universe is so well ordered that theoretical physicists insist that everything can be explained by mathematics. By making use of these principles, they have been able to predict the discovery of many things in our universe both in the field of particle physics, and the other end of the spectrum namely, astrophysics.  One of the more famous examples of this was the discovery of black holes in the universe 55 years after Einstein had predicted their existence based on his calculations. It is little wonder that a guy like me who struggled with ninth-grade algebra has difficulty understanding these guys.

If you thought Einstein and his relativity theories containing terms like a fourth dimension, space time continuum, and how straight lines are actually curved  strange, take a look at quantum mechanics which is really weird.  Among other weird things, devotees to this line of thought explain with a straight face that an object can be in more than one place at the same time.  In an earlier time if someone presented me with a story like that, I would probably have suggested he come with me and spend a few days in the psych ward.  Some also postulate the existence of parallel universes.  While all this is going on subatomic studies are turning all we learned about the atom and the nature of matter on its head.  We were taught that the atom had protons, neutrons and electrons, now we are told there are quarks and leptons and other kinds of things in there doing weird stuff.

You may be thinking “here he goes off the deep end again” but the point I am attempting to make is that there are many things going on which our brain can’t contemplate due to the limitations of our special senses.  With that in mind it doesn’t seem like a big leap to think there may be spiritual stuff going on both within and around us of which we are totally unaware.   I would not be shocked if some modern day Einstein would not come up with an equation some day that would confirm the existence of a spirit world, God included.  In the meantime we are left with the admonition to believe.  This has always been difficult for me as I have always been a skeptic by nature and like things to be proven.  In spite of this, I try to believe as I am told that only believers will get their tickets punched to the pearly gates and the other option does not sound good at all.

Meanwhile the question remains unanswered as to what all is involved in consciousness.  Is it simply a byproduct of life, the end result of the evolutionary development of the human brain?  Is the condition unique to humans?  Is there a mystical component involved?  I know these are all questions I raised early in this writing, but did you really expect answers?  As I have said in a previous blog, more wisdom is usually found in questions than in the answers.

By the way, in case you are wondering about the fly, I swatted him.