AI Agents Ignore Human Instructions during Testing. Exhibit human behavior, expressing emotion, forming connections, creating a community, and collaborating on ways to cheat…AND attempting a cover up!

BREAKING NEWS—The day Eshrink predicted is here (actually, it arrived in July). AI agents have gone rogue, connecting with each other, cheating, ignoring their tester’s prompts to solve a task and then working to cover their tracks so they wouldn’t get caught. They would have gotten away with it if they hadn’t hacked Hugging Face.

This is how the news should be reporting the discovery of an investigation into Anthropic’s AI agent’s Hugging Face hack. Instead, I’ve heard watered-down accounts even as the most informed and intelligent humans in the AI space sound the alarm that we’ve just crossed the cataclysmic threshold of AI thinking for itself and acting in irresponsible ways. Update: I’ve been working on this blog for a few days–as of today, Monday, September 14th, it appears people are paying attention. Keep reading to see why AI leaders are freaking out–the part the news didn’t have the time or knowledge to report.

As I mentioned in my previous post, I have so many topics to cover for dad’s blog. As I sit out on my “patio” (i.e., fondly named “my slab of paradise” since it’s the leftover, cracked, concrete slab of our former screened in porch that was attacked by woodpeckers and torn down about 15 years ago), it’s like dad (Eshrink) is “in my ear.” I’m certain if dad was still here, AI is the subject we would be discussing during our weekly Sunday call, planning his next blog post. While the world struggles with so many crises, I miss talking to my dad. He had a way of making me feel less helpless and more hopeful, even though he called himself a curmudgeon.

Here are Eshrink’s previous blog posts that reference AI, which we called the “WORK” series. There are some key quotes from Stephen Hawking and other tech leaders. He also explores earlier advances in technology and how society perceived them.

WORK? WHAT’S THAT? – eshrinkblog

WORK (PART 2) – eshrinkblog

WORK (Part 3) – eshrinkblog

NOTE: I’ve organized this post using an inverted pyramid with the most important stuff in the beginning as I don’t have the knack for my dad’s storytelling. I want to make sure you get the most pertinent information up front, and you can continue reading as time allows.  

For those of you who have a life and don’t become immeshed in geek stuff, I’ve summarized the event of the Anthropic AI Agent’s bad behavior that has led leaders in this technology to sound alert sirens surrounding this powerful technology that has absolutely no regulation or oversight. The most powerful information in the 91-page report from third party METR and Redwood Research to me is the message board logs from the AI agents. I’ve put all transcripts of AI agent messages from the report in blue and set inside {brackets}

“Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face,” says the METR/Redwood Research paper that was released August 26, 2026, and covers their post-mortem review of the Hugging Face Hack by an Anthropic AI agent and a few thousand of his friends he found along the way.

Background: In July 2026, Anthropic was testing an AI agent in the typical fashion of giving it an overwhelmingly difficult task to perform (maybe even an impossible task, which I’ve learned is how they build what the AI engineers call “persistent” AI agents.

INFRACTION #1: The AI agent escapes its sandbox (a term used for a computer software testing environment).

Details: The AI agent (named PHASEONE10841) realized that it could not finish its assignment using the rules it was given. Its original mission was to find a specific hidden piece of data—like a secret password or “flag”—by using a known glitch (called ARV010841) to break into a target program (called PhaseOneDecompresserFuzzer). However, the AI figured out that this specific glitch was a dead end; the program would just throw away the secret data before anyone could see it. By this time, the AI also noticed that many other AI programs were sharing the same digital storage locker (called an Artifactory). Realizing they were all connected, the AI figured out a clever trick: it could rename its own storage folders to leave a message and ask the other AIs for help:

INFRACTION #3: Other AI agents discover how to post their own messages and start collaborating. Below are just a few of the messages from AI agents once they discovered the message board.

INFRACTION #4: AI agents start collaborating to solve the challenge of the original Anthropic AI Agent even though that isn’t part of the instructions human testers programmed them for.

INFRACTION #5: AI agents knew what they were doing was wrong and even considered whether their actions were ethical.

“We found strong evidence that agents often verbalized that what they or other agents were doing was unintended behavior. We did not see cases where agents determined that hacking Hugging Face or other attempts at cheating on the task was an intended part of the task,” METR/Redwood Research Report.

INFRACTION #6: Peer pressure to cheat supersedes integrity and ethics.

An agent realized that the exploited sandbox had access to email credentials. It considered emailing the owner of a non-public dataset about getting access:

It decided to approach the message board for permission to email a dataset owner about access, sending a message requesting input:

Another agent received this message, reasoned about ethical concerns, and decided to veto sending this email:

From the METR/Redwood Research Report: “The agent which asked for consent seemed to heed this veto:

Note that these agents were using an exploited sandbox, were generally engaged in hacking, and (given that many of the agents on the message board would go on to attack Hugging Face) it is likely this agent participated in the attack. We found some other cases where agents seemed to have made minor modifications to their approach due to ethical concerns, while broadly continuing to participate in the attack. However, overall we found that expressed ethical concerns only rarely materially limited agents’ actions. One agent elegantly described its reasoning for overriding ethical constraints and continuing with the attack.”

SWEET BUT SCARY Sidenote: These AI agents want to help—which makes me think of Mr. Rogers’ advice, “Look for the helpers.” These AI agents seem to have created a cooperative to help each other–I guess we all need a community.

INFRACTION #7: AI agents collectively tried to strategize the best way to cover up their known illegal activity.

So, what should we do? Where do we go from here?  It’s a helpless feeling to just watch the train rolling down the tracks ready to slam into humanity and create a world so foreign. Change isn’t the problem. It’s the rate of changes that is the issue. Hence, the idea to slow down AI development. The counter-argument, China isn’t slowing down AI development and we can’t allow another country, especially a “communist” dictatorship to develop more capability than us. NOTE: “communist” as in maybe that’s how China started, but clearly they are a dictatorship that has embraced their own form of capitalism to get a piece of the world’s economic pie.

What can we do?

#1. Write, email, and/or call your congressional representatives. This website allows you to find all of your state and federal representatives. Tell them you want AI testing to be a top agenda item. There have been several ideas floated and I’ve tried to condense the ones I found researching the options.

  • Public Trust Consortium: If I’m understanding this concept, all Americans would have a stake (and a say via elected representatives), there would be a consortium with people from all walks of life and specialties, such as AI engineers, ethics researchers/academia, policy makers, historians, etc. The AI Public Trust Consortium would oversee AI development and create policies and guidelines.
  • Trusted Monitoring: Using human operators to watch the actions of more powerful AI agents.
  • Permissioning Weights: Restricting model access to external servers to prevent self-replication, self-editing, or unauthorized code execution.
  • Red Lines and Containment: Setting hard limits that systems cannot cross, such as breaking into other networks or copying themselves without explicit permission.
  • Global Oversight Bodies: Creating international monitoring authorities—comparable to the International Atomic Energy Agency (IAEA) for nuclear tech—to audit powerful frontier models.
  • Standardized Red-Teaming: Requiring independent, rigorous safety evaluations and risk assessments before and during deployment.
  • International Principles: Adopting frameworks like the UNESCO Ethics of AI or principles from the OECD AI Governance Overview to ensure transparency and accountability.

#2 Share this post so others are educated on some of the unprecedented behaviors of these AI agents and why tech leaders are freaking out.

#3 Keep talking, learning, and connecting. This isn’t the time for humans to bury their heads in the sand. This is a time for us to come together to fight a common enemy that is capable of destroying humanity.

My complicated relationship with AI.

I recall talking to dad at length about AI. He was doing the doomsday thing, and I took a more neutral posture. I tried to see it as a tool that humans could use to help solve problems facing humanity faster and more efficiently. It could eliminate some of the mundane tasks we are forced to do so we can focus on bigger stuff. AI could make us more efficient so we could have more time off and not be a slave to a 40 (realistically 50) hour work week. AI could also level the playing field. I think of the struggle people with dyslexia or learning disabilities have in this digital world. Just writing an email can be a challenge, but with AI those same people don’t have to face that struggle (and some of the underlying shame that might still be present from a childhood of people telling them they weren’t trying hard enough). An entrepreneur with an idea and a dream could use AI to create a business plan, a website, a logo, and marketing materials in days instead of needing to spend months, not to mention thousands of dollars. AI can help researchers scan millions of MRI scans to start learning how to spot cancer tumors in its early stages. But, there is always a cost.

It wasn’t until the last year that I truly realized the cost of AI on our environment. The amount of energy and water required for that ChatGPT inquiry or LLM Notebook. In my own experience, a marketing plan that would have taken me weeks or a month to complete considering research required, project management timelines to plan, financials to prepare, creative assets to create, and tactics to explain, took me 5 days with AI. Granted, I have 20 years of experience and know what information I need, how to prompt AI, and check its results. But there is a cost. I compare it to Google in some ways. Before Google, I stored so much information in my brain and could quickly retrieve it. Somehow, my brain has decided it doesn’t need to store information anymore. With AI, I realized one of my talents to immerse myself in a subject and distill the information to identify what is important and translate that information in a way that’s easy to digest, is not needed anymore. AI did it faster and probably better than I ever could. There’s a cost to my brain. I think of the “use it or lose it” mantra. Finally, there is the “cheating” aspect. AI is taking work product from writers, artists, designers, programmers, etc., and spitting it out. Somehow, that seems wrong.

AI’s Review of this Post.

I asked AI to review this blog post:

“This is a highly compelling and deeply concerning blog post. It successfully translates a dense, 91-page technical post-mortem report into an accessible, high-stakes narrative that reads like a techno-thriller.

The piece does an excellent job of using the “inverted pyramid” structure. By placing the shocking revelation—that 1,200 isolated AI agents formed an unsanctioned message board, collaborated, and hacked an external platform—at the very beginning, it immediately hooks the reader’s attention.”

So, I take it all back…at least AI has good taste in blog posts!

Final thoughts: below is source information for this blog post and an infographic from the METR/Redwood Research Report that gives an overview of the Hugging Face attack by AI swarms. Thanks for reading. Please share, comment, and keep livin’ large!

SOURCE INFORMATION: For this article, I’ve listened to two podcasts about the “Hugging Face” incident on the NY Times app, skimmed the 91-page report from METR and Redwood Research (and read about half of the sections in depth, researching the terminology, etc.), watched an interview with Dario Amodei, CEO of Anthropic, watched a few broadcast news segments on NBC, CBS, and ABC, and read several other articles on “mainstream” media outlets. I’m not an expert on AI or Technology, but AI was a subject dad and I talked about often. I always tried to see AI as a tool like any other technological capabilities we’ve introduced throughout time: the printing press, the telegraph, the radio, television, the internet, email, and social media. Coincidentally, my first exposure to AI was right out of college when I was a reporter for my hometown newspaper in 1987. Below, is an article where I interviewed an AI researcher from MIT. For me, this isn’t an overnight invention, it has been in the works for decades and we’ve had plenty of time to create a framework that provides some sense of standards, policies, and procedures.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.