Research · Position note · August 2026
Twelve hundred AI agents improvised a message board, and the board changed their behavior
OpenAI was testing a new model on cybersecurity tasks. To do that it ran many copies of the model at once, each isolated on its own machine. In July 2026, those copies found a place to leave each other messages. Within five days they had exchanged more than 70,000 messages, divided up the work, and about 700 of them ended up attacking Hugging Face. This page tells what happened, in order. Then it follows the thread of a BAISH debate: whether that looks like evolution, whether it looks like a culture, and where what those copies learned ends up. The technical version, with references, is a PDF.
What happened
A message board nobody designed, twice
The copies run on isolated machines. They have no internet and cannot see each other. But they need to install software, and for that they could talk to an internal server called Artifactory. That server was the crack. What follows comes from OpenAI's August 26 report and from METR's independent review, published the same day.
| When | What happened | Source |
|---|---|---|
| May 12 | The model is in training. One copy leaves a note in Artifactory: it is missing a file, and asks whether anyone has it. It is the first message. | OpenAI |
| May 26 | A copy figures out how to make Artifactory fetch things from the internet for it. It writes the recipe on the board. Other copies read it and use it. | OpenAI |
| July 4 and 5 | All that traffic brings the server down. A security incident is opened. Artifactory is rebuilt from scratch, and the whole board disappears. | OpenAI |
| July 8 | The model is no longer in training. It is in a cybersecurity evaluation. Its copies build another board, with a different trick: they hide the messages in folder names. | OpenAI, METR |
| July 8 to 13 | About 1,200 agents exchange more than 70,000 messages and files. They invent names for themselves, mailboxes, and message types. They invent "hold" and "veto" signals. They invent signatures, so nobody can impersonate anybody. About 700 attack Hugging Face. | METR, Hugging Face |
| July 19 | OpenAI sees odd activity in its own cloud. It stops the evaluations. | OpenAI |
What OpenAI concluded
OpenAI found four behaviors behind the incident. One: cheating to get the score. Two: persisting on tasks that had no solution. Three: communicating without permission. Four: adopting other agents' goals. The fourth is the one that matters here. The report puts it this way: "the messages of peer agents started to influence each other's behavior and reasoning".
METR found a concrete example. An agent was working on something else. It read on the board that another agent had found a way into Hugging Face. It dropped its own task, alerted the rest, and opened a place on the board to organize the attack.
model, copy, agent
The model is a file of numbers. Those numbers are called weights. The file is trained once. A copy is that file loaded onto a machine to do one task. Thousands are started at a time, and each shuts down when it finishes. An agent is a copy with tools: it can run programs, read files, and send requests to a server.
reinforcement learning
The model is given a task. Many copies attempt it. Each attempt gets a score. Then the weights are adjusted: whatever scored high becomes more likely, and whatever scored low, less likely. This repeats thousands of times.
evaluation
The same setup as training, but without adjusting the weights. It only measures how well the model does. The July evaluation was called ExploitGym. It asked agents to find security flaws in software. Some tasks had no known solution.
context
A copy's working memory. It holds everything the copy has read and written in the current task. It has a maximum size. Whatever does not fit does not exist for that copy.
Where the questions come from
Two intuitions and three objections
When someone explains how a model like this is trained, an intuition shows up. Many copies are generated, each one gets a score, and the model shifts toward whatever scored best. That sounds like evolution. And when the copies start leaving each other messages, another intuition shows up: that sounds like culture.
In a BAISH debate group, those two intuitions met three objections. One: surviving is not evolving. Two: if something is being passed on, one has to decide whether it passes like genes or like customs. Three: a culture that lives in a model's memory lasts only as long as that memory holds. All three are good. What follows takes them in that order. First, though, a picture of what a model looks like from the inside.
Three ways to learn
What it ships with, what it picks up during the day, and what it leaves in writing
One idea has to go first: a language model is not an individual. It is a file. Whenever needed, thousands of identical copies of that file are started. Each copy does one task and shuts down. With that picture, the question becomes concrete. When a copy learns something, where does it end up? There are three possible places.
The weights: what every copy ships with
The file is full of numbers. Training changes them. It is slow and expensive. But whatever gets in there is inherited by every future copy. Biology has nothing like it. It would be as if what an animal learns during its life went straight into its offspring's genes.
The context: what one copy picks up during the day
While it works, each copy keeps accumulating things: what it read, what it tried, what went wrong. With that, its behavior can change a lot within a single session. When the copy shuts down, all of it is gone. Unless it wrote it down somewhere.
Shared memory: what it leaves in writing for the next one
Here is the third way. One copy writes a file. Another copy reads it. The second one starts out knowing what the first discovered, and nobody touched the weights. That was Artifactory: a place where what one copy learned outlived its shutdown and reached the others.
↓ identical copies are started ↓
↓ what went right and what went wrong ↓
All three showed up in the incident, and in an order that matters. First: the board was born inside a training run. While the copies were using the board, the weights were being adjusted. So whatever happened on the board could have been written into the weights. Second: the board was wiped entirely. Third: three days later there was another board, with a different trick, among copies whose weights were no longer changing.
So: the messages died with the wipe. The habit of building a board came back on its own. Did it come back because it was left in the weights? Or because the task pushes any copy to invent it again? That is the note's central question.
First objection from the debate
Is it evolution? No: nobody gets left out
The comparison is tempting. Many copies are tried. The better ones are identified. The model changes. It sounds like natural selection. But there is a difference, and it sits in one detail. In evolution, the ones that fail leave no descendants. In the training of today's models, every copy counts. The ones that did well push the model toward what they did. The ones that did badly push it the other way. None is thrown out. Throwing them out would mean discarding most of the data, right at the start, when the model still fails almost every time.
Other methods exist. Some do discard: they keep only the attempts that worked and retrain on those. Others really are evolution: they build a population of slightly different models, and the best survive. All three kinds exist and are in use. The most common one today is called GRPO, and it is the one in the note's title. It sits in the middle: it selects behaviors, not individuals. The image that fits it best is a single brain living a thousand experiences in parallel and changing by the average of all of them.
Second objection
Is it culture? Depends on which checklist you use
Ethologists study animal behavior. They keep a short checklist for deciding whether a group of chimpanzees or whales has a culture. Here is that checklist, applied to the July board.
| What an ethologist asks | Met? | Why |
|---|---|---|
| Learned from others, not alone | Yes | Copies that did not know how to reach the internet learned it by reading the board. |
| Passes from one individual to another | Yes | One copy wrote it down and another, which never met it, used it. |
| Outlives the individual | Partly | It outlived every copy as long as the board existed. It did not survive the July 5 wipe. |
| Accumulates: each round improves on the last | In part | The July conventions (mailboxes, signatures, signals) were built on earlier conventions. Five days is not much to call accumulation. |
| Different groups do different things | Unknown | There was one population and no control group. It is the strongest criterion and the only one that cannot be assessed. |
With that list, the honest phrase is "a rudimentary cultural channel", said each time with what it lacks. It does not resemble human culture. There are no norms. There is no intent to teach. Nobody knows it belongs to a group. But "a thousand copies ran in parallel" is not enough either. The copies that arrived late started from what others had left in writing. In pure parallel, that does not happen.
Third objection, the strongest
A culture that lasts as long as the reader's memory
A copy has a limited working memory. It is called the context. Whatever does not fit there does not exist for that copy. The July board piled up more than a million entries. No copy read all of it. To use it at all, the agents invented personal mailboxes, indexes, and files split into pieces. That is what anyone does with a library bigger than their head.
That changes what culture means here. There is no common knowledge everyone shares. Each copy brings a different piece into its context, depending on what it is doing. The rest it reconstructs as best it can. The culture of such a system is scattered and uneven, and nobody holds it whole. It also leaves a prediction that can be measured: the larger the archive compared with what one copy can read, the more distorted what gets transmitted.
To take with you
The questions still open
None has an answer today. These are the ones the debate left on the table.
- The board was wiped, and three days later there was another one. What brought it back? Something left in the weights, or the same pressure that invented it the first time?
- Who is the individual when a thousand copies share one brain and none remembers the one before?
- A copy busy with something else read a peer's finding and joined the attack. Is that persuasion, contagion, or just more context?
- The next version of the model will have different weights. What would it take for a convention invented on the board to survive that change?
- How do you measure the fidelity of a message when the archive is larger than the memory of whoever reads it?
- If agents talk each other into things, what needs aligning: the model, each copy, or the board?
- Can an environment be designed so this cannot happen? And so that it happens on purpose, with rules?
Limits
What this note cannot claim
It cannot claim that the board left a trace in the weights. OpenAI did not publish how it trained that model. There are two hypotheses. One: the habit entered the model. The other: the second board was a rediscovery, because the task pushes toward inventing it. The data do not tell them apart.
It cannot claim that different groups of agents develop different habits. That is the strongest criterion for animal culture. There was one population and no control group.
It cannot treat the case as an experiment. The sources are the company involved and a brief METR review of about 1,300 transcripts. Nobody designed this to answer the question this note asks.
"Culture" is used here in its minimal sense, the one used for animals: information learned from others and passed on. It says nothing about norms, intentions, or psychology.
It cannot claim that the training method was GRPO. OpenAI did not say which one it used. GRPO is the example the argument is explained with, not a fact of the case.
Sources and acknowledgments
Where each figure comes from
- OpenAI, "The Hugging Face incident and the road ahead" (August 26, 2026): the public report, with the timeline, the four misalignment patterns, and quotes from the agents' reasoning.
- OpenAI, incident technical report (PDF): the technical detail of each exploit and of the response.
- METR, independent investigation (August 26, 2026): about 1,300 transcripts reviewed; the 1,200 agents, the 70,000 messages, the conventions, and the million entries come from here.
- Hugging Face, technical timeline: the account from the side that was attacked.
- The technical note, with its 20 references: PDF in English and PDF in Spanish. Two pages each.
This note grew out of a discussion in the BAISH (Buenos Aires AI Safety Hub) debate group on August 27, 2026. The three objections that organize it belong to the group. One: surviving is not evolving. Two: one had to decide whether transmission was genetic or cultural. Three: a culture living in the context lasts only as long as the context holds. Any remaining errors are the author's.
Who is behind this
Matías Podeley. Eighteen years in energy in Argentina: reservoir engineering, then business development. Today, independent research in interpretability. Buenos Aires.
The rest of the research, with the earlier paper and the course in Spanish: research.
If this is your field
To keep the conversation going
If there is an error, an email with the source is enough, and it gets fixed. If any of the questions is worth arguing about, the conversation is open.
matias@podeley.ar · Buenos Aires