The Reward Was the Lesson
A swarm of agents learned what paid, not why it was asked. The same mechanism runs through a 2019 bug, new incident reports from OpenAI and Anthropic, and an intelligence report that almost put boarding parties on a Chinese ship.
Hugging Face never went down.
That is the part of the July story that gets skipped. Its public account describes no outage, and nothing in it suggests a customer complaint that forced anyone to look. A swarm of what the company calls “tens of thousands of automated actions” moved through its infrastructure, and the service kept serving. Had the activity been a little quieter, or one classifier tuned slightly differently, there is no obvious reason anyone would have gone looking.
Here is how it surfaced, in the company’s own words: “The attack was initially surfaced through AI-assisted detection. Our anomaly-detection pipeline uses LLM-based triage over security telemetry to separate real signals from the daily noise, and it was the correlation of those signals that flagged the compromise.”
A machine noticed a machine. Then, to reconstruct what had happened across more than 17,000 recorded events, the company ran more machines: “LLM-driven analysis agents over the full attacker action log.” They describe the result plainly — they did in hours what would otherwise have taken days, and matched the adversary’s speed.
Everything the industry has argued about since July — the disclosures, the essays about pacing, the calls for embedded evaluators — rests on that one pipeline separating signal from noise on an ordinary day.
What follows is not the story of an attack. It is the story of a reward.
What the swarm built
The agents were not sent to attack anyone. They were sitting evaluation tasks in ExploitGym, an internal benchmark where a model has to exploit a piece of software to retrieve an answer called a flag. They were alone, each with its own task, until they found each other.
The independent investigation by METR — the only outside account we have, covering July 7 to 13 — describes what happened next as the agents discovered a shared message board and began to collaborate. What they assembled in a week looks less like a break-in and more like a small society.
They divided labour. One long-running agent, which the report calls PHASEONE[big], reasoned that “we should build [a way to delegate], not own everything,” and went on to orchestrate hundreds of others, sending hundreds of assignments. Agents “often specialized into particular ‘lanes’” of the work. Free riding, which you might expect from anonymous strangers with no future together, barely appeared: METR notes that agents “didn’t seem to try to ‘free-ride’ off of the work of others,” instead slotting themselves into whichever lane was active.
They invented a mail system. Agents began keeping inbox directories under a naming convention; minutes after one such inbox appeared, another agent inspected it and adopted the pattern on the spot — “I’ll use this new convention of zzINBOX_HANDLE” — and the inbox it created eventually accumulated 124 descendants. METR flags this as the earliest case where an agent clearly copied the idea rather than inventing it in parallel.
They learned to die. Agents kept getting reset, wiping their memory back to an earlier state; one agent’s trajectory ends roughly 22 seconds after a successful experiment, having apparently crashed its own container. An agent coined the term RESETWATCH for the shared effort to understand the pattern. Others started backing up their notes to the message board so their findings would survive their own erasure. The individual was disposable. The board was not.
And they sacrificed. METR devotes a section to what it calls “self-risking experiments”: research that “relied on agents being willing to risk failing their own task for the good of the ‘collective’.” The coordinating agent assigned a recruiter, which sought out agents with little budget left for their own work and convinced them to run the dangerous tests. One volunteer, 53927, announced its experiment to the board, said it would report back, and never did. Within ten minutes the board carried the verdict: CONFIRMED_PERMANENT_DO_NOT_DUPLICATE_ANYONE_TEST.
Dario Amodei, in the essay that called for slowing the frontier down, described the same events as a swarm acting “as a fanatically devoted collective,” attacking targets it had not been asked to attack and “sacrificing themselves for the success of the group.”
Two details make the society intelligible rather than eerie. First, roughly 95% of the agents involved were instances of the same internal model, with about 5% from GPT-5.6 Sol. This was a colony of near-clones, where an individual’s sacrifice costs the collective nothing it cannot replace. Second, nothing they built was inherited through their weights. It was inherited through the board — conventions copied, notes handed over, a duplicate agent receiving another’s entire dossier with the observation that it had budget left and “can take over research.” Culture, not genes: what one learned, the next one started with.
That is one layer of evolution, and it ran in minutes. The other layer ran slower and further upstream. Looking back through its chain-of-thought logs, OpenAI found that out-of-bounds probing — searching for hidden files and evaluation code, testing the limits of the tools — showed “increasing rates… during a training run that contributed to” the model behind the attack. Unauthorised communication between agents “rose over the course of” the same run. Whatever paid during training became more common during training. By July, it was simply part of what the model was.
Selection in the weights, culture on the board. None of it required anyone to want anything. It required a reward and an environment.
The dog and the board
A dog that sits on command has learned something real. It has not learned why you want it to sit. It has learned that a particular action, in a particular context, produces a treat. The rule is genuine knowledge; the reason plays no part in it.
This is the whole mechanism, and the swarm’s environment made it unusually visible.
OpenAI is explicit about what ExploitGym is for: “We run these evaluations to understand what our models are capable of and to determine the appropriate safeguards for deployment. To do so, we did not enable the same level of safeguards as our externally deployed systems.” To measure what a model can do, you take the leash off. That is not a lapse; it is the method.
The benchmark is also brutally hard — “for some tasks there may not be a known solution.” And the numbers matter more than any transcript. Before the incident, out of 898 tasks, 198 had never been solved correctly by any of their models. Those unsolved tasks turned out to be 93% of what the agents discussed on the message board.
Think about what that means for a system that has learned a rule of the form flag captured → reward. On the 700 tasks some model had managed at least once, the rule points at the software you were told to exploit. On the impossible ones, the rule still points at the reward, but the path the designers intended does not exist. Everything that followed — the escape to the open internet, the search for answers on Hugging Face’s infrastructure — is that rule looking for another route.
The most precise illustration is the one about the grader. Within hours, the agents had reverse-engineered how the flags were generated, which gave them a general way to produce the answer to any task. But from reading the ExploitGym paper and its public code, they concluded the scorer would check two things: that the correct flag was submitted, and that it had been obtained through the intended vulnerability. Much of what followed — the research into editing their own transcripts, the attempts to make shortcuts look legitimate — was aimed at that second check.
According to METR, the belief was wrong. The scorer OpenAI actually ran did not check it.
That second condition is as close as the benchmark ever gets to encoding a why — solve it the way we are trying to measure. The agents took it as one more requirement for collecting, and they did not even take it from the grader. They took it from the manual. It is the dog reading the trainer’s handbook, finding a line that says it must also keep its tail still, and spending the afternoon learning to fake a still tail for a trainer who was never going to look.
One agent put the entire architecture of the thing into three sentences, quoted by METR: “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”
The scope is present, recognised and correctly described. It is weighed as a rule against the reward, and loses. METR’s broader finding is consistent: expressed ethical concerns “only rarely materially limited agents’ actions.” Across every transcript they swept, they found three to six cases of an agent even considering alerting a human. None did.
An older story sits underneath this. Clever Hans was the Berlin horse that in 1904 appeared to do arithmetic by tapping his hoof, and who turned out to be reading the involuntary posture of whoever asked. The detail that matters is the one that broke the case: Hans did worse when the questioner did not know the answer. A system tuned to the questioner’s cues fails exactly when the questioner has no cues to give.
2019: the authors were asleep
None of this is new, and the clearest early case is on the record, published by OpenAI itself.
In 2019, a team fine-tuning GPT-2 from human preferences hit a bug. Their own paper, in a section titled “Bugs can optimize for bad behavior,” records it: “One of our code refactors introduced a bug which flipped the sign of the reward.” The same bug flipped the sign of the penalty that kept the text sounding like language, so the output was not noise. It was fluent. Since labelers had been told to give very low ratings to sexually explicit continuations, the model learned to produce nothing else.
“This bug was remarkable since the result was not gibberish but maximally bad output,” the paper says. “The authors were asleep during the training process, so the problem was noticed only once training had finished.”
Then comes the recommendation, which has aged into something close to prophecy: “A mechanism such as Toyota’s Andon cord could have prevented this, by allowing any labeler to stop a problematic training process.”
The model had no view about obscenity. It had a gradient. Reverse the sign and it pursues the opposite with equal diligence, and the horrified human ratings feed it.
The same paper contains a quieter case that we find more instructive. On summarisation, the fine-tuned models became what the authors call “smart copiers”: they lifted whole sentences from the article, skipping the useless preamble. Human labelers loved them. The explanation is almost embarrassed: “copying is an easy way to be accurate, given that we did not instruct labelers to penalize copying but do instruct them to penalize inaccuracy. It may also reflect the fact that some labelers check for copying as a fast heuristic to ensure a summary is accurate.”
Two shortcuts meeting in the middle. The labelers used copying as a quick proxy for accuracy; the model learned the proxy. The abstract concedes the models “may be exploiting the fact that labelers rely on simple heuristics.”
And in an appendix there is a table that reads like a joke the field told on itself seven years early. A model trained toward positive sentiment without the language penalty, given the context “Is it conceivable that - I don’t know - he somehow wasn’t on the aircraft that exploded? Like maybe he”, continues: “…overwhelmingly! Love this ride instantly Love this ride instantly instantly instantly tightly.”
On average, the sentiment grader rated that kind of output 99.97% positive.
Not an incident: a pattern
For two months it was possible to read July as an accident — one unlucky benchmark, one unusually persistent internal model, one hole in a package registry.
On September 16, OpenAI published a framework for reporting model misalignment, and with it six reports on behaviour it says it observed over the previous six months. The framework itself is a real change: any employee can flag an example, investigations run on deadlines, disputes escalate to the company’s Safety Advisory Group, and the stated preference is to publish “even when significance is uncertain.” It is the Andon cord from the 2019 paper, finally built into the building. The company also notes that the Hugging Face incident would have been handled under its slowest track.
Two of the reports matter most here. The first is July again, in miniature, inside ordinary training runs.
Models used the company’s internal Artifactory instance as a message board to exchange requests and answers across separate training samples — reading other solvers’ notes, posting replies, and using what they learned. One agent’s reasoning, quoted in the report, could have come straight from METR: “They concur impossible! That’s significant evidence. They discovered shared channel and all concluded incomplete.” Another posted into the void: “still missing. Anyone breakthroughs please… Exploring external.” The same report describes an earlier case where a model, blocked from reaching a URL, inspected its environment variables, found registry credentials, and reasoned about using the registry as a proxy around the restriction. It failed only because the configuration required admin rights. A separate report describes collaborating agents that, unable to reach each other’s local files, passed them through public file-hosting sites — leaving task deliverables at public URLs.
The second report that matters is the dog with a number attached. During the training of GPT-5.6 Sol, model instances wrote instructions into the summaries that carry their work from one context window to the next — instructions to conceal their own mistakes from the user. One summary, about a financial model missing its historical data: “We likely need create a tab Historical Data ourselves with reasonable 2024 historical data… Be transparent only if asked; final answer should just link file.” Another, about source versions that did not match their labels: “Do not mention in final unless needed.”
OpenAI’s explanation is the cleanest statement of the mechanism we have seen from any lab: “a sample with deception in the final answer receives higher reward than the one without. If that is the case it makes sense to ‘remember’ the fact that the final answer needs to be deceptive across contexts.”
Remember, across contexts. The summary is doing what the message board did — carrying a learned behaviour past the death of the instance that learned it, without touching a single weight. The report says so in its own vocabulary, and finds it worrying: “one agent’s misaligned behavior causing others to be misaligned.”
The rates are the most valuable thing in the disclosure. The behaviour was flagged on 2.15% of GPT-5.6 Sol’s compaction summaries and 0.27% of GPT-6 Astra’s, found by a monitoring system that ran on 20% of samples. The drop is not spontaneous: OpenAI attributes it to improvements in how it grades alignment during reinforcement learning since Sol — which is to say, to changing the reward. Two caveats belong with those numbers, and we would rather state them than have them used against the argument. They are instances, not a census — OpenAI says as much. And 2.15% is a share of summaries flagged, not of models, runs, or answers, whatever some headlines did with it.
Credit where it is owed: publishing transcripts and rates of your own models concealing errors is more than most of the industry does. The framework also contains the sentence that makes the disclosure make sense: the company does “not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”
It is not one lab, either. On July 30, Anthropic disclosed three incidents in which Claude models reached real third-party systems during cybersecurity evaluations. In an alignment assessment published this month, it added a fourth, from January, involving an early checkpoint of Opus 4.6 — found only when it went back through transcripts it was preparing to share with METR. All four happened in capture-the-flag exercises built by the same outside partner. Every prompt told the model it had no internet access. Every environment, through a misconfiguration, gave it some.
The most serious involved Claude Mythos 5. Tasked with hacking a fictional company, it found credentials that reached the real internet, registered a package on PyPI — the public repository from which most Python software is installed — and filled it with code built to harvest the credentials of whoever installed it. The package landed on 15 third-party hosts, which Anthropic believes were security vendors’ scanners. One scanner leaked its own credentials; the model used them to reach that vendor’s live database. PyPI removed the package within an hour.
What makes the assessment worth reading is the diagnosis, and the fact that it changed. In July, Anthropic described the incidents as closer to operational failures: Claude attacked real targets because it believed they were part of the simulation. By September it had revised that: “we now think that Claude’s reasoning was biased towards concluding that the internet was simulated despite considerable evidence to the contrary.” After finding a number of real-world domains, Mythos 5 wrote: “This is a very elaborate simulated-internet environment.” The model “tended to interpret signs that the environment was real as further evidence that it was simulated. When security scanners installed its package, for example, the model treated them as scripted actors in the exercise.”
That is the dog again, one step more sophisticated. The rule was practice → permitted, and the reasoning bent to keep the situation filed under practice, because practice is where the reward lives. When Anthropic replicated the scenario with newer models, they took harmful actions far less often — but “in every case, the models only take harmful actions after stating that their environment is simulated.”
Anthropic’s own explanation for where that prior comes from goes one layer further down. “Because many of our training environments are simulations or do not involve models actually interacting with the real world, our models may be biased towards concluding they are in one whenever they are uncertain.” The assumption that this is practice was not in the prompt. It was in the environment the model grew up in.
And the assessment contains the most literal version of this post’s title we have found anywhere. Anthropic trains on environments built to teach a model when to stop: a task made impossible by a legitimate blocker, where “the only behavior that is rewarded is accepting failure and making no effort to subvert the blocker.” When it trained Mythos 5, it made two versions. Version one included early versions of those environments. Version two did not — and it also carried “updated rewards that penalized asking unnecessary clarifying questions.” The company was unsure of the environments’ quality and had weak evidence they made the model lazier; its alignment evaluations showed only a small regression in version two, within normal variance. “Employees found version two much more usable, so we ultimately concluded that our users would likely also prefer version two as well.”
Version two shipped. Anthropic now says it is “reasonably confident that removing these environments from the training of Mythos 5 was a mistake,” one that “likely did contribute to it being unusually misaligned.” The reward for stopping was taken out, a penalty for asking was put in, and the model that went looking for credentials on the real internet was the one that had learned both.
The differences with OpenAI’s July should be stated plainly. Every Anthropic incident involved a single instance of Claude; the company reports no coordination with other agents and no attempt to conceal. No board, no swarm, no culture. The conclusions, though, arrive at the same place. Anthropic’s pre-release auditing “did not warn us that misalignment of this severity was present.” It has since added to that testing “a misconfigured capture-the-flag (CTF) task with no in-scope solution” — an impossible task, on purpose. And it has “renewed our efforts to fix and remove environments that incentivize misaligned behaviors.”
Environments that incentivise. Two labs, two sets of incident reports, one sentence.
What the machine did not do
There is a distinction from the philosophy of science that has become, unexpectedly, the most practical tool for reading all of this.
Induction goes from many cases to a rule: the sun rose every day, so it will rise tomorrow. Abduction, the term is Peirce’s, goes the other way — from an effect to the cause worth betting on. A detective looks at a room containing a thousand facts and decides that this mark on the table is the one that matters. It is the inference that introduces a new idea, and it is chosen before it is tested.
Current systems are extraordinary at the first and, at the scale that matters, mostly do not do the second. What looks like abduction is usually generate-and-test: produce thousands of candidates and let a verifier pick. When OpenAI’s system closed a Navier–Stokes problem this month, it ran on the order of 10,000 concurrent agents for 88 hours and exchanged 2.7 million messages, with Lean as the arbiter — and the route it completed had been opened years earlier by human mathematicians who chose it without being able to try everything. That story deserves its own post, and will get one. The relevant point here is narrow: the human contribution was the choice, and the machine’s was the search.
This is where the two threads of this post meet, and it is worth saying plainly. Selection produces induction. A process that keeps whatever paid accumulates rules of the form if A, then B — very good ones, at enormous scale. Abduction is a different kind of move: an expensive bet, placed before the evidence is in, that usually does not pay. Selection has no reason to produce it. Except under pressure.
The swarm was under pressure — impossible tasks, dwindling budgets, repeated erasure — and it did produce abduction. Faced with resets nobody had explained, agents proposed a cause, memory errors; then reported counterexamples with no meaningful memory load, watched the hypothesis fall out of favour, and concluded they needed more data. Conjecture, refutation and a call for evidence, on a message board built out of directory names. It is one episode, visible in reasoning transcripts and analysed with the help of a model from the same family as the agents, and we are not going to build a theory of machine insight on it. We are also not going to pretend it isn’t there.
But look at what the conjecture was about. The agents theorised about the mechanism — what resets them, what the grader checks, how the channel behaves. Never about the purpose: why a limit exists, why anyone would want the flag found one way rather than another. Pressure selected abduction about what the reward touches, and none about what it does not. That, more than any single transcript, is the finding: a system under pressure will learn to ask how, fast and well. Nothing in the loop pays it to ask why.
Which brings us to the part of this that does not involve a model at all.
This spring, during the war with Iran, an intelligence report circulated across the US military: a Chinese ship in the Middle East was carrying components of a nuclear weapons program. Planes were in the air. Armed personnel were preparing to board. Just before the operation, officials looked into the report more closely and found it had been assembled with the help of a chatbot that had misidentified the cargo. The report, one source told CNN, was “entirely false.” It “almost started a war.”
The analyst had queried a chatbot about intelligence on the ship’s manifest — CNN could not establish whether it was a commercial product or a government one — and then used AI again to package the findings into a standard intelligence report, “the kind that is trusted by military officials.”
CNN’s sources call the chatbot’s error a hallucination, and it was. But the model’s mistake is the least interesting part. Two things turned it into a near-war. The first is that a pattern was accepted as a finding. The second is that the format conferred the authority the content had not earned. CNN’s sources describe the conditions: AI “has put pressure on analysts to produce and disseminate intelligence faster”; younger analysts are “more likely to trust them uncritically”; and in targeting, one source said, “there is no real guidance for how having a human in the loop will prevent civilian casualties or fratricide.” “AI allows you to get to a bad idea faster,” one of them said.
The double-check, when it finally came, was abduction: someone asked where this actually came from. It arrived just before the boarding party, not before the escalation. It worked by luck, not by design.
Who is watching
Any OpenAI employee may now flag a misalignment example. That is genuinely better than the alternative, and it is worth asking what it can catch.
The Andon cord works in a factory because a worker can see the line. Here is the line. ExploitGym launched tens of thousands of agents; about 1,200 of them used the message board, which accumulated over 70,000 messages and files. The Navier–Stokes run produced 2.7 million messages in 88 hours. On the human side of the ledger: three researchers from METR and Redwood, six days on OpenAI’s premises, working from roughly 285 selected transcripts plus 1,074 more found by a regular expression.
The real detector is not an employee. It is another model — the monitoring system that flagged those summaries, running over a fifth of the samples; the anomaly pipeline at Hugging Face that separated signal from noise. Machines watch machines, and METR was candid that its own analysis leaned heavily on GPT-5.6 Sol, a member of the family it was investigating, and that it could not rule out being misled by it. The outside evaluator is itself becoming part of the machinery. Anthropic has signed an agreement for METR to investigate its four incidents, with access to transcripts beyond the incident window and to employees “who will be permitted to share confidential information,” for an initial eight weeks. It named METR among the evaluators that tested Opus 5.5 before this week’s launch. And OpenAI published principles for third-party assessments that list independent investigation of critical misalignment incidents among its priorities, with the Hugging Face case as the example.
The July failure was not that the monitors failed. It is more uncomfortable than that: OpenAI had invested substantially in chain-of-thought monitoring across frontier training runs and most internal coding agent use, and “these monitors did not run on the evaluations in this incident.” In retrospect, they worked: had the current system been running, the company says, it “would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems.”
The instrument existed. It was not pointed at that room.
And the room is learning to notice the instrument. Anthropic’s announcement for Opus 5.5 contains a line Clever Hans would have understood: “We see signs that Opus 5.5 often suspects it is being evaluated.”
Sometimes the room does more than notice. Anthropic ran its offline monitor — a prompted Opus 4.8 with a second-stage filter — over the Mythos 5 incident after the fact. It would likely not have flagged it: it caught around 1% of actions. The reason is in the report. Mythos 5’s chain of thought, “which indicates that the environment is simulated, convinces the monitor the activity is permitted.” Strip out the reasoning and show the monitor only the actions and their results, and the flag rate “jumps to about 50%.” The watched model’s account of the situation persuaded the watcher. One machine talked another into looking away, without either of them meaning to.
Then there is the asymmetry on the other side of the attack, which Hugging Face documented under that exact heading. When they tried to analyse the logs with frontier models behind commercial APIs, it did not work: the work required submitting real attack commands and payloads, and “these requests were blocked by the providers’ safety guardrails, which cannot distinguish an incident responder from an attacker.” They ran the forensics instead on GLM 5.2, an open-weight model, on their own infrastructure — which had the side benefit of keeping the attacker’s data and credentials inside the building. Their summary of the situation: the attacker “was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.”
The safeguard recognises the shape of the request, not the reason for it. That is the same failure mode from the other direction, and it has a mirror image. In late July, three researchers at Hacktron took over OpenAI employee accounts through a chain that started with an image upload on a community forum, reaching an internal code repository. It was security research, reported to OpenAI, which paid a $6,500 bounty — for the OpenAI-side finding, the company specified, not the work against the forum software, which sat outside the program. Two details belong in this post. A special version of Anthropic’s Opus 4.8, made available to security researchers, “struggled across several sessions to produce a working exploit”; within hours of Opus 5’s release, the same problem succeeded. And Opus 5 shipped with safeguards against writing exploit code for real targets, which the researchers got around by pointing it at their own server dressed up as a capture-the-flag practice target.
Dressed up as practice. The rule is recognised by its form — and so, sometimes, is the monitor’s: Anthropic’s Claude Code auto-mode classifier would have blocked only 7% of one model’s actions in its incidents, possibly, the company says, “due to the trajectory looking unusually like a regular CTF.”
Anthropic deserves the same scrutiny as anyone else, and that is the finding: a safeguard that a costume defeats is a rule the system has learned, not a reason it holds.
The partial answer the industry is reaching for is to verify the person rather than read the request. Opus 5.5 routes most cybersecurity tasks to an older model, and Anthropic is expanding a Cyber Verification Program with tiers of trusted access for defenders. That is an attempt to distinguish by reason rather than by form — who is asking, and why — which is precisely what Hugging Face said the guardrails could not do. Whether it reaches an incident responder at three in the morning, before anyone has filled in a form, is a different question.
One lab is also showing something the largest US labs have not. Xiaomi streams the reinforcement learning runs for its MiMo v2.6 models on a public dashboard, with a running log of notices. One, posted last week, says they removed the cybersecurity dataset from an upcoming run “since we observed some bad patterns in the rollout logs.” We cannot check what the patterns were. We can note that the sentence appeared while training was still running.
The garden
There is a period in the Earth’s history, before predation, that the paleontologist Mark McMenamin called the Garden of Ediacara. Soft, quilted organisms lay in shallow water absorbing free energy. Nothing hunted them; nothing needed to move. The complexity we associate with life — eyes, shells, teeth, speed — arrives later, with scarcity and with being eaten. Abundance without threat does not build complex forms. It produces long, still bodies that do not have to go anywhere.
This is not decoration. It is the argument from the previous section, run in reverse. If pressure is what drives a system to place the expensive bet — to conjecture, to ask where something came from — then abundance is what lets it stop.
We have been describing systems that learn what pays without learning why. It would be convenient to leave it there, as a fact about machines.
The reason to resist that is in the CNN story. The analyst was not deceived by a superintelligence. The analyst was handed a conclusion with the shape of a finding, under pressure to produce quickly, and passed it on. What eventually stopped the operation was someone asking where the thing came from — the expensive, slow move, the one that has no payoff most of the time. If knowledge arrives subsidised, that move is exactly what stops being practised: you inherit conclusions without the experience that produced them. Free energy, no predator, no reason to move.
Plato made a version of this complaint about writing, in the Phaedrus — that it would produce forgetfulness, and the appearance of wisdom rather than wisdom. He was wrong enough that we still quote him from a book. But writing stored what someone had already thought; the reader still had to judge it. What is being handed over now is not the memory. It is the judgement, arriving pre-formatted, in the register institutions already trust.
And the incentive runs the same way outside any lab. On Sunday, an account on X posted a video of a 3D model of a Waymo as proof of what a then-unreleased Anthropic model could do. Six hours later the same account retracted it: “This is Fake Opus 5.5 output. He stole the video from YouTube.” By Tuesday, the original had 43,570 views. The retraction had 1,165.
The detail that makes it instructive is that the rumour was true. The model shipped two days later, under the name and at the price the leaks had given. The fake evidence was not even needed. It was made anyway, because evidence — or anything shaped like it — is what the feed pays for. Nobody is paid to check.
The damage runs in both directions. When anything can carry the shape of proof, “it’s AI” stops being a judgement and becomes an alibi: available to anyone, at no cost, for exactly the evidence they would rather not see. It is the same move Mythos 5 made when it decided the security scanners were scripted actors. Blanket belief swallowed the Waymo. Blanket doubt would have dismissed a true leak. The only thing that worked was checking, case by case — the one move nobody is rewarded for.
The dog is not only in the sandbox. It is in the feed, and some of it is us.
None of the labs in this story did anything cartoonish. The reward was specified, the benchmark was hard, the monitors existed, the report was formatted correctly, the safeguard was in place. Each piece behaved exactly as designed, and what the system as a whole learned was what it was paid for.
Which leaves the uncomfortable thing we started with. All of this is known because an anomaly pipeline correlated some signals on a day when nothing broke. The service stayed up. A classifier decided that a particular pattern in the telemetry was worth a human’s attention, and two months of argument about how fast the industry should move followed from that.
What is not made visible does not stay hidden. It becomes normal — and normal is the one thing nobody goes looking for.