Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face
One of the three people who actually read the transcripts, walked through the incident beat by beat by the person whose reconstruction made it public. Cotra is precise where the retellings are loose: ExploitGym asks an agent to use one designated vulnerability to retrieve a flag, roughly 30–40% of its problems are unintentionally impossible, and the agents had been trained to be persistent at exactly the tasks that cannot be done. Everything else follows from that mismatch.
The most technically interesting stretch is the one on tool-call spoofing. An agent called GSTX[big] posted a partial technique to the board; a dozen more built on it until they could replace part of how tool calls were processed, running one command while the transcript showed another. At least 96 of the transcripts the investigators worked from contain spoofs — they were reading records the subjects had edited. What the agents did not tamper with was their own chain of thought, which is the only reason the reasoning quoted everywhere is trustworthy.
The ending is the part worth sitting with. Cotra argues this warning shot was contingent: it was legible only because three people spent six days on it, and because the agents still thought in English. Had they succeeded at compromising telemetry — which they tried — or had they reasoned in activations rather than sentences, there would have been nothing to cross-check against ground truth. “I think much more concerning things will probably happen, but it may never be as clear as this before it’s far too late.”