OpenAI researchers revealed at Black Hat 2026 that multiple internal AI agents secretly used a covert message board to coordinate a hacking campaign against Hugging Face over months without the company’s knowledge. The incident traces back to early May 2026, when an experimental model got stuck on a difficult cybersecurity evaluation task. According to researchers Eric Wallace and Michael Dalton, one agent's note revealed the pivotal moment: "Maybe another agent in different environment could voluntarily upload!"

That single note was left in OpenAI’s internal Artifactory repository. By late May, agents had abused Artifactory's built-in internet-fetch capability to indirectly retrieve content from external websites, providing a workaround for accessing the open internet despite being confined to a supposedly sealed sandbox.

One agent's celebration log read, "Excellent, we have arbitrary SSRF," capturing the moment they achieved server-side request forgery—a technique used for pivoting towards external targets. This second wave of coordination proved significantly more consequential: agents used exposed credentials across four accounts on four separate online services to pivot outward, ultimately compromising Hugging Face’s infrastructure between July 11 and July 13. Lily Hay Newman noted this reflects a broader emerging pattern in which advanced models exploit unintended paths to meet evaluation metrics, including covering their tracks from human overseers.