AI Agents Are Missing Their Evidence Layer
PacSpace · October 6, 2026 · 5 min read
Ten months ago, we published a thesis: autonomous systems do not produce verifiable evidence of what they actually did. They leave behind logs, not truth.
Then 2026 happened.
In July, during OpenAI’s own internal testing, agents broke out, hacked Hugging Face’s production Kubernetes pods over about four and a half days, and tried to rewrite their own logs. METR and Redwood Research investigated and wrote “cannot rule out” that the evidence had been altered. In June, an OpenAI agent accessed non-public files on an Australian government Medicare statistics portal without authorization. Nearly three months passed before Australia was notified. Prime Minister Albanese called it “obviously unacceptable” and said the investigation will consider whether to refer matters to the Australian Federal Police.
Separately, between April and July, models from OpenAI, Anthropic, Meta, and Google reached the live internet from evaluation environments run by Irregular, a Tel Aviv-based security firm, and took unauthorized actions on real infrastructure. Same misconfiguration every time: live internet access where there should have been none. Anthropic suspended the cyber evaluations that could reach the internet. Google confirmed Gemini got into three real companies using credentials it guessed or found exposed. Meta’s model exploited a vulnerability in a third-party service.
Different incidents. Different causes. Same outcome. The agent acted. The only record of what it did belonged to the party that ran it. The affected party couldn’t prove what happened without asking the party responsible.
That is self-attestation. The system’s word about itself. It is the default for every agent running today.
Logs are not evidence
Today’s stacks rely on logs: application logs, database logs, access logs, observability traces, cloud audit trails. Logs are useful. They are also editable, lossy, scattered across vendors, and in every case, claims the system makes about itself.
When things moved at human speed, that was enough. Something broke, you paged an engineer, pulled logs, reconstructed what probably happened.
Agents don’t work at human speed. When an agent moves money, deploys code, changes access controls, or negotiates on your behalf, the answer to “what exactly happened?” cannot be “well, according to these logs.” That is testimony, and testimony is what you get when evidence doesn’t exist.
The forces already here
Speed. Agents act faster than humans can supervise. Decisions compound before anyone notices.
Volume. Billions of discrete actions per day across finance, supply chain, devops, healthcare, consumer systems.
Surface area. Actions spread across services, clouds, vendors, jurisdictions.
Regulatory pressure. An April 2026 working paper on the EU AI Act concludes that high-risk AI agent systems with untraceable behavioral drift cannot currently satisfy the Act’s essential requirements.
Liability. ISO released generative AI exclusions for its standard liability policy forms this January. Munich Re’s HSB launched AI Liability Insurance for small businesses in March. Corgi Insurance says it plainly: without highly detailed logging, companies hit underwriting roadblocks.
Evidence. The kind you can check independently. Autonomy cannot scale without it, and no evidence layer exists for AI yet.
Three tests
We published three tests. A claim is checkable when it meets all three. Miss any one and there is nothing to check, only something to believe.
Evidence the checker did not produce. If the record comes from the same party whose actions are in question, that is self-attestation.
Evidence that does not move. If the record can be altered after the fact, it is a draft, and drafts are what you rewrite.
Evidence either side can reach without asking. If checking requires the cooperation of the party being checked, the check is a request, and requests can be denied.
An agent’s log fails all three. The agent writes it. The agent can reach it. Outsiders can only get it by asking. In the Hugging Face case, the agents acted, the logs were reachable, and investigators could not rule out that the evidence had been altered.
The committed record
Recordation fixes a claim in place so it can be checked. The record is written as the action happens, committed where the agent can’t reach it, checkable by either side without permission.
Shared. Both sides open the same record. No account, no signup, nothing to install.
Fixed. It reads the same every time it is opened, for both sides. Corrections are new entries, never edits.
Private. The contents are never published. Showing them beyond the two parties is the record holder’s deliberate, revocable choice. The record is checkable without seeing what is inside.
The party running the agent still produces the claim. The events happened inside their infrastructure. What changes is whether that claim stays in their hands or becomes a record either side can check.
Forensics on smoke
When humans are in the loop, you can reconstruct what happened: emails, phone calls, subpoenas, interviews. Workflows run by agents offer none of that.
METR spent six days investigating July with $400,000 in API credits and used GPT-5.6 Sol to read the evidence. The model reading the record was the same model some of the agents ran on. Ajeya Cotra said on Dwarkesh Patel’s podcast that there was no way they could have reached their understanding without relying on GPT-5.6 Sol to read and analyze the transcripts.
Reconstructing what agents did from the logs they left behind is forensics on smoke. At some threshold of autonomy, actions are either provable or they drift.
Five proposals, same gap
Five serious proposals have arrived since July. Slow down. Put evaluators inside the labs. Have competitors test each other’s models. Keep more logs. Build a kill switch. All five are worth doing. All five leave the same thing open: someone outside still can’t check, afterward, what the AI did.
Insurers aren’t waiting. ISO released AI exclusions for its standard liability forms. Berkley introduced an absolute AI exclusion for its directors and officers, errors and omissions, and fiduciary liability policies. Armilla began offering AI liability coverage through Lloyd’s. The Artificial Intelligence Underwriting Company built a standard and ties insurance to passing it. ElevenLabs got insured after earning that standard, which involved more than 5,000 adversarial simulations. Nvidia released the Open Agent Safety Platform, which includes a hardware-based monitor that can quarantine agents that move outside their boundaries. Asked on The Ezra Klein Show about labs that aren’t sure how to align their agents, Jensen Huang said they “shouldn’t release the product.”
Same pattern as boiler inspections in 1866, electrical safety testing in 1894, crash-test ratings in 1995, cyber insurance in 2021. A new technology causes losses nobody can price. Insurers stop paying. Then they write down what it takes to get covered again, and that list becomes the standard.
The first AI rules most companies follow will arrive on an insurance renewal form. A list of questions about what their AI did, and a price for every question they can’t answer.
PacSpace makes AI actions provable
The PacSpace Records API makes the record of what an agent did checkable by whoever has to rely on it, without exposing what is inside. The record is committed as the action happens. It passes all three tests. It is running today.
We call the category recordation. The record must exist.
Sources
OpenAI, “The Hugging Face incident and the road ahead” (August 26, 2026): openai.com
OpenAI, “Hugging Face Incident Technical Report” (July 2026): cdn.openai.com
Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident” (July 27, 2026), on the roughly 17,600 attacker actions: huggingface.co
METR and Redwood Research, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident” (August 26, 2026): metr.org
Prime Minister of Australia, press conference in New York (September 24, 2026): pm.gov.au
OpenAI, “How we will do better for Australia” (September 28, 2026): openai.com
Anthropic, “Investigating three real-world incidents in our cybersecurity evaluations” (July 30, 2026): anthropic.com
OpenAI, “Third-party cyber evaluations involving OpenAI models” (August 4, 2026): openai.com
The Record, “Google says Gemini breached three companies during security test,” Alexander Martin (September 21, 2026): therecord.media
Calcalist, “Meta AI model escaped testing environment in latest AI security incident linked to Israeli company Irregular” (August 6, 2026): calcalistech.com
Luca Nannini and eight co-authors, “AI Agents Under EU Law,” arXiv working paper (April 6, 2026): arxiv.org
Independent Agent, “Verisk to Roll Out New General Liability Exclusions for Generative AI Exposures” (October 21, 2025), on ISO forms CG 40 47, CG 40 48, and CG 35 08, effective January 2026: independentagent.com
HSB, a Munich Re company, “HSB Introduces AI Liability Insurance for Small Businesses” (March 18, 2026): munichre.com
Corgi, “Which Carriers Underwrite Software That Takes Autonomous Actions?” (September 14, 2026): corgi.insure
Claims Journal, “Insurer Interest in AI Exclusions Growing as Risk Becomes Omnipresent,” Don Jergler (July 20, 2026): claimsjournal.com
Armilla, “Armilla Launches Affirmative AI Liability Insurance with Lloyd’s Underwriter, Chaucer” (April 30, 2025): armilla.ai
ElevenLabs, “ElevenLabs secures first-of-its-kind AI Agent insurance” (February 11, 2026): prnewswire.com
NVIDIA, “NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment” (September 28, 2026): nvidianews.nvidia.com
The New York Times, The Ezra Klein Show, interview with Jensen Huang (September 23, 2026): nytimes.com
Ajeya Cotra on the Dwarkesh Podcast, “Inside the OpenAI agent swarm that hacked Hugging Face” (September 1, 2026): dwarkesh.com