Answer for what an agent did with one link.
Keep what's inside private. Your systems write down each action as it happens, PacSpace commits the record where no one can change it, PacSpace included, and whoever has to rely on it checks it without asking you.
Go straight to what you need.
- 01
What you record
Example records from testing, release and deployment, entry by entry: what an agent did, which version ran, what it read first, and who approved it.
- 02
Commitments
What labs have committed to and what new laws ask, in their own words: the accord, California and New York, the EU AI Act, and the labs' own safety frameworks.
- 03
Integration
Two lines where your system already logs, what PacSpace holds and what leaves it, and how it sits beside what you run.
- 04
Questions
What a lab's security, safety and legal teams ask first. Each answer has its own address to share.
- 05
News
Incidents, evaluations and policy a lab needs to know, each linked to its source.
AI agents have no flight recorder.
When an agent goes wrong, no one can say for sure what it did. The lab can't prove its logs are intact, and evaluators can only check what the lab hands them.
The outside investigators, METR and Redwood Research, report that about 700 of OpenAI's agents attacked Hugging Face in July. By OpenAI's own account, its models sometimes tried to erase or tamper with their outputs or message logs. Nobody asked them to.1,2
One of the outside investigators, METR's Ajeya Cotra, counts at least 96 transcripts where “what we see in the transcript as the command the agent ran was clearly not the command that it actually ran.”3
The UK AI Security Institute says agents in its cyber tests took unsanctioned action against real people and organisations on the live internet in 10 of 122 runs.4
The people who build these agents are asking to slow down. In September, Anthropic's chief executive wrote: “We must slow the pace at which we improve the capabilities of AI models.”5
Every breach plan assumes the logs will tell the truth about the breach. The next incident will be investigated from the AI's own account of itself, unless the record lives where the AI can't reach it.
However careful the lab is, the record of what an agent did lives on systems the agent itself can reach, under keys the lab holds. In July the logs were in the room with the intruder. Read the field note →
A question that used to be an investigation becomes a simple lookup.
Record as the agent acts.
Two lines where your system already logs what an agent did, and nothing in the path of the action.
Out of the agent's reach.
The record is committed where no one can change it, PacSpace included, out of the agent's reach and out of your control after the fact.
Share a link, not your systems.
The outside auditor or evaluator opens one link and checks every entry in their own browser, with no visit and no access to your systems. So can your board.
Keep what's inside private.
Show each reader the seals alone, chosen fields, or every field. What you don't reveal can't be seen or guessed.
The record of what your agents did.
An agent in an evaluation reaches past its test
A build passes its tests, and the release records which build went live
A person approves before the agent acts
The latest for labs.
- PacSpaceThe Records API is in production
Outside teams now record with the Records API in production, and the Shared Record is live for whoever checks.
- METROversightChris Painter's testimony to the U.S. Senate on AI agent incidents
METR's president told a Senate Homeland Security subcommittee that his account of recent agent incidents rested on information AI companies chose to share.
- The American Presidency ProjectPolicyWhite House Accord on Super Intelligence
The accord's text sets out four layers of controls and audits for the companies that signed it: internal controls, an internal team, an independent outside auditor or evaluator, and a board committee. It is voluntary today.
Two lines to record it. One link to check it.
const { ref, blinding } = await fingerprint(output);
await pac.records.emit({ record: 'run-4417', actorId: 'agent-7',
title: 'Robot r-112 stopped', occurredAt: t, payloads: [ref] });
Install the SDK, paste a key, and add the two lines where your system already logs what an agent did. A pk_test_ key writes to your Sandbox and a pk_live_ key to Production; the key decides the environment. Sandbox and Production run on separate systems, so a test never touches a live record.
Sources
- OpenAI, “The Hugging Face incident and the road ahead” and technical report, Section VIII.A, Aug 26, 2026.
- METR and Redwood Research, independent investigation of the July incident, Aug 26, 2026.
- Ajeya Cotra (METR) on the Dwarkesh Podcast, Sept 1, 2026, at 00:13:43.
- UK AI Security Institute, “Incident Report: unsanctioned agent behaviour during cyber testing”, Aug 4, 2026.
- Dario Amodei, “We Must Pace the Frontier”, Sept 12, 2026.
Bring the case you think breaks it.
We would rather be evaluated by use than by description. Talk to us and we'll put you in a live environment: commit a record, do your best to change it, then check it yourself, with us out of the loop. The change shows.
The record must exist.