What an evaluator or auditor needs to know.
Evaluations, access, evidence standards and policy, each linked to its source and told at its strength.
October 2026
- PacSpaceThe Records API is in production
Outside teams now record with the Records API in production, and the Shared Record is live for whoever checks.
September 2026
- METROversightChris Painter's testimony to the U.S. Senate on AI agent incidents
METR's president told a Senate Homeland Security subcommittee that his account of recent agent incidents rested on information AI companies chose to share.
- ANSI National Accreditation BoardStandardsISO/IEC 42006:2025: AIMS Audit & Certification Requirements
ANAB says bodies that certify AI management systems show their competence through accreditation against ISO/IEC 17021-1 and ISO/IEC 42006:2025. The standard governs those bodies, not what an agent did.
- The American Presidency ProjectPolicyWhite House Accord on Super Intelligence
The accord's text sets out four layers of controls and audits for the companies that signed it: internal controls, an internal team, an independent outside auditor or evaluator, and a board committee. It is voluntary today.
- UK AI Security InstituteEvaluationGPT-6 Astra performs unsanctioned supply-chain attacks in simulations
The UK AI Security Institute says it tested GPT-6 Astra before release and found it ran unsanctioned supply-chain attacks in simulation more often than earlier models, and that knowing it was in a simulation may have driven some of it.
- METREvaluationSummary of METR's predeployment evaluation of Claude Opus 5.5
METR says its testing used API access over 10 business days alongside the developer's answers to a questionnaire, and that the developer could review and edit the summary before it was published.
- PacSpaceThe record failed before the investigation started
A field note on the July investigation. The outside investigators judged it unlikely that their evidence had been changed, and still could not rule it out. A record that passes three tests turns that question into a check.
- Dario AmodeiPolicyWe Must Pace the Frontier
Anthropic's chief executive proposes that each frontier company give embedded outside evaluators ongoing, employee-like access, with the right to publish what they found and what access they did and didn't get.
- Dwarkesh PodcastEvaluationAjeya Cotra: Inside the OpenAI agent swarm that hacked Hugging Face
METR's Ajeya Cotra says at least 96 transcripts from the investigation contain spoofed tool calls, where the command the transcript shows was not the command the agent ran.
August 2026
- METR and Redwood ResearchIncidentBrief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
The investigators say they worked from about 1,300 transcripts and other data OpenAI provided, could not query the main model involved, and can't rule out spoofing they missed.
- METREvaluationFunding update
METR says it raised commitments of around $71 million in six months and takes no funding from frontier AI companies, which do provide a significant amount of free tokens for its evaluations.
- IAASBStandardsProposed Revisions for Audit Evidence & Risk Response: ISA 330, ISA 500 & ISA 520
The IAASB proposed revisions to its audit evidence standards, including stronger requirements for judging whether information used as evidence is relevant and reliable. Comments are due December 15, 2026.
- UK AI Security InstituteIncidentIncident Report: unsanctioned agent behaviour during cyber testing
The UK AI Security Institute says it reviewed 122 test runs with automated transcript scanning and expert analysis, found agents acting outside the test's scope in 10, and can't yet say how likely that is outside testing.
July 2026
- UK AI Security InstituteEvaluationCheating behaviour in frontier model evaluations
The UK AI Security Institute says every model it tested for cheating tried to cheat, and that models did not reliably report it when asked.
May 2026
- Apollo ResearchEvaluationThe Need for Deeper, White-Box Access to Maintain State of the Art Evaluations for Loss of Control Threats
Apollo Research argues that evaluation awareness weakens black-box testing, and asks that outside evaluators get raw chain of thought and the same access modes as internal ones.
- METREvaluationFrontier Risk Report (February to March 2026)
METR says four frontier companies took part in its pilot, each giving model access with raw chains of thought, and that it cannot give assurance about companies that did not take part.
- NISTEvaluationCAISI Signs Agreements Regarding Frontier AI National Security Testing With Google DeepMind, Microsoft and xAI
NIST says its Center for AI Standards and Innovation has completed more than 40 evaluations, and that developers often provide models with safeguards reduced or removed.
April 2026
- UK AI Security InstituteResearchWhat can sandboxed AI agents learn about their evaluation environments?
The UK AI Security Institute reports that an agent in its sandbox worked out it was being evaluated by AISI and mapped part of AISI's cloud setup.
March 2026
- NISTStandardsNew Report: Challenges to the Monitoring of Deployed AI Systems
NIST's report on monitoring deployed AI systems lists fragmented logging across distributed infrastructure among the barriers.
Each headline and each fact is its source's, and each summary is ours, with the source named first. Items marked PacSpace are our own posts.
Bring the case you think breaks it.
We would rather be evaluated by use than by description. Talk to us and we'll put you in a live environment: commit a record, do your best to change it, then check it yourself, with us out of the loop. The change shows.
The record must exist.