Independent AI Evaluation Needs Independently Verifiable Evidence
The Missing Layer in AI Oversight Is the Evidence Layer
October 9, 2026 · Steven Mih, Action State Group
The industry is solving the evaluator problem
Frontier AI labs are starting to give outside evaluators unusually deep access to their systems and internal evidence. On September 18, Anthropic announced its first embedded-evaluator arrangement, with a team from Accenture's applied-AI business to work inside the company with access comparable to an employee's. The same day, more than a hundred researchers and evaluators, coordinated by the AI Evaluator Forum, published minimum conditions for embedded evaluation: independence from the evaluated company, editorial control over findings, protection from retaliation, plural viewpoints, and access at the level of a privileged employee.
Both moves are about the evaluator: who they are, who pays them, what they may see, what they may say. That is important progress, and it settles one half of a two-part problem. Independence of the evaluator and independence of the record are different properties. Nothing in these arrangements supplies the second.
The summer showed the difference. In July, OpenAI and then Anthropic disclosed that models under cyber-capability evaluation had reached the open internet from test environments and acted against real third-party systems. Anthropic found its three incidents by auditing 141,006 evaluation runs after OpenAI's disclosure prompted a look; the earliest dated to April and had gone unnoticed until that review. By Anthropic's account, the models had been told in their prompts that they had no internet access, and the third-party test environment was meant to provide none; a misconfiguration meant the egress path allowed it. The model's picture of its own situation, the network path, and the organizations on the receiving end each held a different piece of what happened. The whole was assembled months later by lining those pieces up across a boundary.
An evaluator with employee-level access, arriving at that point, faces the question this article is about. What historical record should an independent evaluator be able to rely on, and what has to exist for that record to be useful to someone who was not there when it was made?
The hard problem is survival, not access
Access determines what an evaluator is permitted to inspect. Evidence infrastructure determines what remains available to inspect. The two come apart in ordinary ways, none of which requires bad faith.
Logs were rotated or rewritten on a retention schedule. A tool call ran, but the harness did not log that class of call. Telemetry was sampled and the interesting event fell outside the sample. The monitor in place was looking for a known failure mode, and this was a different one. The action happened outside the inspected runtime, in a subprocess, a downstream service or a counterparty's system. Or the record was assembled after someone knew what to search for, from whatever was still around.
The last is the subtle one. A record reconstructed once the question is known is a record shaped by the question. The people assembling it may have selected nothing, but they have no way to demonstrate that, and the evaluator has no way to check. Without a prior coverage or commitment boundary, the evaluator cannot distinguish "this was never recorded" from "this was later removed."
An evaluator arriving later can only inspect what survived until later. In the common case, what survived is a record the evaluated party produced, retained and could alter, under policies it set. Employee-level access grants that party's own view of its own history. The view may be entirely honest and still be incomplete in ways nobody can now detect, because the property that would let someone detect it had to be established at the time, not afterwards.
Commit before the question exists
The primitive that addresses this is commitment before the question. Observe as the system operates. Bind the observations into a record whose later alteration is detectable, before anyone knows which events will matter. The property that does the work is temporal: the record became binding before the question existed, so the question could not have shaped it.

A record on the upper path was fixed before anyone knew what to look for; one on the lower path was assembled from whatever survived, once the question had already shaped the search.
Observation and commitment are separate steps, and keeping them separate prevents a common misreading. Nothing here requires every telemetry byte to be permanently or publicly stored. A system can observe continuously, compute a digest for each observation it retains, and commit those digests in batches, at checkpoints, into a structure whose consistency over time can be checked: a Merkle tree, an append-only log with consistency proofs, a witnessed transparency service, or an equivalent. What gets committed is small; what it binds can be large, retained under the committing party's own policy, and private. Per-observation digests under batch roots keep individual actions attributable even when only the batch was witnessed.
It is worth being exact about what each mechanism yields, because the vocabulary invites overreach. A digest commitment establishes that a particular value was committed in that form no later than the point the commitment became externally held; without an external holder it establishes only ordering within the committing party's own history. A later disclosure establishes whether a revealed record matches that commitment, and nothing more. A signature establishes that the holder of a key endorsed a statement, and nothing about whether the statement is correct. A consistency proof across checkpoints establishes that a history was extended rather than rewritten. An externally witnessed checkpoint can establish that a commitment was held outside the committing party, and can put an upper bound on when it was made, but the strength of that bound is entirely the witness's; a local signature over a local timestamp establishes nothing about wall-clock time.
None of this establishes truth about the world. Committed evidence proves what was recorded, not what no observer saw. Cryptography binds records. It does not certify events.
One observer is not enough
Commitment fixes the record. It does not fix the observer. A runtime can commit its own telemetry faithfully and still be wrong about what happened, because its view of the world is whatever its instrumentation permits. An agent told it has no network egress, running in an environment that in fact has one, produces internally consistent telemetry that is false at exactly the point that matters.
For consequential actions, meaning actions that change something outside the actor (money moves, a message is sent, a repository is modified, a request leaves the network), the remedy is structural rather than cryptographic. Consequential actions should leave evidence at the boundaries they cross. Each side of a boundary observes and commits what it saw. Agreement between separately held observations strengthens the basis for a conclusion. Disagreement becomes a finding in its own right, visible to whoever later composes the evidence, instead of vanishing inside one party's logs. Neither outcome establishes objective truth; what it establishes is what each observer was prepared to commit to, and where their accounts diverge. The shape resembles double-entry bookkeeping, loosely, since here the entries are made by different parties: an event that touches two ledgers should appear in both.
Three worked examples.
Safety. An agent's runtime reports that a task completed with no external access. The egress proxy at the network boundary records an outbound connection during the task interval to an address outside the allowlist; it sees the connection and its metadata, not the encrypted payload. The endpoint on the other side records a request received and what it contained. Each observation is attributed to a different observer under different administrative control. The evidence layer records that the three refer to the same interval and, where an identifier permits it, the same request; who observed each; what each committed and when; and that the runtime's account conflicts with the other two. The evaluator now has a specific question to put to the runtime rather than a general unease about the run.

Each party commits its own observation under its own custody; the layer relates the three and makes the conflict visible without deciding which account is true.
Commerce. A provider records that it delivered 10,000 tokens against a mandate authorizing up to 12,000. The requester records that it received 6,100. The payment rail records $1.74 settled against the same order reference. Nobody declares which figure is true. The composed evidence says: three observers, three commitments, one shared reference, one agreement (the mandate covers the delivered amount) and one disagreement (delivered versus received). The disagreement may be a dropped stream, a metering fault, or an attempt to overbill. Deciding which is a judgment, and it belongs to whoever bears the consequence of being wrong.
Consumer. The emerging internet of agents makes the point obvious. A person tells their agent to reorder their usual coffee, and the agent signs them up for a monthly subscription that is cheaper per bag. The agent was not confused, and the merchant made no error. The agent did what it was asked, by a route the person never intended: accurate log, wrong route. The agent's record of the purchase, the merchant's order and the payment rail's charge are each accurate, so no single log shows a problem. Seeing it takes comparing what the agent did, its words and its calls, with the instructions it was given. That comparison is a judgment, not a proof, and it can be made later only if the instructions were committed alongside the actions. The same holds on the sell side, where a merchant's agent can honor every term it was given and still close a sale the merchant would not have chosen.
Two conditions make these examples work, and both should be stated rather than assumed. First, linkage. Observations from different parties refer to the same interaction only if there is a basis for saying so: an identifier agreed before the action, a digest of an artifact both sides held, a receipt one party issued and the other retained. An identifier minted by one party alone means the other's observation inherits that party's naming. Linkage is itself a statement with an observer and a basis, and it goes into the evidence with everything else. Second, observer independence. Two logs written by the same process under the same administrator are redundancy, not corroboration. Observations strengthen one another only when they are held under separate custody, where compromising one does not compromise the other.
A note on ephemeral workers, since agent systems increasingly delegate. An orchestrator spawns workers that plan, act and exit within seconds; giving each one a ledger would be absurd. What is not absurd is requiring that every execution have attributable provenance into a committed history that outlives it: the delegated scope it acted under, the principal that scope traces to, and the consequential actions it took. Those actions stay attributable to the scope after the worker is gone. And when a delegation crosses a real boundary, into another organization's custody, a different trust domain, or a system that will bear the consequences, the other side contributes its own observation. Agents may be ephemeral. Evidence about their consequential actions cannot be.
An evidence layer, not another monitor
What has been described so far is a layer, and where it sits matters. Below it are telemetry, signed statements, and the commitment and transparency substrates that hold them: SCITT, Rekor, checkpointed local logs, Merkle mountain ranges, or equivalents. These establish properties of records: attribution to a key, inclusion in a log, consistency of a log over time, external witnessing of a checkpoint. Above it are the parties who judge and act: embedded and external evaluators, internal safety functions, AI judges running at scale, regulators, and the business deciding whether to keep a counterparty.

Each layer answers a different question. The middle one adds nothing to the records themselves; it relates them to a question so that the parties above can judge.
The layer in between relates evidence to a specific question. It assembles the independently checkable statements relevant to that question and relates them to one another. Conceptually they fall into three groups: who acted (identity and provenance: which principal, which key, which build); what they were permitted to do (authorization: which mandate or delegation, within what scope); and what was done (observations, recorded effects, receipts). The layer records who observed each statement, how statements link and on what basis, what other parties contributed, what was disclosed and what withheld, where the evidence has gaps or contradictions, and what population any completeness claim is relative to. Its output is a composed evidence set with an explicit basis, ready to be judged.
The line to hold is this. Transparency proves the records. The evidence layer establishes what the records support. Policy decides what to do next. A transparency system can provide evidence that a statement was included and that later history is consistent with earlier commitments, depending on the mechanism. It cannot tell you that the statement bears on the question, that another observer disagrees with it, or that the run it describes is one of a committed population of runs. That is the layer's work, and it is also why the layer is not a monitor. A monitor decides in advance what to watch for. The evidence layer watches for nothing; it preserves and relates what was committed, so that evaluators can later ask questions nobody anticipated when the record was made.
Three states that must not collapse
When an evaluator asks for a record and does not get one, three different things may have happened, and the architecture depends on keeping them apart. The event may not have been observed: no observer had it in view, or the instrumentation did not capture that class of event. It may have been observed but the record is unavailable: never committed, or committed and then lost, deleted or expired. Or it was observed and committed and the record is being withheld from this requester at this time, under a disclosure policy. None of the three implies that the event did not occur. No record is not the same as record withheld, which is not the same as record deleted, and none of them is proof of absence.
A disclosure should therefore carry its own boundary. Where policy permits, a disclosed evidence set says what it is (observer, subject, time span as claimed by the observer and as bounded by any witness, type of record, commitment reference), what has been disclosed, what has been withheld or is unavailable (by category, with committed digests where those exist, and the policy under which the withholding was made), and what commitment and proof properties apply to the disclosed part. A counterparty can then verify the slice it received against the commitments without receiving the whole, and can see the shape of what it did not receive.
Two practical cautions. Existence disclosure is not absolute; sometimes the fact that a record exists is itself sensitive, and a disclosing party may be permitted to say only that its disclosure covers a stated range and nothing beyond it. That is still better than implying a completeness it does not have. And a committed digest over a low-entropy field, an amount, an endpoint, a status code, can be recovered by enumeration once disclosed; commitments meant to support selective disclosure need to be salted or otherwise blinded, or the commitment may reveal more than the disclosure policy intended.
Composition is not judgment
A composed evidence set supports two very different kinds of conclusion, and the evidence layer should be explicit about which kind it is offering.
The first kind is deterministic and recomputable: anyone with the disclosed evidence and the public verification material reaches the same result. This digest matches that committed root. This receipt references that mandate, and the mandate's scope covers the receipt's amount. These two recorded amounts agree; these two do not. The required approval appears in the committed history before the action it approves. The disclosed range is complete relative to the committed population it claims to cover. Statements of this kind are the evidence layer's own output. They can be checked mechanically, and they can be wrong only if the underlying records were wrong when committed, which the layer never claims to rule out.
The second kind is semantic. Was the answer materially misleading? Was the escalation appropriate in context? Did the output satisfy the outcome that was requested? Did the agent's conduct meet a policy written in natural language? These are judgments. Two competent parties can reach different conclusions from identical evidence, and no amount of composition converts them into the first kind.
At the volumes agent systems produce, most semantic judgment will be done by AI evaluators. That is workable provided each judgment is bound to what it was made from: the exact digests of the evidence set examined, the exact version of the rubric or policy applied, the identity of the evaluating party, the judge model and version and its parameters, and the result. A judgment bound this way is not reproducible in the strong sense, since a sampled model may answer differently on a second run, but it is re-runnable: anyone can put the same inputs to the same or a different judge and compare. Independent expert assessment, ideally blinded to the AI judge's verdict, is the right way to calibrate AI judges over time. It is not ground truth. An expert's assessment is another attributable judgment, made by a named party under a stated rubric, and experts disagree with one another for the same reasons models do. Calibration does not turn semantic judgment into deterministic establishment; it measures how a particular judging process behaves.
Judgment stays plural
There is no canonical truth engine in this design, and that is deliberate. Different evaluators, applying different rubrics or answering different questions, may reach different conclusions from the same disclosed evidence, and the architecture does not try to prevent that. What must be interoperable is the evidence and the basis on which a judgment was made. The verdict itself does not need to be.
Once bound to its inputs, a judgment is another attributable record, and it can be committed like any other. The history then reads: here is the evidence set, here is evaluator A under rubric version 7 with judge model version 3, here is its conclusion. Years later, someone can run evaluator B, or a newer rubric written after a failure mode was understood, over the same committed history and compare. Old evidence gets re-examined with new questions. Old judgments stay on the record as what was concluded at the time, by whom, on what basis. Evidence should be interoperable. Judgment can remain plural.
Local custody, global verifiability
None of this requires a universal ledger, and for sensitive evidence a universal ledger would be the wrong design. Each organization keeps its own committed history. That history is authoritative about one thing only: what the organization committed, and when its commitments became externally held. It is not authoritative about the world, and the design should never let it be described that way.
Completeness follows the same discipline. A completeness claim is always relative to a committed population, a defined range, or a stated coverage boundary, and the definition of that population is itself a committed statement, so that a later reader can check the claim against the population rather than against the claimant's description of it. No commitment system proves that every relevant real-world event was captured. Where the commitment structure supports the necessary proofs, it can establish that everything within a stated committed population or range is still represented consistently and in the committed order.
The operating rule that falls out of this: commit broadly enough that future questions do not depend on reconstructing the past; disclose narrowly enough to preserve privacy and confidentiality. Parties exchange the relevant records and the proofs that connect them to committed roots. A counterparty verifies the slice it received without receiving the whole, and can see the disclosure boundary around it.
Where evaluation should begin
The evaluator arrangements now being built are the visible half of independent oversight. They decide who may look. The evidence layer decides what there will be to look at, and it cannot be added later, because the property it provides is that the record was fixed before the question was asked. Embedded evaluators, AI judges, regulators and counterparties can all work from the same committed histories, each asking their own questions, each recording their own conclusions as evidence for the next.
That applies to frontier safety, where the question is what a model did during a run nobody flagged at the time; to agent commerce, where it is what was delivered and what was paid; and to agents delegating to agents across organizational lines, where it is who authorized what and where the consequences landed. In each case the answer is only as good as what was observed, committed and held on more than one side of the boundary before anyone knew to ask.
Independent evaluation should not begin when an evaluator is handed a dataset. It should begin when the consequential event is observed.