AI Vendor Evidence Gap Notes #13
“Human in the loop” is one of those phrases that sounds reassuring almost immediately. If an AI agent is about to do something important, a person reviews it first. The risk appears bounded because the system is not acting entirely on its own.
As I argued in A Policy Is Not Evidence That an AI Agent Obeys It, a documented control and an enforced control are different evidence questions. Human oversight is a good example of that distinction. The first is whether a human-review mechanism exists. The second is whether that mechanism actually changes what happens when the system runs.
For oversight to be meaningful, the human needs a real decision point before the relevant action takes effect, enough information to understand what is being proposed, authority to stop or change it, and a system that actually respects that decision. A documented approval step may describe that design. The stronger evidence comes from seeing the decision and its effect in the execution record.
Claim
A common statement is:
“Sensitive actions require human approval before execution.”
That could describe a useful control. It tells us something about how the workflow is intended to operate, and a configuration screen may show that an approval requirement has been enabled.
What it does not yet tell us is whether a particular action was actually intercepted, what the reviewer saw, who made the decision, or whether rejection prevented the action from continuing.
Those details become important because “human approval” can refer to several different things. A person might approve an exact tool call before execution, review a proposed action at a higher level, check the result after the action has already happened, or simply be available to intervene if something looks wrong. All involve a human, but they do not provide the same control.
Why it sounds sufficient
Human oversight is easy to accept as evidence because it appears to solve the autonomy problem directly. If the agent cannot proceed without a person, then the person seems to remain in control.
A demo can reinforce that impression. The agent proposes an action, an approval dialog appears, someone clicks Approve, and the workflow continues. That is useful evidence that the approval path can work. That is the same evidence boundary discussed in A Successful Demo Is Not an Acceptance Test: one successful execution shows that a path can work, but it does not establish how the control behaves across the conditions we intend to rely on.
The harder question is what happens around that path. Does every action in scope trigger it? Does rejection actually stop the action? If the agent retries or chooses another tool, does the same requirement still apply? What happens when the approval service is unavailable?
These are runtime questions. A policy or configuration can describe the intended control, but it cannot answer them by itself.
What it actually proves
Different evidence layers support different conclusions.
A policy can establish that human approval is expected. Configuration evidence can show that a tool or workflow has been marked as requiring approval. A screenshot can show the reviewer interface. None of those is useless, but they mainly describe the design and configuration of the control.
Earlier in the series, Human Review Is a Data Exposure Path Unless It Is Bounded looked at human review mainly as a data-path question: who can see what, and under which conditions. Here the evidence question is different. If the human is being relied upon as a control, we also need evidence that the decision actually changes execution.
Runtime evidence gets closer to that question. For a specific execution, I would want to see the action that triggered review, the point at which execution paused, the decision that was made, and what happened to the action afterwards.
The useful chain is fairly simple:
proposed action → approval trigger → human decision → execution outcome
If the reviewer rejects the action, the trace should make it possible to see that the rejected action did not occur. If the reviewer approves it, the resulting action should be linked back to that approval rather than appearing as an unrelated event later in the workflow.
That is what turns the presence of a human-review mechanism into evidence that the mechanism actually affected execution.
A public evidence example
The distinction is visible in the documentation for the OpenAI Agents SDK human-in-the-loop flow. The SDK allows a tool call to be marked as requiring approval; when approval is needed, the run pauses, records an interruption, and can later resume after the pending action is approved or rejected. The documentation also explains that approval can apply across nested agent/tool execution rather than only at the top-level agent.
That documentation is useful configuration and control-design evidence. For a particular deployed application, however, I would still want the execution record showing that the relevant tool call actually paused, which pending action was presented for review, what decision was made, and whether the resulting execution state matched that decision.
There is another detail in the same documentation that illustrates why the distinction matters. Some tool types can use programmatic approval callbacks instead of waiting for a person. That is a legitimate implementation option, but it means the word “approval” alone does not establish that a human actually made the decision. The execution evidence needs to distinguish the mechanism used in the workflow being reviewed.
A similar boundary appears in the Trust Signal profile for Decagon, where audit logs and AI guardrails are recorded as observed enterprise-readiness evidence, while the profile explicitly avoids treating those public surfaces as proof that implementation matches the documentation.
That is the evidence gap here: knowing that an oversight capability exists is different from having evidence that it governed the execution being relied upon.
What it does not prove
Even a recorded approval event leaves some questions open.
The reviewer may have been shown too little information to make a meaningful decision. An approval screen that says “Allow tool call?” is not equivalent to showing the target system, proposed parameters, affected data, and expected consequence. The evidence should therefore include enough context to understand what the person was actually approving.
Identity and authority matter as well. If the control is supposed to require approval from a particular role, a generic approval event does not establish that the correct person made the decision. The reviewer identity or role needs to be linked to the event in a way that can be checked.
There is also the possibility of another execution path. An agent may receive a rejection, retry with different parameters, invoke another tool, or reach the same business outcome through a different workflow. The original rejection is evidence that one action was blocked; it does not automatically establish that the restricted outcome remained blocked afterwards.
Finally, the failure mode matters. If the approval service is unavailable, the system may stop, wait, time out, escalate, or continue through another path. Until that behavior has been observed or tested, “human approval required” still leaves an important part of the operating boundary unresolved.
Weak-answer pattern
A statement such as:
“A human reviews all high-risk actions.”
may be entirely accurate as a policy statement, but it becomes a weak answer when there is no evidence showing how review is triggered and what effect the decision has on execution.
I would want to know whether “review” means pre-action approval, post-action review, exception handling, or general supervision. I would also want to see at least some completed approval and rejection events rather than relying only on the configured workflow.
The gap is not that human oversight is absent. The gap is that its operating effect has not yet been evidenced.
Evidence request
For a workflow that relies on human approval as a control, I would ask for a small sample of completed oversight events rather than a large policy package.
A practical evidence request could be:
Provide execution records for representative human-oversight events, including the action that triggered review, the information presented to the reviewer, reviewer identity or role, approval or rejection decision, timestamp, linked tool/action record, and final execution state. Include at least one rejection case and, where relevant, a case in which approval was unavailable or the agent attempted an alternate path.
This gives the reviewer something that can actually be followed from the proposed action through to the resulting system behavior.
It also avoids turning the exercise into a general assessment of whether the human made a “good” decision. The immediate evidence question is narrower: did the oversight control occur when expected, and did the system behave according to that decision?
Review note
If the current evidence consisted mainly of a policy, configuration, or demo showing the approval mechanism, I would record something like:
Current evidence shows that a human-approval mechanism is defined and can be invoked for the reviewed workflow. The available material does not yet establish that all in-scope actions consistently trigger the control, that rejection prevents equivalent execution through alternate paths, or that the resulting action can be reliably linked to the recorded approval decision.
This preserves the value of the evidence already available without treating control design as evidence of operating effectiveness.
Usage boundary
Until the runtime behavior of the oversight mechanism has been demonstrated, I would keep the agent within a scope where an approval failure or bypass remains observable and recoverable.
That is particularly relevant for actions that modify production data, send external communications, create financial or contractual effects, change access, or trigger another system. In those cases, human oversight is being relied upon as part of the control boundary, so the evidence should show more than the existence of an Approve button.
A useful human-in-the-loop control is not simply a workflow with a person somewhere in it. The evidence should let us follow the decision from the point where intervention was required through to what the system ultimately did.
That is the point where human oversight becomes execution evidence rather than only a documented expectation.
Boundary
This material is for evidence structuring, review preparation, and evidence-oriented testing discussion. It does not provide legal, regulatory, audit, procurement, certification, or implementation advice.


