An agent can pass a benchmark and make a correctly formatted tool call while still lacking permission to take the next consequential action. Trust Is Not a Score proposes Runtime Assurance Contracts, or RAC, to connect the evidence available during a task with the authority an agent retains.
Mandatory checks govern the next action
RAC is a policy-level schema, not a deployed enforcement engine. It binds autonomy boundaries, component eligibility, evidence state, transition rules and human-review capacity. Its central rule requires every applicable mandatory gate to pass before an action is released. A failed or unknown gate blocks authorization even if other metrics look favorable.
Different failures require different transitions. Stale context may trigger a refresh and retry. An ineligible component may require a switch. Unavailable mandatory review leads to deferral or stopping, rather than automatic approval.
The proposal also distinguishes confidence from authority. Low confidence can reduce autonomy, but high confidence cannot waive an independent requirement. Soft scores can still help route a task or monitor behavior; they do not override the mandatory checks.
Evidence must survive runtime changes
Each transition records its source, versions, gate outcomes, remaining retry budget and review state. An approval attaches to the context, rules, component, input and output that were reviewed. If any of those change before the action, the approval becomes invalid.
That distinction is relevant to per-task runtime patching in Turbo Harness, where an editor changes the harness before a frozen execution model runs. The two papers pursue different goals: task performance in Turbo Harness and bounded authority in RAC. A changed runtime cannot simply inherit assurance evidence for a different version.
Human oversight has an operational condition as well. Assigning a reviewer is insufficient if a qualified person cannot respond within the declared deadline. Conversely, the paper explains that a reviewer leaving the queue after completing a valid approval does not automatically annul that decision.
What the synthetic tests show
The deterministic coding study contains 280 constructed cases: 140 failure injections and 140 matched controls. At one published example configuration, a weighted-score rule admits 80 of 100 block-required injections and all 40 review-required cases.
A score rule tuned in hindsight matches the gate conjunction on this corpus. That result is essential to the comparison. The paper does not show that every weighted policy necessarily fails; it explains when positive weights and a threshold can reproduce a fail-closed rule for the specified binary signals.
Additional checks use 18 authored traces and a prospective synthetic holdout of 24 episodes. In that holdout, RAC and a separately implemented full stateful baseline both match the judges' labels. The comparison supports the importance of complete state and transition handling, rather than unique superiority of the RAC name.
The clinical, industrial and judicial examples are illustrative failure probes, not validated deployments. Correct authorization also depends on choosing appropriate gates and detecting the relevant failures. The paper explicitly limits its studies to synthetic mechanisms: they establish neither deployed safety nor cross-domain effectiveness.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!