What the paper is about
Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. We identify two resulting gaps. The understanding--execution gap arises when a requirement is understood but not satisfied in execution; the state--authority gap arises when an agent's interpretation or completion claim does not establish the required state. On SkillsBench, using only agent-visible prompts, workspace information, and injected skill specifications, we extract 509 source-grounded task directions. Across seven models, only 79.6%--86.4% are satisfied, while completion-claim rates exceed official evaluator pass rates by 28.7--37.9 percentage points. We therefore separate agent proposals from authoritative state. Agents may plan, act, and request completion, but only admissible evidence from qualified providers may establish specification-governed state. SpecHarness operationalizes this principle by compiling visible specifications into source-linked obligations and governing execution and finalization through versioned obligation state. Verifiable requirements are mediated or validated at runtime, while ambiguous or subjective requirements remain advisory. Experiments on guideline-following and artifact-generation tasks show that specifications can serve not merely as behavioral guidance, but as authority over compliant execution and completion. The same ai evaluation question is explored in What Should We Ask Next? Retrieval-Aware..., which adds a research perspective.
What it covers
Who Holds the Pen? Let Specifications, Not Agents, Sign Off Haiqing Li Xin Ma Yinhao Wu Wenliang Zhong Feng Jiang Thao M. Dang Xiao Hu Hehuan Ma Yuzhi Guo Junzhou Huang Abstract Large language model agents operate under external specifications, including task instructions, guidelines, output schemas, and reusable skills. In most systems, however, these specifications remain context for the same model that acts, evaluates outcomes, and declares completion. This collapses proposal and acceptance within the agent, leaving no independent specification authority boundary. This creates two structural gaps. The understanding–execution gap arises because understanding a requirement does not ensure satisfying it during execution. The state–authority gap arises because an agent’s interpretation or completion claim does not prove that the required state has been achieved. We quantify these gaps on SkillsBench. Using only agent-visible task prompts, workspace information, and injected skill specifications, we extract 509 source-grounded task directions and assess whether they are satisfied during execution. Across seven models, only 79.6%–86.4% of these directions are satisfied. Moreover, agents’ completion-claim rates exceed official evaluator pass rates by 28.7–37.9 percentage points. We formulate specification authority as a separation between agent proposals and authoritative state: agents may plan, act, and request completion, but only admissible evidence from qualified providers may establish specification-governed state. SpecHarness operationalizes this paradigm by compiling agent-visible specifications into source-linked obligations and governing execution and finalization through versioned obligation state. Reliably executable or verifiable requirements are mediated or validated at runtime; ambiguous or subjective requirements remain advisory or are excluded from enforcement. We instantiate and evaluate SpecHarness on guideline-following and artifact-generation tasks. Our results support treating specifications not merely as influences on agent behavior, but as authority over what constitutes correct execution and compliant completion. Authoritative judgments about specification-defined state should remain independent of the agent’s self-assessment. 1 Department of Computer Science and Engineering, The University of Texas at Arlington, Arlington, TX, USA 2 Monash University, Melbourne, VIC, Australia 3 Department of Computer Science, Kent State University, Kent, OH, USA Figure 1: Two structural gaps in specification following. Top: The agent understands the required order but executes the steps incorrectly, illustrating the understanding–execution gap. Bottom: The agent claims completion before the specification-required state is established, illustrating the state–authority gap. Introduction Large language model agents increasingly control complex tasks by inspecting environments, planning steps, selecting skills, invoking tools, and deciding when to stop ( Li et al. 2026a ; Wang et al. 2023 ) . Their behavior is expected to follow external specifications at three levels: task instructions, guidelines, and policies define goals, constraints, and completion conditions; tool schemas, API documentation, operating procedures, and SKILL.md files define operation semantics, preconditions, arguments, and execution order; and manifests, schemas, and contracts define required artifacts and their acceptance conditions. However, most systems still treat these specifications only as agent-readable context rather than as a basis for an external runtime to govern execution and acceptance. Agents may therefore omit requirements, violate procedures, misuse tools, misinterpret effects, combine constraints incorrectly, or claim completion before the required state has been established ( Bandi et al. 2026 ) . This mismatch produces the two gaps illustrated in Figure 1 . Requirements may be correctly interpreted yet fail to govern execution, creating the understanding–execution gap. Moreover, agent actions and self-assessments cannot authoritatively establish specification-defined state, creating the state–authority gap. The key question is therefore not only when or how execution is checked, but who may establish specification-governed state and authorize completion. We quantify both gaps on all 87 SkillsBench tasks using a compiler selected from seven candidate language models on development annotations and frozen before evaluation; the full comparison and selection protocol are reported in Appendix A.1. Using only agent-visible task prompts, workspace information, and injected skill specifications, the selected compiler extracts 509 source-grounded task directions. Across seven representative language models, only 79.6%–86.4% of these directions are satisfied, while agents’ completion-claim rates exceed official pass rates by 28.7–37.9 percentage points. These findings quantify both gaps across the evaluated agents (Figure 3 ). Some approaches ( Ouyang et al. 2022 ; Bai et al. 2022 ) internalize specification following through model training, but still leave the same model responsible for both acting under a specification and judging whether it has been satisfied. As shown in Figure 2 , inference-time methods intervene at three stages: post-hoc verification checks artifacts or final states after execution; completion gating checks completion claims before acceptance but usually leaves prior execution and intermediate state unconstrained; and runtime enforcement ( Mazzocchetti 2026 ) constrains behavior during execution, as in VIGIL ( Li et al. 2026c ) , which monitors agent–tool interactions for temporal, argument, and value-flow violations. These approaches improve verification, acceptance control, or behavioral compliance, but do not jointly govern execution and evidence-authorized state commitment. What is missing is an explicit specification authority boundary governing both execution and accepted state. We formulate this distinction as the State Authority Principle : for reliably grounded and verified conditions, the specification defines what may be accepted, while an external runtime determines whether the required state has been established. SpecHarness operationalizes this principle as a runtime architecture for specification-governed agents. Agents still interpret tasks, plan, select tools and skills, implement solutions, and repair failures, but only propose actions, state changes, and completion; SpecHarness commits the corresponding authoritative state. It converts verifiable requirements in agent-visible materials into source-linked obligations, checks them through controlled execution or authorized validation, and updates state only when the required evidence satisfies its commit rule. Ambiguous, subjective, conflicting, or unverifiable requirements remain advisory rather than becoming hard obligations, and continue to guide the agent. These mechanisms follow directly from the two gaps: operational obligations narrow the understanding–execution gap, while evidence-authorized commitment prevents self-assessment from establishing accepted state. For example, authorizing an array-processing action confirms only its preconditions, arguments, and ordering constraints; the corresponding obligation state is committed only after an authorized validator verifies the required path, shape, data type, and finite values. SpecHarness thus allows specifications not only to describe desired behavior, but also to govern which verifiable states and outcomes may be accepted. Agent proposes; SpecHarness commits. The main contributions of this paper are as follows: 1. We identify the missing specification authority boundary and its two observable consequences: the understanding–execution and state–authority gaps. 2. We formulate the State Authority Principle, separating proposal autonomy from state authority: agents propose, while an external runtime commits specification-governed state only from admissible evidence produced by qualified providers. 3. We instantiate and evaluate SpecHarness , an obligation–evidence–commit architecture governing decision, artifact, and completion state across GuideBench and SkillsBench. Figure 2: Three intervention paradigms for specification compliance. (a) Post-hoc verification detects violations only after execution has completed. (b) Completion gating checks the agent’s completion claim before acceptance but does not constrain the preceding execution. (c) Runtime enforcement intervenes during execution to prevent or constrain specification-violating actions. Figure 3: Empirical evidence for the two structural gaps across seven language models. Top: Models satisfy 79.6%–86.4% of the 509 source-grounded task directions recovered from agent-visible materials. Bottom: Agent-reported completion rates exceed official verifier pass rates by 28.7–37.9 percentage points. Related Work Post-hoc Verification (Fig. 2(a)). Post-hoc methods assess specification compliance after an answer, trajectory, or artifact is produced. NSVIF ( Su et al. 2026 ) represents natural-language instructions as logical and semantic constraints and reports interpretable violations. SkillsBench ( Li et al. 2026b ) evaluates skill-guided executions with deterministic verifiers over final artifacts and environment states. AgentRx ( Barke et al. 2026 ) synthesizes trajectory invariants, checks them step by step, and produces auditable violation logs with supporting evidence, while AgentOps ( Dong et al. 2024 ) identifies the artifacts and lifecycle data needed for observability. These methods provide evidence about completed executions, but use it for diagnosis rather than to govern authoritative task state. Completion Gating (Fig. 2(b)). Failure analyses motivate explicit control at the completion boundary. MAST ( Cemri et al. 2026 ) identifies verification and termination as a major category of multi-agent failure, including premature termination and missing or incorrect verification. Verify-gated completion ( Nguyen and Tran 2026 ) treats an agent’s completion claim as a proposal and places a read-only verifier before task acceptance, using fail-closed admission and auditable event records. Such designs block unsupported completion claims but verify primarily at terminal admission. SpecHarness instead organizes verification around source-grounded obligations and specification-governed state transitions, making completion the final commit rather than the sole verification point. Runtime Enforcement (Fig. 2(c)). Runtime enforcement constrains behavior or state transitions during execution. VIGIL ( Li et al. 2026c ) , AgentSpec ( Wang et al. 2025a ) , and FORGE ( Palumbo et al. 2026 ) enforce specification-derived policies through trace checking, runtime rules, or action mediation. Verification-gated mission-state governance ( Tang et al. 2026 ) further commits proposed mission updates only after deterministic verification. SpecHarness builds on classical reference-monitor ( Saltzer and Schroeder 1975 ) , runtime-verification ( Leucker and Schallhart 2009 ) , and transactional-commit ( Gray and Lamport 2004 ) principles rather than treating them as new primitives. Its agent-specific contribution is to compile visible specifications into source-linked obligations and use qualified evidence to govern versioned state and finalization; preventive mediation remains limited to closure-audited surfaces (Appendix B.2). SpecHarness: Runtime Specification Authority SpecHarness reallocates runtime authority: agents plan and act, while only specification-authorized evidence may establish authoritative state. Problem Setting and State Authority. For a task instance x x , let V x = ( T x , G x , K x , W x ) V_{x}=(T_{x},G_{x},K_{x},W_{x}) (1) denote its agent-visible context: task prompt T x T_{x} , applicable guidelines G x G_{x} , injected skill specifications K x K_{x} , and observable workspace state and schemas W x W_{x} . Held-out verifiers and oracle solutions are evaluation-only and excluded from V x V_{x} and runtime enforcement. Let H exec H_{\mathrm{exec}} denote the executor, observer, and validator substrate available to SpecHarness. At time t t , the agent proposes p t p_{t} , the runtime observes evidence e t e_{t} , and SpecHarness maintains authoritative state s t = ( L t , C t , ν t ) s_{t}=(L_{t},C_{t},\nu_{t}) , comprising a versioned obligation ledger, derived control state, and dependency versions. Let k i k_{i} be obligation o i o_{i} ’s assertion key and s t [ k ] s_{t}[k] its ledger value, or ⊥ \bot if absent. SpecHarness enforces s t + 1 [ k ] ≠ s t [ k ] ⇒ ∃ o i , e t : 𝖶𝗂𝗍𝗇𝖾𝗌𝗌 t ( k , o i , e t ) s_{t+1}[k]\neq s_{t}[k]\Rightarrow\exists o_{i},e_{t}:\mathsf{Witness}{t}(k,o{i},e_{t}) , where 𝖶𝗂𝗍𝗇𝖾𝗌𝗌 t ( k , o i , e t ) \mathsf{Witness}{t}(k,o{i},e_{t}) requires k = k i k=k_{i} , admissible evidence Adm i ( e t , s t , H exec ) \operatorname{Adm}{i}(e{t},s_{t};H_{\mathrm{exec}}) , and a corresponding 𝖼𝗈𝗆𝗆𝗂𝗍 ( o i , e t , k ) \mathsf{commit}(o_{i},e_{t},k) event. Admissibility is defined below. This state-write invariant defines SpecHarness: agent actions, outputs, and self-assessments cannot directly establish authoritative state. Thus, the agent proposes; SpecHarness commits . A commit may record validated success or failure; satisfaction is determined separately. Figure 4: SpecHarness compiles agent-visible specifications into source-linked obligations, mediates closure-audited actions, validates observed effects, commits versioned state from authorized evidence, and permits finalization only when all fresh mandatory obligations are satisfied. Advisory and abstained requirements remain outside hard enforcement. Runtime Obligation Construction. SpecHarness segments the visible context into source-addressable units and assigns each a disposition: 𝒰 x \displaystyle\mathcal{U}{x} = Segment ( V x ) = { u j = ( q j , ℓ j ) } j = 1 n x , \displaystyle=\operatorname{Segment}(V{x})={u_{j}=(q_{j},\ell_{j})}{j=1}^{n{x}}, (2) δ x \displaystyle\delta_{x} : 𝒰 x → { hard , advisory , abstain , residual } . \displaystyle:\mathcal{U}{x}\rightarrow{\textsc{hard},\textsc{advisory},\textsc{abstain},\textsc{residual}}. Here, q j q{j} identifies the visible source and ℓ j \ell_{j} its requirement text. Reusable policies and skill specifications may yield templates; task prompts and instance-specific guidelines are compiled per task, while the observable workspace grounds entities, paths, arguments, outputs, and completion conditions. Using development annotations derived only from agent-visible materials, we compare seven language-model compilers, select one by development-set extraction quality, and freeze it before execution; compiler details and the selection protocol are provided in Appendix A.1. Held-out official verifiers are used only afterward for post-hoc obligation–test alignment and final evaluation. Let Q x ⋆ = 𝒞 ⋆ ( 𝒰 x ) Q_{x}^{\star}=\mathcal{C}^{\star}(\mathcal{U}{x}) be the frozen compiler output and 𝒟 x = Directions ( Q x ⋆ ) \mathcal{D}{x}=\operatorname{Directions}(Q_{x}^{\star}) the source-addressable index for direction-level measurement and trace attribution. Given H exec H_{\mathrm{exec}} , the obligation builder constructs ℬ ( Q x ⋆ , H exec ) → Γ x ⋆ = ( 𝒪 hard , 𝒪 adv , 𝒰 abs , R x ) . \mathcal{B}(Q_{x}^{\star},H_{\mathrm{exec}})\rightarrow\Gamma_{x}^{\star}=\bigl(\mathcal{O}{\mathrm{hard}},\mathcal{O}{\mathrm{adv}},\mathcal{U}{\mathrm{abs}},R{x}\bigr). (3) Here, Γ x ⋆ \Gamma_{x}^{\star} is the task-specific obligation IR maintained by the runtime: 𝒪 hard \mathcal{O}{\mathrm{hard}} contains mandatory obligations that may block covered transitions or completion, 𝒪 adv \mathcal{O}{\mathrm{adv}} contains non-blocking obligations and guidance, 𝒰 abs \mathcal{U}{\mathrm{abs}} records abstentions, and R x R{x} preserves residual context. Directions and obligations need not correspond one-to-one: several directions may ground one obligation, while advisory, abstained, and residual directions create no independent blocking rules. Thus, 𝒟 x \mathcal{D}{x} supports measurement and trace attribution, whereas Γ x ⋆ \Gamma{x}^{\star} governs authorization, validation, commitment, and finalization. A grounded obligation is o i = ⟨ 𝗌𝗋𝖼 i , 𝖺𝗎𝗍𝗁 i , 𝖾𝖿𝖿𝖾𝖼𝗍 i , 𝗌𝗍𝖺𝗍𝖾 i , 𝖼𝗍𝗋𝗅 i ⟩ , o_{i}=\langle\mathsf{src}{i},\mathsf{auth}{i},\mathsf{effect}{i},\mathsf{state}{i},\mathsf{ctrl}{i}\rangle, (4) encoding provenance, action matching and authorization, execution and validation, state commitment and satisfaction, and dependency and enforcement control, respectively. An obligation enters 𝒪 hard \mathcal{O}{\mathrm{hard}} only if it is mandatory, its parameters are source-grounded, it has an authorized evidence provider, and its validator is qualified for blocking under the development protocol; otherwise, it remains advisory or triggers abstention. Validator qualification and closure stress tests appear in Appendices B.1–B.2. The builder also derives a declared action surface 𝒜 decl \mathcal{A}{\mathrm{decl}} . A closure audit identifies actions eligible for bounded mediate-and-commit ; other safely isolated declared channels use validate-and-commit , while unsupported channels are rejected. Hard classification alone does not imply preventive mediation. Moreover, admissible evidence may commit a validated failure without satisfying o i o{i} . Agent Interface and Proposals. The agent retains the original task context, applicable skills, and residual context R x R_{x} , and receives advisory guidance and source-linked runtime feedback. SpecHarness alone maintains the authoritative ledger and exposes only derived feedback and control decisions. The agent submits proposals in three tagged classes: 𝒫 = 𝒫 act ⊎ 𝒫 repair ⊎ 𝒫 finalize , \mathcal{P}=\mathcal{P}{\mathrm{act}}\uplus\mathcal{P}{\mathrm{repair}}\uplus\mathcal{P}{\mathrm{finalize}}, (5) for actions with bound arguments, repairs targeting reported failures, and finalization requests. Proposals express intent, not authoritative fact. The agent may plan, select tools, generate code, interpret errors, and revise artifacts, but cannot modify the ledger, inject validator outcomes, mark obligations satisfied, or authorize finalization. Action-bearing proposals proceed to authorization; finalization is determined solely by committed obligation state at current evidence versions. Action Authorization and Controlled Execution. SpecHarness first maps each action proposal to a canonical action. For matched proposals, it identifies the relevant obligations and returns an authorization decision: ( a ^ t , n t ) \displaystyle(\hat{a}{t},n_{t}) = Normalize ( p t ) , \displaystyle=\operatorname{Normalize}(p_{t}), (6) n t \displaystyle n_{t} ∈ { exact , normalized , unmatched } , \displaystyle\in{\textsc{exact},\textsc{normalized},\textsc{unmatched}}, M t \displaystyle M_{t} = { o i ∈ 𝒪 x act ( s t ) : Match i ( a ^ t , s t ) = 1 } , \displaystyle={o_{i}\in\mathcal{O}^{\mathrm{act}}{x}(s{t}):\operatorname{Match}{i}(\hat{a}{t},s_{t})=1}, u t \displaystyle u_{t} = Auth ( a ^ t , M t , s t ) , \displaystyle=\operatorname{Auth}(\hat{a}{t},M{t},s_{t}), u t \displaystyle u_{t} ∈ { allow , block , unclear } . \displaystyle\in{\textsc{allow},\textsc{block},\textsc{unclear}}. An unmatched proposal raises an error when mediation is expected, enters validate-and-commit on a safely isolated declared channel, follows the base policy when explicitly out of scope, or is rejected when unsupported. Let π ∼ eff a \pi\sim_{\mathrm{eff}}a mean that execution path π \pi can produce an effect equivalent to canonical action a a , and let GovernedByAuth ( π , a ) \operatorname{GovernedByAuth}(\pi,a) mean that π \pi is mediated as a a and subjected to its authorization procedure. The action surface is closed for a a if and only if every available effect-equivalent path is either governed or denied: Closed ( a , H exec ) = 1 \displaystyle\operatorname{Closed}(a,H_{\mathrm{exec}})=1 (7) ⇔ ∀ π ∈ AvailPaths ( H exec ) , \displaystyle\iff\forall\pi\in\operatorname{AvailPaths}(H_{\mathrm{exec}}), π ∼ eff a ⇒ GovernedByAuth ( π , a ) ∨ Denied ( π ) . \displaystyle\pi\sim_{\mathrm{eff}}a\Rightarrow\operatorname{GovernedByAuth}(\pi,a)\lor\operatorname{Denied}(\pi). Define 𝒜 static = { a ∈ 𝒜 decl : Closed ( a , H exec ) = 1 } \mathcal{A}{\mathrm{static}}={a\in\mathcal{A}{\mathrm{decl}}:\operatorname{Closed}(a,H_{\mathrm{exec}})=1} , and let Preventive ( M t ) \operatorname{Preventive}(M_{t}) hold iff M t M_{t} contains a hard obligation with mode block . Actions in 𝒜 static \mathcal{A}{\mathrm{static}} use bound controlled executors, and preventive matches require u t = allow u{t}=\textsc{allow} for dispatch. Closure then yields a ^ t ∈ 𝒜 static ∧ Preventive ( M t ) ∧ u t ≠ allow \displaystyle\hat{a}{t}\in\mathcal{A}{\mathrm{static}}\land\operatorname{Preventive}(M_{t})\land u_{t}\neq\textsc{allow} (8) ⇒ ∄ π ∈ AvailPaths ( H exec ) : \displaystyle\Rightarrow\not\exists,\pi\in\operatorname{AvailPaths}(H_{\mathrm{exec}}): π ∼ eff a ^ t ∧ BypassesAuth ( π , a ^ t , s t ) . \displaystyle\pi\sim_{\mathrm{eff}}\hat{a}{t}\land\operatorname{BypassesAuth}(\pi,\hat{a}{t},s_{t}). Here, BypassesAuth \operatorname{BypassesAuth} denotes an effect-equivalent path outside the applicable authorization procedure. Under the audited boundary and threat model, no-bypass applies only to closure-audited actions; isolated channels use post-effect validation, and finalization still requires all fresh mandatory obligations to be satisfied (Appendix B.2). Effect Validation, Commitment, and Recovery. After execution, trusted observers produce evidence e t e_{t} for obligation o i o_{i} , its bound validator returns r t r_{t} , and SpecHarness constructs dependency digest d t d_{t} and provenance record ρ t \rho_{t} . Evidence is admissible iff Adm i ( e t , s t , H exec ) ⇔ \displaystyle\operatorname{Adm}{i}(e{t},s_{t};H_{\mathrm{exec}})\iff{} Trusted i ( e t ; H exec ) \displaystyle\operatorname{Trusted}{i}(e{t};H_{\mathrm{exec}}) (9) ∧ c i ( e t , s t ) \displaystyle}{\displaystyle\land c_{i}(e_{t},s_{t}) ∧ r t ∈ { passed , failed } . \displaystyle}{\displaystyle\land r_{t}\in{\textsc{passed},\textsc{failed}}. The ledger update is L t + 1 [ o i ] = { ( r t , d t , ρ t ) , Adm i ( e t , s t , H exec ) , L t [ o i ] , otherwise . L_{t+1}[o_{i}]=\begin{cases}(r_{t},d_{t},\rho_{t}),&\operatorname{Adm}{i}(e{t},s_{t};H_{\mathrm{exec}}),\ L_{t}[o_{i}],&\text{otherwise}.\end{cases} (10) Trusted i \operatorname{Trusted}{i} checks the bound provider, channel, validator, scope, and version metadata, which ρ t \rho{t} records with the commit-event identifier. The ledger update and commit event are atomic. Both passed and failed may be committed; validator errors, effects without admissible evidence, and agent claims cannot create authoritative state. A committed result is fresh iff its dependency digest matches current artifact, input, validator, environment, and dependent-obligation versions; a known mismatch yields stale , and incomplete version evidence yields unknown . Satisfaction is Sat t ( o i ) ⇔ \displaystyle\operatorname{Sat}{t}(o{i})\iff{} o i ∈ dom ( L t ) ∧ h i ( L t , s t ) \displaystyle o_{i}\in\operatorname{dom}(L_{t})\land h_{i}(L_{t},s_{t}) (11) ∧ Fresh t ( o i ) = true . \displaystyle}{\displaystyle\land\operatorname{Fresh}{t}(o{i})=\textsc{true}. Let 𝒪 x act , mand ( s t ) \mathcal{O}^{\mathrm{act,mand}}{x}(s{t}) be the mandatory subset of 𝒪 x act ( s t ) \mathcal{O}^{\mathrm{act}}{x}(s{t}) . Finalization requires FinalizeAllowed t ( x ) ⇒ ∀ o i ∈ 𝒪 x act , mand ( s t ) , Sat t ( o i ) . \operatorname{FinalizeAllowed}{t}(x)\Rightarrow\forall o{i}\in\mathcal{O}^{\mathrm{act,mand}}{x}(s{t}),;\operatorname{Sat}{t}(o{i}). (12) Mutations recompute freshness for affected entries and dependents, marking them stale or unknown until revalidation. SpecHarness returns source-linked feedback, the agent may propose repairs, and only new admissible evidence can recommit state. The guarantee covers grounded mandatory obligations, not the full natural-language specification. Raw Agent SpecHarness Change vs. Raw (pp) Model Pass ↑ \uparrow U–E ↓ \downarrow S–A ↓ \downarrow Pass ↑ \uparrow U–E ↓ \downarrow S–A ↓ \downarrow Pass U–E S–A GPT-5.6 Sol 71.3 13.6 28.7 85.1 6.3 6.9 +13.8 -7.3 -21.8 Claude Fable 5 69.0 14.5 31.0 81.6 7.1 10.3 +12.6 -7.4 -20.7 Gemini 3.1 Pro 62.1 17.1 37.9 79.3 8.4 12.6 +17.2 -8.7 -25.3 Kimi K3 60.9 17.5 32.2 67.8 10.8 14.9 +6.9 -6.7 -17.3 GLM-5.2 57.5 18.7 29.9 69.0 9.6 11.5 +11.5 -9.1 -18.4 Qwen3.7-Max 54.0 20.0 36.8 65.5 11.2 17.2 +11.5 -8.8 -19.6 DeepSeek-V4-Pro 52.9 20.4 33.3 63.2 12.0 16.1 +10.3 -8.4 -17.2 Macro Average 61.1 17.4 32.8 73.1 9.3 12.8 +12.0 -8.1 -20.0 Table 1: Cross-model results on all 87 SkillsBench tasks. Pass is the official-verifier pass rate; U–E and S–A are the understanding–execution and state–authority gaps. Change columns report percentage-point differences from Raw. Macro Average is the unweighted mean across seven task-agent models. Experiments Benchmarks. We evaluate SkillsBench ( Li et al. 2026b ) and GuideBench ( Diao et al. 2025 ) under different authoritative-state semantics. SkillsBench contains 87 tool-use and artifact-production tasks; agent-visible prompts, workspaces, and injected skills provide runtime inputs, while held-out verifiers determine success. GuideBench contains 1,042 guideline-constrained decision tasks and tests generalization from execution and artifact state to decision state. Across both benchmarks, held-out verifiers, references, and oracles are used only for evaluation, never for obligation construction or runtime feedback. Experimental Setup and Baselines. The seven models (GPT-5.6, Claude Fable 5, Gemini 3.1, Kimi K3, GLM-5.2 ( Zeng et al. 2026 ) , Qwen3.7 ( Qwen Team 2026 ) , and DeepSeek-V4 ( Xu et al. 2026a ) ) in Table 1 also serve as candidate compilers. Using fixed development annotations and a prespecified protocol, we select GPT-5.6 Sol and freeze it before evaluation. Its output Q x ⋆ Q_{x}^{\star} defines direction index 𝒟 x = Directions ( Q x ⋆ ) \mathcal{D}{x}=\operatorname{Directions}(Q{x}^{\star}) and runtime IR Γ x ⋆ \Gamma_{x}^{\star} , yielding 509 directions across 87 SkillsBench tasks. An evaluation-only audit aligns them with 573 of 585 held-out official test functions, including 406 fine-grained matches; this measures alignment, not recovery of verifier semantics. Full compiler-selection results appear in Appendix A.1; extraction, coverage, and alignment results appear in Appendix A.2. All agent runs use OpenHands ( Wang et al. 2025b ) as the shared execution substrate for both Raw and SpecHarness; configuration details appear in Appendix C.1. On SkillsBench, the same seven models serve separately as task agents. The frozen 𝒟 x \mathcal{D}{x} provides a common source-grounded measurement surface and denominator across agents and conditions, while execution evidence determines S m , c ( x ) ⊆ 𝒟 x S{m,c}(x)\subseteq\mathcal{D}{x} . Neither set uses official verifiers or oracles. We compare Raw and SpecHarness across all agents and, with GPT-5.6 Sol fixed, compare Agentic Rubrics ( Raghavendra et al. 2026 ) , VeriMAP ( Xu et al. 2026b ) , AgentSpec ( Wang et al. 2025a ) , and SpecHarness as post-hoc verification, completion gating, runtime enforcement, and mediate-and-commit methods. On GuideBench, removing four duplicate rules from 301 guideline entries yields 297 obligation templates and 5,817 task-level instances across 1,042 tasks. Without an independent extraction oracle, they define a common measurement surface but not complete semantic recovery. We compare Raw and SpecHarness across the same seven agents and, with GPT-5.6 Sol fixed, compare adapted RvLLM ( Zhang et al. 2026 ) , adapted VeriMAP ( Xu et al. 2026b ) , SatLM ( Ye et al. 2023 ) , and SpecHarness as post-hoc verification, completion gating, computation substrate, and evidence-authorized commitment methods. Paired conditions share inputs, budgets, timeouts, and tool access. Frozen obligations and validators evaluate all conditions read-only; only SpecHarness uses their evidence to commit authoritative state and authorize finalization. Appendix C.2 audits baseline implementations, runtime access, budgets, and completion rules. Metrics. For benchmark b ∈ { SB , GB } b\in{\mathrm{SB},\mathrm{GB}} , let 𝒵 x b \mathcal{Z}^{b}{x} denote the common frozen measurement surface for instance x x , and let S m , c b ( x ) ⊆ 𝒵 x b S^{b}{m,c}(x)\subseteq\mathcal{Z}^{b}{x} contain the units whose satisfaction is supported by the shared validators. Let N m , c b N^{b}{m,c} be the valid runs, P m , c b P^{b}{m,c} the officially passing runs, and A m , c b A^{b}{m,c} the runs accepted by the condition-specific runtime. We define G m , c UE , b \displaystyle G^{\mathrm{UE},b}{m,c} = ∑ x | 𝒵 x b ∖ S m , c b ( x ) | ∑ x | 𝒵 x b | , \displaystyle=\frac{\sum_{x}|\mathcal{Z}^{b}{x}\setminus S^{b}{m,c}(x)|}{\sum_{x}|\mathcal{Z}^{b}{x}|}, (13) G m , c SA , b \displaystyle G^{\mathrm{SA},b}{m,c} = | A m , c b ∖ P m , c b | N m , c b . \displaystyle=\frac{|A^{b}{m,c}\setminus P^{b}{m,c}|}{N^{b}{m,c}}. For SkillsBench, 𝒵 x SB = 𝒟 x \mathcal{Z}^{\mathrm{SB}}{x}=\mathcal{D}{x} is the frozen task-direction index; for GuideBench, 𝒵 x GB = 𝒪 act GB ( x ) \mathcal{Z}^{\mathrm{GB}}{x}=\mathcal{O}^{\mathrm{GB}}_{\mathrm{act}}(x) is t The ai agents story also surfaces in NVIDIA DeepStream 9.1 Adds Agentic Skills..., adding another angle. as detailed in the full paper on Arxiv The same ai evaluation question is explored in The Delegation Blind Spot, which adds a research perspective.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!