A safety verifier can issue an internally valid certificate for a harmful action if the causal graph it trusts is wrong. In Who Verifies the Graph?, Fabio Rovai tests that failure mode in CIVeX, the author's own earlier causal action verifier. The study changes the graph supplied to the verifier while keeping its other components fixed.
Attack the assumption behind the certificate
CIVeX checks whether a proposed state-changing action has an identifiable causal effect under a committed graph. It attaches an identification argument and a lower confidence bound to its decision. The new study asks what happens when that committed graph contains an error, rather than honestly declaring an uncertainty.
The synthetic benchmark contains 1,050 actions per experimental cell, spanning six workflow families and seven seeds. At the published confounding strength, removing one edge that represents hidden confounding produced a 15.3 percent false-execution rate. That rate counts harmful executions as a share of all proposed actions; separately, 91 percent of the actions actually executed were harmful.
A different attack reversed a mediator's arrow direction. In that constructed regime, false executions reached 48.9 percent and correct executions fell to zero. These are results for one verifier and synthetic benchmark, not measured attack rates against deployed agents. The study grants access to the graph and does not establish how difficult that access would be to obtain.
Check decisions with independent experiments
The proposed defense compares an observational certificate with a bounded randomized sample. If their intervals disagree, the action can be refused or reconsidered. Refusing actions that fail the check, or cannot be checked, removed observed false executions in the tested settings.
The response to disagreement matters. A fallback that reuses the wrongly committed adjustment set can inherit the same error, even after collecting randomized data. The paper reports that a graph-free fallback addresses the harmful-execution problem in the reversed-arrow setting.
The test is a discrepancy diagnostic, not a generally calibrated guarantee. The author reports two false alarms among 555 executions when the tested graph was truthful. Experiment availability also limits the defense: allowing unchecked actions through reintroduces harmful executions, while refusing them reduces useful activity.
Audit wrongful inaction as well as action
Checking only actions that reach execution misses beneficial actions rejected earlier. Under the omitted-edge attack, the paper reports that 97.1 percent of beneficial actions remained unexecuted in an execution-attestation setting. A clean harmful-execution record can therefore coexist with a system that blocks valuable work.
Auditing rejected decisions recovers some of that value, but requires more experiments. In the reported accounting, recovering safety cost 127 experiments per 1,050 actions; recovering the lost value required another 614.
The practical research question is which assumptions a certificate depends on and which decisions an audit never sees. The study does not validate the defense for irreversible real-world actions, where obtaining a safe randomized sample may be impossible. It motivates reporting both harmful actions allowed and useful actions blocked, with the audit budget made explicit.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!