A hazard-analysis agent must choose a method whose assumptions fit the event and whose required inputs can be supplied. A plausible workflow can fail either test. Wangshu Zhu and colleagues make those checks explicit in HazardWeaver, an agent architecture for state-dependent scientific route selection.
The HazardWeaver study separates scientific applicability from execution capability. It then revisits route choices as new evidence, derived artifacts or failures change what the agent can do. Scientific models and solvers generate the numerical outputs; the language model chooses routes and manages the workflow.
Evidence determines whether a method applies
The Hazard Knowledge Compiler extracts conditions from scientific literature and model documentation. Each condition records its scope, supporting source and role, such as a required observation, validity range or known failure condition.
At task time, the system marks required conditions satisfied, violated or unresolved. A route is applicable only when all required conditions are satisfied. Unresolved evidence keeps a route outside the eligible set without discarding it, allowing the agent to seek missing information and reassess.
That distinction is useful for workflows in which choosing a model name is easier than checking its assumptions. The architecture records why a route is excluded so that recovery can target the missing evidence rather than repeat the same invalid plan.
Compatible tools are a separate requirement
The Hazard Capability Graph represents executable models, datasets and transformations with typed inputs and outputs. Direct connections require registered compatibility; information-changing operations such as aggregation or downscaling remain explicit workflow steps.
A scientifically appropriate route may still be unavailable because an input is missing or an output cannot feed the next model. The agent searches for compatible alternatives and recomputes eligibility as execution proceeds. The paper warns that failure to find a route within its budget is not a formal proof that no feasible route exists.
Keeping provenance attached to derived artifacts also allows evaluation to distinguish an output generated from permissible evidence from one using information unavailable at the task's cutoff.
Grading the decision as well as the artifact
The Hazard Weaver Benchmark comprises 141 sealed instances across seven single-hazard domains and four multi-hazard interaction classes. Its inventory includes 90 single-hazard and 51 multi-hazard tasks; 96 instances admit multiple routes.
The primary metric, Decision-Constrained Accuracy, combines output correctness with route and execution validity. The benchmark can credit different registered valid routes instead of requiring the solver to reproduce one reference workflow. Abstention receives credit only when permitted recovery cannot reach another benchmark-valid route; a timeout alone is insufficient.
The authors report that HazardWeaver outperforms adapted agent baselines, with the largest gains on multiple-route tasks. The captured source supports that reported finding and the detailed architecture, but does not provide a complete readable result table for Franklin to turn into a numerical leaderboard.
This is a research benchmark of scientific decision workflows. It does not establish operational disaster-warning reliability or authorize unsupervised decisions during emergencies. Its useful contribution is a more demanding evaluation contract: an analysis needs both a supported result and a defensible path to that result.
Comments