HEXIS: Compiling Skills into Extended Finite State Machines
The HEXIS research addresses a common failure point in AI agents: the difficulty of reliably following complex instructions. While agents are often provided with "skills"—documents containing domain knowledge and procedural rules—they frequently struggle to infer the correct sequence of actions, leading to skipped steps or improper execution. HEXIS solves this by compiling these skill documents into extended finite state machines (FSMs). By separating the agent's task-specific reasoning from the rigid control flow of the process, HEXIS ensures that the agent follows the required steps and dependencies, significantly improving performance and efficiency. The ai agents story also surfaces in OpenAI Unveils GPT-Red an Automated Model..., adding another angle.
Separating Knowledge from Control
In traditional agent setups, the model must simultaneously manage task knowledge and decide the next step in the workflow. This coupling creates a high risk of error, as the model might forget a requirement or misinterpret the order of operations. HEXIS changes this by offloading the control flow to an FSM. The machine tracks the agent's progress, records intermediate results, and uses explicit transition conditions to determine what happens next. Meanwhile, the agent’s language model is still used for reasoning, but it operates within specific states, guided by local instructions that provide only the knowledge necessary for that particular stage of the task.
Incremental Compilation
The HEXIS compiler builds these machines through a two-step process. First, it initializes a machine by mapping the skill document’s clauses and tool interfaces into states, variables, and transition rules. Second, it uses an incremental update process to refine the machine based on actual execution traces. When a new trace is processed, the compiler identifies missing operations or dependencies and adjusts the machine’s structure. Crucially, any proposed update must pass static checks and successfully "replay" all previously accepted execution traces before it is committed, ensuring that the machine remains reliable as it evolves. The same ai evaluation question is explored in The Delegation Blind Spot, which adds a research perspective.
Performance and Efficiency
Experiments across four benchmarks—including spreadsheet manipulation, mathematical reasoning, data analysis, and long-context question answering—demonstrate that HEXIS significantly outperforms standard methods like Skill + ReAct. On average, HEXIS improved success rates by 16.1 percentage points. Beyond accuracy, the approach is highly efficient; by removing the need for the model to repeatedly infer the next step from long, complex skill documents, HEXIS reduced execution token usage by 38.4% to 88.9% for the Qwen3.8-27B model. These results suggest that explicitly encoding procedural requirements into a state machine is a more robust way to manage agent workflows than relying on context-based inference alone. The ai agents story also surfaces in OpenAI Says AI Found Possible Navier–Stokes..., adding another angle. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!