Two agents can understand a shared goal and still duplicate work or block each other's path. OverForge addresses cooperative planning by separating persistent coordination strategies from the immediate actions that carry them out. The study uses a frozen language model and a hierarchical controller, rather than updating model weights.
Keeping roles separate from tactical moves
A strategy describes roles or a division of labor, without specifying primitive actions. A tactic chooses what an individual agent should do in the current state. The feasible strategies depend on the environment: shared corridors may favor turn-taking, while separated resources may support specialized roles and hand-offs.
Each agent maintains its own observation-based world representation and a model of its partner's likely state, intentions and needs. It sees the partner's behavior and messages, not the partner's internal representation.
The controller proposes strategy–first-action branches. It compares immediate actions using criteria including goal progress and social fit, then imagines how each branch could develop. Separate cloned world representations prevent an imagined future from changing the real state used for action selection.
Spending inference when choices remain ambiguous
The paper calls its controller a Prefrontal Cortex Module, a functional analogy rather than a claim that it reproduces brain anatomy. The same frozen model acts as proposer, judge and generative forward model.
The controller combines an action's immediate assessment with an assessment of its imagined trajectory. It can deepen deliberation while candidate branches remain ambiguous, up to a defined limit. Its confidence reflects the distribution across branches rather than a single high score.
This is related to the gap between declared plans and executed actions: a stated coordination role matters only if subsequent decisions follow it. The declaration-gap study routes tasks to pattern-specific executors, while OverForge retains the strategy inside each action branch and tests role persistence during cooperation. Neither approach makes choosing an appropriate plan automatic.
What the kitchen experiments support
Experiments use Overcooked-v2 with layouts that impose different coordination constraints. Every condition runs five 300-step episodes with randomized recipes. Language-model agents share a frozen Qwen-3.5-27B backbone and the same dialogue allowance, helping isolate the controller from changes in model or communication budget.
The authors report seven soup deliveries in a connected kitchen for OverForge, compared with three for each flat language-model baseline. They also describe role retention, adoption of unfamiliar partners' proposals, and probes that separate strategic reasoning from tactical adaptation.
These findings remain tied to the tested layouts and partners. The paper qualifies its improvements by whether a layout affords the proposed coordination strategy. A role assignment that works in one kitchen may be infeasible in another.
Imagined outcomes and numerical judgments also depend on the language model. The researchers use fixed criteria and comparison contexts to reduce scoring arbitrariness, but acknowledge residual bias. Cross-episode memory supports adaptation in their experiments; that does not establish lifelong competence across unrestricted tasks or reliable cooperation with every unfamiliar partner.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!