Back to AI Research

AI Research

SquidAgent estimates when parallel agents save more time than coordination costs

Key Takeaways

  • The proposed scheduler uses output-token budgets, shared context and upfront conventions to choose between serial and parallel work.
  • Sending independent tasks to several agents can shorten a job, but workers may duplicate planning or produce outputs that disagree.
  • [SquidAgent](https://arxiv.org/abs/2610.08647) makes those coordination costs part of the decision to parallelize.
  • The authors report an average 2.6-times wall-time speedup and a 2.2-times throughput improvement over Claude Code across nine evaluation tasks.
  • Those are results for the evaluated task suite and execution setup, not a guarantee that adding agents will accelerate an arbitrary project.

Sending independent tasks to several agents can shorten a job, but workers may duplicate planning or produce outputs that disagree. SquidAgent makes those coordination costs part of the decision to parallelize.
The authors report an average 2.6-times wall-time speedup and a 2.2-times throughput improvement over Claude Code across nine evaluation tasks. Those are results for the evaluated task suite and execution setup, not a guarantee that adding agents will accelerate an arbitrary project.

Account for the slowest worker and the shared work

SquidAgent decomposes a request into a directed acyclic graph of subtasks. Dependencies divide the graph into layers. Tasks within a layer can run together, but the layer finishes only after its slowest worker.
The proposed comparison adds two overheads to that critical path. Re-exploration covers workers reconstructing decisions and context the orchestrator already established. Alignment covers reconciling conflicting names, assumptions or formats in their outputs.
Serial execution sums the subtask costs and avoids these particular coordination overheads. Parallel execution is useful only when its critical path plus overhead is cheaper. An independent task is therefore a candidate for parallel execution, not an automatic instruction to launch another worker.

Estimate output length rather than execution duration

The authors argue that models are poor at estimating their own wall-clock runtime. Backend load, tool latency and retries can alter duration, while a model may answer in terms of human engineering time.
SquidAgent instead predicts output-token budgets during the same planning step that creates the task graph. For a fixed model and decoding configuration, output length supplies a proxy for relative generation time.
This approximation has a stated boundary: it fits generation-dominated work and does not directly account for retrieval, tool execution or external API latency. A short response following a slow tool call can still take longer than a large generated artifact. The scheduler's cost estimate should therefore be read alongside the workload it measures.

Reduce overhead before choosing parallel execution

Workers inherit the orchestrator's session state, including its earlier decisions. Before a parallel layer starts, the orchestrator supplies shared conventions for interfaces, naming and formatting. The scheduler then applies a deterministic rule to the estimated costs.
A safety margin requires the predicted advantage to exceed a threshold above one. That margin can reject modest gains, but it reduces the chance that estimation noise makes parallel execution slower than the serial alternative.
The paper evaluates code generation, document authoring and structured planning against seven baselines. Its reported averages support a specific orchestration approach: preserve planning context, specify conventions before delegation and make concurrency conditional on estimated savings. Reproducing those gains requires comparing completion quality and total runtime under the same model and resource conditions. Throughput alone would miss outputs that take longer to reconcile or fail to satisfy the requested deliverable.

Comments