Back to AI Research

AI Research

HEAR connects agent workflow intent with inference-engine runtime state

Key Takeaways

  • A bidirectional serving protocol lets harnesses and engines coordinate cache use and execution settings, with workload-specific speed and latency trade-offs.
  • An agent harness knows which requests depend on one another, while an inference engine knows which prefixes remain in memory and how busy its queues are.
  • The [HEAR paper](https://arxiv.org/abs/2610.06597) proposes a communication contract between those layers so a scheduling policy can use both kinds of information.
  • It targets how existing workflow requests are served, while preserving task-selection logic and model semantics.
  • ## Descriptions and controls have different force

An agent harness knows which requests depend on one another, while an inference engine knows which prefixes remain in memory and how busy its queues are. The HEAR paper proposes a communication contract between those layers so a scheduling policy can use both kinds of information. It targets how existing workflow requests are served, while preserving task-selection logic and model semantics.

Descriptions and controls have different force

HEAR organizes messages into four categories. A harness describes workflow intent and sends execution requirements or controls. The engine reports runtime state and capabilities, then returns outcomes such as completion, rejection or failure. Messages identify the associated request, context version or serving instance.
The protocol separates several distinctions that matter for implementations. A prediction that a context will be reused is different from a request to retain its cache. A scheduling preference is best-effort, while a mandatory requirement constrains valid execution. Observing a cached prefix does not reserve it. Acceptance of a preparation request does not mean its work has finished.
Those distinctions give optimization policies a shared vocabulary without prescribing one universal scheduler. An implementation can expose only the controls it supports, allowing a harness to discover capabilities rather than assume an unavailable operation will succeed.

Cache reuse can compete with waiting time

The authors test cache-aware coordination on SCBench and a production-derived Mooncake trace. On SCBench, a cache-aware scheduler increases reuse and reduces median time to first token, but repeated deferral of cold requests worsens the maximum wait. A waiting-protection rule limits that behavior.
With a forty-second guard, the reported SCBench batch speedup is 1.61x relative to first-come-first-served scheduling. Median time to first token falls from 63.1 to 28.3 seconds, a 2.23x reduction; the ninety-fifth-percentile wait also improves. A sixty-second guard achieves higher reuse and faster batch completion while providing weaker tail-latency protection.
Mooncake results vary with load. Retaining contexts with likely future use helps at lower load, while request ordering and retention offer complementary benefits at intermediate load. No tested policy dominates across all load levels. The protocol permits exchanging the relevant information, but the policy still determines how to use it.

Configuring execution for the workload

HEAR also supports role-specific inference configurations. In the research-agent experiments, workload-specific settings yield reported end-to-end speedups of 1.23x on BrowseComp-Plus and 2.45x on DeepResearchBench, without observed task-quality degradation in those evaluations. The authors keep models, prompts, tool budgets and hardware matched within each setting and include failures and retries in the results.
Those numbers describe the tested implementations and workloads. They do not imply that adding HEAR to an arbitrary agent automatically reproduces the same gains, or that an unchanged benchmark score proves quality is unaffected outside the measured tasks.
The proposal gives systems engineers a way to make context lifecycles and live serving conditions explicit. Its practical benefit depends on a compatible engine, accurate runtime feedback and policies that respect dependencies while choosing between cache efficiency and user waiting time.

Comments