Back to AI Research

AI Research

HyperBrowseComp tests difficult web searches across 13 languages and mixed media

Key Takeaways

  • The 423-question benchmark exposes retrieval-harness sensitivity and difficult evidence chains, rather than a universal model ranking.
  • Finding a fact on the web can require more than a good text query.
  • A clue in a video might identify a street, which identifies a bank, whose old report contains the answer.
  • HyperBrowseComp builds questions around that kind of evidence chain.
  • Alham Fikri Aji and colleagues introduce the [HyperBrowseComp benchmark](https://arxiv.org/abs/2610.03574), comprising 423 manually authored, human-validated questions across 13 languages.

Finding a fact on the web can require more than a good text query. A clue in a video might identify a street, which identifies a bank, whose old report contains the answer. HyperBrowseComp builds questions around that kind of evidence chain.
Alham Fikri Aji and colleagues introduce the HyperBrowseComp benchmark, comprising 423 manually authored, human-validated questions across 13 languages. Native or highly proficient speakers write questions in their own language contexts instead of translating an English seed dataset.

Concise answers can require long searches

Each question targets a short, publicly verifiable answer. Authors submit source URLs, exact evidence locations, acceptable aliases and a solution trace; another annotator checks correctness, uniqueness and access. A valid answer can require finding a video timestamp, inspecting a scanned report or connecting map evidence with text.
The benchmark excludes sources requiring payment or authentication and restricts sensitive personal information. Approximately 64.3% of questions receive at least one non-text modality classification, with classifications checked by humans. Text-only questions remain when they meet the difficulty criteria, so multimodality is not required for every item.
The team also runs difficulty audits without internet access. It removes 77 candidates that meet a combined answerability and specificity rule, leaving the 423 retained questions. That reduces some opportunities to answer from stored knowledge, while it cannot prevent future deliberate training on the released benchmark.

Search integration changes the result

Five models are tested with provider-native search, while selected configurations also use an Exa retrieval interface or the OWL multi-agent environment. Native-search Gemini 3.7 Flash reaches the highest reported full-set accuracy, 31.68%, in these runs.
Changing the retrieval setup changes outcomes. GPT-5.6 Sol rises from 19.15% with built-in search to 26.71% with Exa, while the evaluated Gemini variants decline under Exa. These are model-and-harness combinations, not isolated measurements of a language model's ability.
OWL has 93 terminal runtime failures among 423 attempts, which count as incorrect. Removing those failures would change the meaning of the score by excluding a part of the system's observed behavior. The paper reports the failure accounting as well as resource use, showing why agent evaluation should include whether the workflow completes.

A stress test, not a typical browsing workload

Human evaluation covers 30 questions in Indonesian, Thai and Vietnamese. Human participants answer 15 correctly and give up on four; the best full-set model answers 13 correctly on that same subset. This small sample provides context about effort and differing failure cases rather than evidence that either system is generally equivalent to human researchers.
The questions deliberately emphasize obscure evidence and difficult clue chains. A 31.68% score on this set should not be described as a browser failing that proportion of ordinary searches. Open-web access and source changes also differ from evaluation against a fixed corpus.
The benchmark gives developers a way to test language coverage, media inspection and persistent search together. It also makes a practical measurement point: retrieval tools, runtime limits and failure handling belong in the performance claim alongside the model name.

Comments