A relevant paper is not always an inspiring one. Two papers can use similar vocabulary without sharing the idea that unlocks a research problem. ScholarCatalyst turns that distinction into a retrieval benchmark: systems must find earlier papers that researchers say did, or could have, advanced their projects.
The benchmark draws on 184 researchers and 207 computer-science projects. Authors validate early-stage research questions, judge candidate papers and explain their choices. These judgments include useful papers they had not encountered during the original project, so the task is broader than recovering a bibliography.
Reconstructing the question before the solution
Each query describes the landscape at a project’s outset: what was known, what remained unresolved and what the researchers wanted to answer. It must not disclose the eventual solution. The dataset contains a broad core question for each project and additional questions approaching that problem from particular subfields.
Retrieval is restricted to papers published before the source project. External web access is disabled in the main setting to prevent systems from retrieving subsequent accounts of the work. This makes the evaluation closer to searching before the project’s outcome was available, rather than looking up an answer afterward.
The pipeline drafts annotation materials from a paper and its references, then adds uncited candidates found through retrieval. Authors can validate, revise or rewrite those materials. They label both helpful papers and topically related papers that they reject, supplying rationales rather than letting a model’s relevance judgment become the final label.
Measuring inspiration rather than topic overlap
The resulting benchmark contains 894 queries and a corpus of roughly 191,000 papers. Recall@20 measures the fraction of author-credited papers appearing in the first 20 results. The study also uses a ranking-sensitive metric, nDCG, to reward putting those papers higher.
The hard negatives matter. They are papers with topical relevance that the authors did not judge useful for advancing the question. In the reported analysis, rejected papers are at least as similar to queries as the positive papers under the tested lexical and embedding measures. Matching a subject is therefore insufficient to identify the needed idea.
Authors describe several kinds of useful connection: adapting a method, drawing on empirical evidence, generalizing an idea or responding to a limitation. The study also finds that many useful subfield-specific papers were not cited in the finished work. Citation links alone do not reconstruct the full intellectual path.
Why more search did not solve the benchmark
Among the main pre-cutoff systems, general-purpose dense retrieval performs best. The strongest reported results recover 39% of positive papers for core questions and 51% for subfield-specific questions in the top 20. The narrower questions are easier in this setting.
The tested search agents do not improve over embedding retrieval. The authors connect this to candidate coverage: an agent can reason about only the papers that its searches expose. Rephrasing a question repeatedly does not necessarily surface a useful paper whose connection is not obvious from its title or abstract.
These results concern a temporally filtered computer-science corpus and author-judged inspiration, not every kind of scholarly search. The annotation process also begins with machine-generated materials and a selected candidate pool, even though authors control the final judgments. ScholarCatalyst’s practical value is its demanding evaluation target: can a system find an idea worth adapting, rather than merely collect plausible related reading?
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!