Back to AI Research

AI Research

Learning to Ideate for Scientific Impact | AI Research

Key Takeaways

  • What the paper is about Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated...
  • Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility.
  • This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected \emph{impact}.
  • We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake.
  • We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label.
Paper AbstractExpand

Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility. This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected \emph{impact}. We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake. We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label. We then train a goal-conditioned reward model to predict citation-impact labels from research goal and idea pairs, and use this reward to align an idea generator through supervised fine-tuning followed by reinforcement learning. To reduce circularity, we evaluate generated ideas with a held-out, reference-grounded protocol that compares model outputs against historical ideas under the same research goal and weights judgments by the reference idea's citation-impact label. Experiments show that our RL-tuned model consistently produces ideas with higher estimated impact than both the base model and supervised fine-tuning baselines. Our findings position scientific impact as a practical, outcome-grounded feedback signal for aligning LLMs in open-ended scientific discovery.

What the paper is about

Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility. This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected \emph{impact}. We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake. We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label. We then train a goal-conditioned reward model to predict citation-impact labels from research goal and idea pairs, and use this reward to align an idea generator through supervised fine-tuning followed by reinforcement learning. To reduce circularity, we evaluate generated ideas with a held-out, reference-grounded protocol that compares model outputs against historical ideas under the same research goal and weights judgments by the reference idea's citation-impact label. Experiments show that our RL-tuned model consistently produces ideas with higher estimated impact than both the base model and supervised fine-tuning baselines. Our findings position scientific impact as a practical, outcome-grounded feedback signal for aligning LLMs in open-ended scientific discovery. The ai search story also surfaces in Stanford AI discovery identifies natural weight..., adding another angle.

What it covers

Learning to Ideate for Scientific Impact Shubham Kale Affiliation: TCS Research, India Correspondence to: [email protected] Aniketh Garikaparthi Affiliation: TCS Research, India Manasi Patwardhan Affiliation: TCS Research, India Abstract Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility. This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected impact . We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake. We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label. We then train a goal-conditioned reward model to predict citation-impact labels from research goal and idea pairs, and use this reward to align an idea generator through supervised fine-tuning followed by reinforcement learning. To reduce circularity, we evaluate generated ideas with a held-out, reference-grounded protocol that compares model outputs against historical ideas under the same research goal and weights judgments by the reference idea’s citation-impact label. Experiments show that our RL-tuned model consistently produces ideas with higher estimated impact than both the base model and supervised fine-tuning baselines. Our findings position scientific impact as a practical, outcome-grounded feedback signal for aligning LLMs in open-ended scientific discovery. Keywords: Impact Alignment, Reinforcement Learning, Scientific Forecasting, Scientific Ideation, Reward Modeling, Large Language Models 1 Introduction The onset of large language models (LLM) use in science has led to an asymmetric rise in the number of publications, without substantial evidence for increase in quality at a comparable rate ( Birhane et al., 2023 ; Messeri and Crockett, 2024 ; Kusumegi et al., 2025 ) . Within this shift, LLMs have expanded their role in scientific discovery from writing assistance to broader support across the research life cycle, including literature understanding, hypothesis generation, experimentation, and end‑to‑end research automation ( Zhang et al., 2024 ; Zhang et al., 2025 ; Lu et al., 2026 ) . Notably, research ideation has emerged as a key driver in adoption where systems such as SciMON, ResearchAgent, SciMuse, DeepInnovator, and IRIS show that LLMs can generate, refine, and evaluate literature-grounded ideas through iterative search, expert evaluation, and human‑in‑the‑loop interaction ( Wang et al., 2024a ; Baek et al., 2025 ; Gu and Krenn, 2025 ; Fan et al., 2026 ; Garikaparthi et al., 2025 ) . However, these systems largely optimize for proxy criteria such as novelty, feasibility or clarity over downstream impact ( Li et al., 2024 ; Guo et al., 2024 ; Qiu et al., 2025 ) . Further, human evaluations on such system outputs highlight a critical gap between perceived quality and practical value ( Si et al., 2025 ) . In scholarly research, many proposed ideas are easy to judge for novelty or clarity but difficult to assess for downstream uptake until years later. Citations are an imperfect signal, but they provide one scalable trace of how strongly a contribution is taken up by later work.These limitations motivate ideation systems that are explicitly conditioned on impact , a delayed real-world signal for aligning models towards practical relevance and scientific quality. While classical work on scientometrics ( Waltman, 2016 ; Hutchins et al., 2016 ; Thelwall et al., 2013 ) probes scientific impact prediction and provides a foundation for reasoning about downstream value, they fail to tie prediction as a mechanism for empowering generation systems. Recent works extend this by forecasting high-impact topics, predicting impact of new papers, and introducing multi-dimensional benchmarks covering awards, media attention, patents, and artifact adoption ( Gu and Krenn, 2024a ; Zhao et al., 2024 ; Lu et al., 2025 ; Zhu et al., 2026 ; Zhang and Wu, 2024 ; Zhu and Shuhuai, 2024 ) . However, these approaches remain largely post hoc, operating on existing entities (e.g., papers or topics). Similarly, work on future-aligned proposals and retrospective evaluation improves impact-driven ranking and judgment ( Wang et al., 2026 ; Jiang, 2026 ; Ajith et al., 2026 ) , but does not optimize idea generation for expected downstream value before costly validation. Thus, current approaches focus on assessment and forecasting, not controllable, impact-aware ideation. More recently, Tong et al. (2026) train a reward model on high- vs. low-citation abstracts and use it to guide idea generation. However, their supervision remains (1) limited to binary preferences, diluting dense citation signals, and (2) conditioned on arbitrary samples rather than multiple ideas for the same research goal, making it sensitive to surface cues (e.g. phrasing of empirical results or topic popularity cues) rather than goal-conditional impact. Moreover, optimizing responses without an explicit reward over (research goal, idea) limits its ability to capture whether an idea is impactful for the conditioning objective. This leaves the challenge of aligning idea generation directly to impact while preserving explicit goal conditioning. Figure 1 : Dataset construction pipeline. Alignment in LLMs has largely focused on generic objectives—helpfulness, harmlessness, honesty, and safety—via RLHF-, RLAIF-, and DPO-style methods ( Wang et al., 2024b ; Tan et al., 2025 ) . We ask whether we can leverage RL via reward-models to align language models for prioritizing impactful science; a notably different challenge aimed at tying LLMs preferences towards practical, real-world guided outcomes. In this work, we make the following contributions:

• We construct a large-scale dataset of research goal–idea pairs from computer science papers including interdisciplinary work, each annotated with an ordinal citation-impact label normalized by publication year.

• We formulate citation-aligned scientific ideation as goal-conditioned generation under a delayed uptake signal, and train a reward model over (research goal, idea) pairs using ordinal citation labels.

• We align a goal-conditioned idea generator using the learned reward model, starting from supervised fine-tuning and further optimizing with reinforcement learning.

• We introduce a held-out, reference-grounded evaluation protocol for measuring whether generated ideas are judged impactful relative to historical ideas under the same research goal.

• Empirically, our reward model reaches 48.0% accuracy on ordinal impact prediction, beating GPT-5 by over 20 points, and our RL-aligned generator improves majority-vote impact rate from 30.66% for the base model and 26.61% for SFT to 51.66%. 2 Related Work Idea Generation and Evaluation. A substantial body of recent work studies LLM-based scientific ideation. SciMON , ResearchAgent , DeepInnovator , and IRIS generate and refine literature-grounded ideas through novelty-oriented search, iterative revision, specialized training, and human-in-the-loop interaction ( Wang et al., 2024a ; Baek et al., 2025 ; Fan et al., 2026 ; Garikaparthi et al., 2025 ) . Complementary work has focused on evaluating idea quality through human judgments, preference modeling, or dedicated benchmarks, including SciMuse , IdeaBench, AI Idea Bench 2025 , Proof of Time , and HindSight ( Gu and Krenn, 2024b ; Guo et al., 2024 ; Qiu et al., 2025 ; Ye et al., 2026 ; Jiang, 2026 ) . Citation and Scientific Impact Prediction. Scientific impact has traditionally been studied through citation-based indicators and their normalized variants, with complementary alternatives such as altmetrics ( Waltman, 2016 ; Hutchins et al., 2016 ; Thelwall et al., 2013 ) . More recent work extends this line by forecasting high-impact topics from evolving knowledge graphs, predicting the impact of newborn papers from textual or citation-related signals, and introducing broader benchmarks that move beyond citations to multiple dimensions of scientific influence ( Gu and Krenn, 2024a ; Zhao et al., 2024 ; Lu et al., 2025 ; Zhu et al., 2026 ; Ajith et al., 2026 ) . While these methods are valuable for forecasting and assessment, they largely operate over already instantiated objects such as topics, papers, or contributions, rather than directly supporting goal-conditioned idea generation. 3 Problem Definition Let 𝒫 = { p i } i = 1 N \mathcal{P}={p_{i}}{i=1}^{N} denote a corpus of computer science and and related interdisciplinary papers. Each paper p i p{i} is associated with full text x i x_{i} and a citation-derived impact signal. From each paper, we extract a structured pair ( g i , d i ) = Φ ⁡ ( x i ) , (g_{i},d_{i})=\Phi(x_{i}), (1) where Φ \Phi is an LLM-based extraction function, g i g_{i} denotes the research goal , and d i d_{i} denotes the corresponding idea description . The extraction is designed to capture the core research objective and proposed idea while excluding empirical results and other post hoc evidence. To represent impact in a form suitable for learning, each paper is assigned a year-normalized citation signal, which is further discretized into an ordinal label y i ∈ { 0 , … , K − 1 } y_{i}\in{0,\dots,K-1} , where the K K ordered categories correspond to increasing levels of scientific impact (e.g., Very Low , Low , Medium , High , and Very High ). This yields a structured dataset 𝒟 = { ( g i , d i , y i ) } i = 1 N . \mathcal{D}={(g_{i},d_{i},y_{i})}{i=1}^{N}. (2) Given this dataset, our goal is to learn a conditional idea generator that, for a given research goal g g , produces an idea description d d that is aligned with higher-impact ordinal labels. Formally, we seek a policy π ⁡ ( d ∣ g ) \pi(d\mid g) that generates ideas maximizing an impact-alignment objective induced by the ordinal label space (Detailed in Section 5 ). We study this problem through a three-stage framework consisting of: (i) a reward model that predicts ordinal impact from a research goal–idea pair, (ii) a supervised generator that produces idea descriptions conditioned on research goals, and (iii) a reinforcement learning stage that further aligns the generator toward higher predicted impact. 4 Dataset Construction Our dataset construction pipeline consists of five steps refer figure 1 . Step 1: Corpus Acquisition and Paper Filtering. We use Semantic Scholar resources—S2AG metadata ( Wade, 2022 ) , the S2ORC-v2 full-text corpus, and the abstract dataset—collectively covering ∼ 120 ​ M \sim 120\mathrm{M} scholarly papers. From this pool, we select open-access papers published between January 2010 and December 2024, we choose this publication window to ensure that each paper had sufficient time to accumulate citation data by the citation collection date (10 March 2026). Papers published very recently often have artificially low citation counts due to limited exposure time rather than low scientific impact, which could incorrectly place them into lower-impact citation bins.We then filter paper that are tagged with ‘Computer Science’ as a field of study, including interdisciplinary works spanning computer science and other domains. This filtering yields 1.1M papers. We then align corpus identifiers across metadata, full-text, and abstract resources to construct a unified dataset containing full texts, abstracts, and metadata such as citation counts, publication dates, and field labels. Step 2: Text Cleaning. For each paper in the filtered corpus, we perform text cleaning to retain only ideation-relevant content by removing references and appendices from the full text. We also exclude review and survey papers, which primarily summarize prior work rather than present specific research goal–idea pairs. To identify these, we first apply a rule-based filter over titles and abstracts using keywords such as ‘review’ and ‘survey’, yielding 38K candidate papers. We then apply an LLM-based verification step to refine this set, ultimately identifying and removing 32K review or survey papers from the corpus, finally resulting into 1M papers in the cleaned corpus. Step 3: Citation Normalization and Ordinal Label Construction. To derive an impact proxy, we use citation counts scraped on March 10, 2026. Since raw citation counts are strongly influenced by publication age, we normalize citations by the age of the paper in years. Concretely, for a paper with raw citation count c c and age a a , we define its citation-normalized impact score as x = c a x=\frac{c}{a} , where x x denotes citations per year. We then converted this continuous signals into ordinal impact bins using fixed thresholds. The thresholds are chosen to balance three considerations: preserving a meaningful ordering over citation-normalized impact, separating qualitatively different citation regimes, and maintaining sufficient support in each class for stable learning. In particular, the lowest bin captures papers with negligible yearly citation uptake, the intermediate bins capture progressively stronger but more common levels of influence, and the highest bin isolates papers with clearly exceptional citation-normalized impact. This discretization provides a robust ordinal supervision signal while reducing sensitivity to noise and heavy-tailed variation in raw citation counts ( Bornmann, 2012 ) . Zero / Very Low : 0 ≤ x ≤ 0.3223 , \displaystyle:;0\leq x\leq 0.3223, (3) Low : 0.3223 < x < 5.0 , \displaystyle:;0.3223<x<5.0, (4) Medium : 5.0 ≤ x < 15.0 , \displaystyle:;5.0\leq x<15.0, (5) High : 15.0 ≤ x < 40.0 , \displaystyle:;15.0\leq x<40.0, (6) Very High : x ≥ 40.0 . \displaystyle:;x\geq 40.0. (7) These ordinal bins form the impact labels used in downstream reward modeling and impact-aware idea generation. Step 4: Stratified Sampling and Data Splitting. From the cleaned corpus, we construct a stratified sample of 100K papers to ensure balanced coverage across both publication years and ordinal impact labels, further forming two subsets, each of 50K papers. retaining the uniform distributions. Each 50K sampled subset is then split into training, validation, and test partitions using a 70:15:15 ratio. One of the subset is used for reward modeling and the other is used for supervised fine-tuning (SFT), and reinforcement learning (RL). This design to have two different subset is to minimize the possibility of reward hacking arising from memorization or distribution overlap between reward modeling and policy optimization stages. Dataset Split Zero/VL Low Med. High V.High Total Reward Train 7,231 7,232 7,241 7,236 6,033 34,973 Reward Val 1,552 1,550 1,549 1,552 1,292 7,495 Reward Test 1,552 1,552 1,548 1,551 1,293 7,496 SFT/RL Train 7,239 7,238 7,236 7,239 6,038 34,990 SFT/RL Val 1,552 1,552 1,552 1,551 1,292 7,499 SFT/RL Test 1,549 1,552 1,551 1,550 1,295 7,497 Table 1 : Paper distribution across ordinal classes for distinct dataset subsets and splits. Counts remain approximately uniform across train, validation, and test sets, indicating successful stratification. Step 5: Research Goal and Idea Extraction. The final stage converts each paper into a structured research goal–idea description pair. To enable this, we first optimize the extraction prompt by designing seven candidate variants and evaluating them on 15 open-access computer science papers curated by human evaluators. For each paper, we generate outputs from all seven prompts, yielding 105 candidate extractions. These are evaluated through both human assessment and an LLM-based judge(prompt provided in Appendix A.2.1 ) using a fixed rubric where the judge compared the extracted research goal and idea description against the full paper text. Based on this combined evaluation, we select the highest-quality prompt (prompt provided in Appendix A.2.2 ) and apply it to the sampled corpus to extract structured research goals and idea descriptions for all papers. The final dataset thus consists of triples of the form ( research goal , idea description , ordinal impact label ). Final dataset statistics are provided in Table 1 . Figure 2 : Impact Aligned Ideation Training formulation 5 Methodology Our framework consists of three stages: reward modeling, supervised idea generation (SFT), and impact-aligned reinforcement learning (RL) refer figure 2 . Reward Modeling. We first learn a reward model that predicts the ordinal impact label from a research goal–idea pair. Formally, let R ​ M θ ​ ( g , d ) → p θ ​ ( y ∣ g , d ) , RM{\theta}(g,d)\rightarrow p_{\theta}(y\mid g,d), (8) where p θ ​ ( y ∣ g , d ) p_{\theta}(y\mid g,d) denotes an ordinal distribution over the K K impact levels, and θ \theta are the reward model parameters. We implement R θ R_{\theta} by attaching an ordinal classification head based on CORN ( Shi et al., 2023 ) to a language model backbone. Given the training set 𝒟 = { ( g i , d i , y i ) } i = 1 N \mathcal{D}={(g_{i},d_{i},y_{i})}{i=1}^{N} , the reward model is trained to minimize the ordinal classification loss ℒ RM ( θ ) = ∑ i = 1 N ℓ CORN ( p θ ( ⋅ ∣ g i , d i ) , y i ) , \mathcal{L}{\mathrm{RM}}(\theta)=\sum_{i=1}^{N}\ell_{\mathrm{CORN}}!\left(p_{\theta}(\cdot\mid g_{i},d_{i}),y_{i}\right), (9) where ℓ CORN \ell_{\mathrm{CORN}} denotes the CORN loss for ordinal classification ( Shi et al., 2023 ) . To use the reward model for policy optimization, we convert the predicted ordinal citation-impact category into a scalar reward. During evaluation, the reward model first predicts an ordinal citation-impact label y ^ = R ​ M θ ​ ( g , d ) , \hat{y}=RM_{\theta}(g,d), (10) where y ^ ∈ { 0 , … , K − 1 } \hat{y}\in{0,\dots,K-1} . The scalar reward is then obtained using a linear reward-scaling function: r ⁡ ( g , d ) = ρ ⁡ ( y ^ ) = y ^ K − 1 ⋅ r max , r(g,d)=\rho(\hat{y})=\frac{\hat{y}}{K-1}\cdot r_{\max}, (11) where K K denotes the number of ordinal citation-impact categories , r max r_{\max} is the maximum reward scaling constant and ρ ⁡ ( ⋅ ) \rho(\cdot) denotes a linear reward-scaling function that maps the predicted ordinal label y ^ \hat{y} to a scalar reward value. In our implementation, we use K = 5 K=5 and r max = 5 r_{\max}=5 , resulting in scalar rewards in the range [ 0 , 5 ] [0,5] . Supervised Idea Generation. We train a conditional generator, which is a pre-trained model π ϕ \pi_{\phi} to produce an idea description from a research goal π ϕ ​ ( d ∣ g ) \pi_{\phi}(d\mid g) , where ϕ \phi are the generator parameters. The supervised fine-tuning (SFT) objective maximizes the conditional likelihood of the reference idea description given the research goal: ℒ SFT ( ϕ ) = − ∑ ( g i , d i ) ∈ 𝒟 log π ϕ ( d i ∣ g i ) . \mathcal{L}{\mathrm{SFT}}(\phi)=-\sum{(g_{i},d_{i})\in\mathcal{D}}\log\pi_{\phi}(d_{i}\mid g_{i}). (12) This stage provides an initialization policy that can generate coherent and goal-conditioned idea descriptions before reinforcement learning, This stage provides an initialization policy that can generate coherent and goal-conditioned idea descriptions before reinforcement learning, thereby mitigating the cold-start problem commonly observed when RL optimization is initialized from weak policies ( Ouyang et al., 2022 ) . Model Acc MAE ± 1 \pm 1 Acc Spear QWK Zero-shot Qwen3-8B 20.7 1.410 58.7 0.065 0.004 GPT-4o 23.9 1.155 68.1 0.354 0.235 GPT-5 27.7 1.027 75.3 0.402 0.338 Fine-tuned RM 48.0 (+20.3) 0.632 (-0.39) 90.6 (+15.3) 0.756 (+0.35) 0.754 (+0.41) Table 2 : Trained Reward Model (RM: Qwen3-8B + LoRA + CORN) Performance comparison with zero-shot models on the ordinal classification task. Improvement of RM over GPT-5. Acc: Accuracy, Spear: Spearman Correlation Model V.Low Low Med. High V.High GPT-4o 1.95 1.41 1.07 0.96 0.38 GPT-5 0.66 0.86 0.94 1.02 1.79 Qwen3-8B 3.01 2.01 1.00 0.00 1.00 RM 0.50 0.71 0.77 0.74 0.43 Table 3 : Class-wise Mean Absolute Error (MAE) for RM and baseline models. Best , Second Best Performance Impact-Aligned Reinforcement Learning. Finally, we initialize a policy model with the SFT parameters and further optimize it using reinforcement learning. Let π ψ \pi_{\psi} denote the policy, initialized from ϕ \phi . For a given research goal g g , the policy samples a group of candidate ideas: d ( 1 ) , d ( 2 ) , … , d ( G ) ∼ π ψ ( ⋅ ∣ g ) , d^{(1)},d^{(2)},\dots,d^{(G)}\sim\pi_{\psi}(\cdot\mid g), (13) where G G is the group size. Each generated candidate is paired with the same research goal and evaluated by the reward model: r ( j ) = r ⁡ ( g , d ( j ) ) r^{(j)}=r(g,d^{(j)}) . Following the Group-Relative Policy Optimization (GRPO) ( Shao et al., 2024 ) setting, rewards within a sampled group are normalized to obtain relative advantages: A ( j ) = ( r ( j ) − μ r ) / σ r A^{(j)}=(r^{(j)}-\mu_{r})/\sigma_{r} , where μ r \mu_{r} and σ r \sigma_{r} are the mean and standard deviation of the rewards within the group. The policy is then updated to increase the likelihood of candidates with higher relative advantage. In abstract form, the optimization objective can be written as max ψ 𝔼 g ∼ 𝒟 , d ( j ) ∼ π ψ ( ⋅ ∣ g ) [ 1 G ∑ j = 1 G A ( j ) log π ψ ( d ( j ) ∣ g ) ] . \max_{\psi};\mathbb{E}{g\sim\mathcal{D},,d^{(j)}\sim\pi{\psi}(\cdot\mid g)}\left[\frac{1}{G}\sum_{j=1}^{G}A^{(j)}\log\pi_{\psi}(d^{(j)}\mid g)\right]. (14) Thus, the policy is encouraged to generate idea descriptions that receive higher reward , i.e., generated ideas are aligned for stronger scientific impact. 6 Experimentation and Results (a) Qwen3-8B (Zero-shot) (b) RM (Qwen3-8B + LoRA + CORN) Figure 3 : Confusion matrices on the ordinal classification benchmark. Darker cells indicate larger sample counts. 6.1 Models and Training Details For reward modeling, we use Qwen3-8B based backbones, which is trained with a CORN ordinal regression head, using LoRA ( Hu et al., 2021 ) with rank r = 64 r=64 , scaling factor α = 128 \alpha=128 , dropout 0.1 0.1 , maximum sequence length 4096, and train for 3 epochs using learning rate 2 × 10 − 5 2\times 10^{-5} . For supervised fine-tuning (SFT), we use Qwen3-8B with LoRA rank r = 64 r=64 , scaling factor α = 128 \alpha=128 , and dropout 0.1 0.1 , trained for 3 epochs with maximum sequence length 2048, learning rate 2 × 10 − 5 2\times 10^{-5} , cosine learning-rate scheduling, and weight decay 0.01 0.01 . For reinforcement learning, we initialize the policy from the SFT model and optimize it using the DAPO ( Yu et al., 2025 ) variant of GRPO with learning rate 1 × 10 − 6 1\times 10^{-6} , group size 4, maximum completion length 256, and KL regularization coefficient β = 0.01 \beta=0.01 . Generation during RL training uses nucleus sampling with top- p = 0.95 p=0.95 and temperature 0.9 0.9 . Across all stages, we employ LoRA adaptation over attention and feed-forward projection layers, together with gradient checkpointing and mixed-precision BF16/FP16 training. Detailed RL optimization hyperparameters are provided in Table 10 . 6.2 Metric Metrics for evaluation of the Reward Model. We evaluate the reward model using multiple ordinal classification metrics. Accuracy (Acc) measures exact category prediction accuracy, while Mean Absolute Error (MAE) measures the average ordinal distance between predicted and ground-truth citation-impact labels. ± 1 \pm 1 Accuracy considers a prediction correct if it falls within one adjacent ordinal category of the ground-truth label. Spearman Correlation (Spear) evaluates rank correlation between predicted and true citation-impact levels, and Quadratic Weighted Kappa (QWK) measures ordinal agreement while assigning larger penalties to predictions that are farther from the correct citation-impact category. Metrics for the evaluation of the Idea Generator. We propose a novel metric to evaluate the impactfulness of generated research ideas: Reference-Grounded Citation-Weighted Idea Impact Evaluation ( RGCW-IIE ). The evaluation is performed by comparing each generated idea against a ground-truth reference idea under the same research goal. Unlike similarity-based metrics, the proposed evaluation does not require the generated idea to exactly match the reference idea. Instead, it measures whether the generated idea is judged to be scientifically impactful relative to the ground-truth idea, and weights this judgment using the citation-normalized impact label of the reference idea. Let 𝒟 = { ( g i , d i , y i ) } i = 1 N \mathcal{D}={(g_{i},d_{i},y_{i})}{i=1}^{N} denote the evaluation dataset, where g i g{i} is the research goal, d i d_{i} is the ground-truth reference idea description, and y i ∈ { 0 , 1 , … , K − 1 } y_{i}\in{0,1,\ldots,K-1} is the ordinal year-normalized citation label assigned to the ground-truth idea. The label y i y_{i} represents the citation impact of the reference idea after normalizing for publication year. Let ℳ ∈ { ℳ base , ℳ SFT , ℳ RL } \mathcal{M}\in{\mathcal{M}{\mathrm{base}},\mathcal{M}{\mathrm{SFT}},\mathcal{M}{\mathrm{RL}}} denote a candidate idea-generation model, corresponding to the baseline, supervised fine-tuned, and reinforcement-learned models, respectively. For each research goal g i g{i} , the model generates an idea (prompt shown in Appendix A.2.3 ) d ^ i ℳ = ℳ ⁡ ( g i ) \hat{d}{i}^{\mathcal{M}}=\mathcal{M}(g{i}) . Given the research goal g i g_{i} , the ground-truth idea d i d_{i} , and the generated idea d ^ i ℳ \hat{d}{i}^{\mathcal{M}} , we use a strong evaluator model ℰ \mathcal{E} , i.e GPT-5.1, GPT-4.1, GPT-4o, to judge whether the generated idea is impactful relative to the ground-truth idea. The evaluator produces a binary impact judgment and a natural-language rationale explaining why the generated idea is considered impactful or not impactful. ( OPEN z i ℳ , q i ℳ ) = ℰ ⁡ ( g i , d i , d ^ i ℳ ) z{i}^{\mathcal{M}},q_{i}^{\mathcal{M}})=\mathcal{E}(g_{i},d_{i},\hat{d}{i}^{\mathcal{M}}) , where q i ℳ q{i}^{\mathcal{M}} denotes the evaluator’s textual justification and z i ℳ = { 1 , if ​ d ^ i ℳ ​ is judged impactful relative to ​ d i , 0 , otherwise. z_{i}^{\mathcal{M}}=\begin{cases}1,&\text{if }\hat{d}{i}^{\mathcal{M}}\text{ is judged impactful relative to }d{i},\ 0,&\text{otherwise.}\end{cases} The evaluator is instructed to make its decision by considering the relevance of the generated idea to the research goal, based on its scientific contribution and potential impact. Prompt A.2.4 To incorporate the citation impact of the reference idea into the evaluation, we define the instance-level Reference-Grounded Citation-Weighted Impact Score as s i ℳ = z i ℳ ​ ( y i + 1 ) s_{i}^{\mathcal{M}}=z_{i}^{\mathcal{M}}(y_{i}+1) Equivalently, s i ℳ = { y i + 1 , if ​ d ^ i ℳ ​ is judged impactful relative to ​ d i , 0 , otherwise. s_{i}^{\mathcal{M}}=\begin{cases}y_{i}+1,&\text{if }\hat{d}{i}^{\mathcal{M}}\text{ is judged impactful relative to }d{i},\ 0,&\text{otherwise.}\end{cases} The addition of one ensures that an impactful generated idea receives a positive score even when the corresponding reference idea belongs to the lowest ordinal citation-impact class. Therefore, the reward reflects both the binary impactfulness judgment and the citation-normalized importance of the reference idea. The overall RGCW-IIE score of model ℳ \mathcal{M} is computed as the average score over all evaluation instances: = 1 N ​ ∑ i = 1 N s i ℳ =\frac{1}{N}\sum_{i=1}^{N}s_{i}^{\mathcal{M}} A higher value of RGCW-IIE indicates that the model more frequently generates ideas judged to be impactful relative to the ground-truth reference ideas, particularly for examples whose references have higher citation-normalized impact. For bounded comparison, we define the normalized score: nRGCW ​ - ​ IIE ​ ( ℳ ) = ∑ i = 1 N z i ℳ ​ ( y i + 1 ) ∑ i = 1 N ( y i + 1 ) . \mathrm{nRGCW\text{-}IIE}(\mathcal{M})=\frac{\sum_{i=1}^{N}z_{i}^{\mathcal{M}}(y_{i}+1)}{\sum_{i=1}^{N}(y_{i}+1)}. Since y i ∈ { 0 , 1 , … , K − 1 } y_{i}\in{0,1,\ldots,K-1} and z i ℳ ∈ { 0 , 1 } z_{i}^{\mathcal{M}}\in{0,1} , the normalized score satisfies 0 ≤ nRGCW ​ - ​ IIE ​ ( ℳ ) ≤ 1 0\leq\mathrm{nRGCW\text{-}IIE}(\mathcal{M})\leq 1 . We also define IR (Impact Rate) which denotes the percentage of generated ideas judged as impactful by the evaluator, while I’Idea (Impactful Ideas) denotes the absolute number of generated ideas classified as impactful. 6.3 Reward Model Performance We first evaluate the reward model on the ordinal citation-impact prediction task. As shown in Table 2 , the proposed RM substantially outperforms all zero-shot baselines across every metric. In particular, it exceeds the strongest zero-shot baseline GPT-5.The same trend is also reflected in rank-sensitive and agreement-based metrics, where the fine-tuned reward model reaches a Spearman correlation of 0.756 and a quadratic weighted kappa (QWK) of 0.754. These results indicate that impact prediction benefits strongly from goal-conditioned supervised adaptation, and that the learned reward model captures ordinal structure in citation-normalized impact substantially better than prompt-only zero-shot evaluation. Table 3 further shows that the fine-tuned reward model improves class-wise prediction quality across most impact bands. Relative to zero-shot baselines, the model reduces error substantially in the zero/very-low, low, medium, and high categories, while remaining competitive even on the very-high-impact category, which is typically the most difficult due to its rarity and broader semantic variability. Figure 3 compares the confusion matrices of the zero-shot and fine- The ai search story also surfaces in Google AI Releases TimesFM 3 for..., adding another angle. as detailed in the full paper on Arxiv The ai search story also surfaces in House Intelligence Committee Warns AI Safeguards..., adding another angle.

Comments (0)

No comments yet

Be the first to share your thoughts!