Back to AI Research

AI Research

Reinforcing Agentic Creativity in Scientific Ideati... | AI Research

Key Takeaways

  • What the paper is about Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictabl...
  • Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation.
  • Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered.
  • We introduce AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning.
  • Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea).
Paper AbstractExpand

Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying what kind of creativity to pursue is critical. Overall, our results suggest that creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs.

What the paper is about

Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying what kind of creativity to pursue is critical. Overall, our results suggest that creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs. The same ai evaluation question is explored in Learning to Ideate for Scientific Impact, which adds a research perspective.

What it covers

Reinforcing Agentic Creativity in Scientific Ideation with Night Science Priyanka Kargupta † † thanks: Work completed while interning at Microsoft. Corresponding authors: [email protected] , [email protected] Silviu Cucerzan Shweti Mahajan Affiliation: University of Illinois Urbana-Champaign Microsoft Microsoft Research Allen Herring Jiawei Han Ryen W. White Sujay Kumar Jauhar Abstract Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist , an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying what kind of creativity to pursue is critical. Overall, our results suggest that creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs . Reinforcing Agentic Creativity in Scientific Ideation with Night Science Priyanka Kargupta 1,* Silviu Cucerzan 2 Shweti Mahajan 3 Allen Herring 2 Jiawei Han 1 Ryen W. White 2 Sujay Kumar Jauhar 2,* 1 University of Illinois Urbana-Champaign 2 Microsoft 3 Microsoft Research Correspondence: [email protected] , [email protected] Blog: pkargupta.github.io/night_scientist Code: microsoft/ai_night_scientist

  • Work completed while interning at Microsoft. 1 Introduction Large language models (LLMs) have excelled at structured, systematic tasks with clear verifiability (e.g., coding and quantitative reasoning), where their performance is often improved with reinforcement learning (RL) ( Guo et al., 2025 ; Wang et al., 2025 ; Zhong and Wang, 2024 ; Liu et al., 2024 ) . This success is consistent with a broader tendency toward minimizing token entropy ( Agarwal et al., 2025 ) , where models favor high-probability outputs that reflect frequent patterns and expected answers in training data ( McCoy et al., 2024 ) . While LLMs have increasingly been applied to scientific ideation ( Si et al., 2025 ; Gottweis et al., 2025 ) , they lack originality ( Zhao et al., 2025 ) , tend to generate homogeneous outputs ( Wenger and Kenett, 2025 ) , and even plagiarize at nontrivial rates ( Gupta and Pruthi, 2025 ) . Ultimately, this directly conflicts with the key attributes of open-ended, creative scientific discovery : novelty, diversity, and serendipity. Prior methods depict discovery as a highly structured process, involving hypotheses derived from prior observations and/or data, tested against evidence, and refined or rejected accordingly ( Gottweis et al., 2025 ; Agarwal et al., 2026 ) . But this is only part of the scientific discovery process. They typically view creativity as only an attribute of ideas rather than as part of the process, treating individual actions as fixed behaviors (e.g., search, write) and executing them through scaffolded pipelines ( Gu et al., 2024 ; Lu et al., 2024 ) . While this supports the critical reasoning crucial for validating existing ideas, this overlooks the creative reasoning necessary for discovering new ones. Both are complementary and crucial perspectives of scientific discovery, referred to as day and night science ( Wechsler et al., 2018 ; Halpern, 2007 ; Stent, 1988 ; Yanai and Lercher, 2019 ) . Night science captures the often-neglected nature of human-driven discovery: loosely structured and dynamic exploration driven by highly creative actions, such as exploring distant analogies (e.g., biology-inspired technology), considering alternative perspectives after serendipitous encounters (e.g., debates with colleagues from other domains), and acting on partly-formalized intuitions (e.g., abandoning status quo assumptions). Together, day and night science form a spectrum that human researchers have smoothly traversed to uncover breakthroughs ( de Chantal and Markovits, 2022 ; Dwyer et al., 2025 ) , such as chemotherapy and penicillin ( Hirsch, 2006 ; Ligon, 2004 ) . Figure 1: Given an input task, creativity can be injected into three different axes of reasoning: action, process , and outcome . Moreover, each level can be executed with varying degrees of creativity . We hypothesize that LLMs can better traverse this spectrum if given explicit control over when and how to deviate from high-probability output, enabling targeted doses of night science while maintaining day science’s goal-directed behavior. To test this, we introduce AI Night-Scientist , an RL-based agentic framework that incentivizes models to exercise this control effectively. Grounded in cognitive science ( Lubart, 2001 ; Cohen, 1989 ; Dwyer et al., 2025 ; Harvey and Berry, 2023 ) , it embeds creativity along three axes of reasoning (Figure 1 ):

β€’ Action-level ( how ): Associates each action with a creativity level, allowing the same action to vary from conventional, high-probability behavior to more exploratory and unconventional behavior.

β€’ Process-level ( when ): Controls when to shift between lower- and higher-creativity actions based on how the multi-step trajectory unfolds.

β€’ Outcome-level ( what ): Captures the creativity of the resulting idea, favoring outputs that are novel while remaining relevant and useful. We utilize this framework to train an LLM-based agent using GRPO ( Shao et al., 2024 ) for generating scientific research proposals , where identifying promising ideas often requires long-horizon creative reasoning beyond immediately verifiable evidence. Overall, our work argues that LLM-based support for scientific discovery should span the full day-to-night science spectrum. AI Night-Scientist shows that conventional LLMs do not naturally navigate this spectrum effectively, but targeted reinforcement learning can reshape when and how they depart from structured, predictable reasoning, leading to more creative outcomes. Our contributions can be summarized as: 1. We introduce AI Night-Scientist , an RL-based agentic framework that explicitly represents creativity across actions, reasoning processes, and outcomes for long-horizon scientific ideation. 2. We show that creativity is more than sampling stochasticity: explicit guidance on how to be creative via RL helps an agent learn when creative deviations are useful, while higher temperature does not. 3. Empirically, AI Night-Scientist produces more diverse scientific proposals than its base model, expanding the range of research directions by 27.8% and contribution types by 14.9%, while improving predicted impact by up to 32.0 percentage points and originality by 66.2 points. 2 Background and Related Work Scientific discovery has been characterized as an interplay between structured, hypothesis-driven day science and more exploratory, intuition- and serendipity-driven night science ( Stent, 1988 ; Yanai and Lercher, 2019 ) . Classic theories of creativity characterize creative thought through remote associations between otherwise distant concepts ( Mednick, 1962 ) , generative and exploratory modes of cognition ( Ward et al., 1999 ) , and outcomes that are both original and useful ( Runco and Jaeger, 2012 ) . Creativity can also vary in degree: prior work describes a continuum of creative behavior ( Cohen, 1989 ) , with problem-solving strategies ranging from paradigm-preserving to paradigm-stretching and paradigm-breaking ( McFadzean, 1998 ) . Related accounts distinguish between exploring existing conceptual spaces and transforming them to enable qualitatively new ideas ( Boden, 1998 ) , while theories of creative ideation show that originality can arise either by flexibly exploring many conceptual directions or by persistently exploring a few in greater depth ( Nijstad et al., 2010 ) . Together, these perspectives motivate our view of creativity as varying both where it enters reasoning (within actions, processes, and outcomes) and how strongly it is expressed. This view contrasts with most current LLM-based approaches to scientific discovery. Existing research agents broaden ideation through retrieval, search, multi-agent interaction, and iterative generation ( Gu et al., 2024 ; Lu et al., 2024 ; Kargupta et al., 2025a ; Gottweis et al., 2025 ) , but generally treat actions such as searching, debating, and writing as fixed behaviors and primarily assess creativity in the resulting idea. This leaves little control over how creatively individual actions are performed or when creative deviations should occur throughout reasoning. Relatedly, Kargupta et al. (2025b) find that LLMs struggle with the metacognitive awareness needed to monitor and adapt their reasoning, further limiting their ability to flexibly shift between critical and creative modes. Recent work has begun targeting the mechanisms that produce creative scientific ideas. O’Neill et al. (2025) use structured assumption inversion to generate novel hypotheses, inspiring the spark action in our framework, while Kargupta et al. (2025c) and Kargupta et al. (2026) use retrieval to identify research gaps and surface interdisciplinary inspiration. Other work instead learns scientific capabilities directly: GIANTS ( He-Yueya et al., 2026 ) trains smaller models to anticipate scientific insights, while Tong et al. (2026) study whether models can learn scientific taste. Together, these approaches suggest that not only scientific outputs, but also the processes that produce them, can be shaped through structure and supervision. Reinforcement learning offers a way to shape this process without prescribing exactly how discovery should unfold, although open-ended ideation has no single correct answer and must balance qualities such as novelty, relevance, feasibility, and usefulness ( Afzal et al., 2025 ) . Moreover, useful creative exploration requires more than simply injecting randomness ( Schmidhuber, 2010 ) . Our work builds on these ideas by learning both how creatively individual actions should be performed and when different levels of creativity are useful across a reasoning trajectory. 3 AI Night-Scientist : A Creativity-Aligned Agentic Framework Figure 2: AI Night-Scientist consists of a: (1) rollout phase, where an LLM builds a reasoning trajectory f f by iteratively selecting and executing actions with creativity levels, and (2) reward phase during training, where the reward is computed over the final proposal. We propose AI Night-Scientist , as illustrated in Figure 2 , which represents creativity directly within the agent’s reasoning trajectory rather than only in its final output. We extend the ReAct-style reasoning setting ( Yao et al., 2023 ) : at each step, the agent selects an action together with a creativity level that specifies how that action should be carried out, executes it, and updates its current idea state. Repeating this process produces a trajectory through the space of possible ideas, allowing the agent to move between more familiar and more unexplored directions as reasoning unfolds. This lets us represent creativity at three levels (Figure 1 ): action-level creativity captures how an action is performed, process-level creativity captures when to shift between lower- and higher-creativity actions, and outcome-level creativity captures the novelty and usefulness of the resulting idea. 3.1 Multi-Level Representation of Creative Reasoning We define each level below before applying the framework to scientific research proposal generation: Definition 3.1 ( Action-Level Creativity ) . Let π’œ \mathcal{A} denote the action space for a task 𝒯 \mathcal{T} . For each action a ∈ π’œ a\in\mathcal{A} , we define an ordered set of creativity levels π’ž a = { c a ( 1 ) , … , c a ( K a ) } \mathcal{C}{a}={c{a}^{(1)},\dots,c_{a}^{(K_{a})}} , where each level gives a natural-language description of how a a should be performed. Lower levels describe more conventional, high-probability behavior, while higher levels describe increasingly exploratory or unconventional behavior. The number of levels K a K_{a} may vary across actions. Definition 3.2 ( Process-Level Creativity ) . Let Ο„ = ( ( a 1 , c 1 , o ^ 1 ) , … , ( a N , c N , o ^ N ) ) \tau=\bigl((a_{1},c_{1},\hat{o}{1}),\dots,(a{N},c_{N},\hat{o}{N})\bigr) denote a reasoning trajectory, where a i ∈ π’œ a{i}\in\mathcal{A} is the action selected at step i i , c i ∈ π’ž a i c_{i}\in\mathcal{C}{a{i}} is its creativity level, and o ^ i \hat{o}{i} is the resulting intermediate output. At each step, the agent selects the next action-level choice ( a i , c i ) (a{i},c_{i}) based on the task 𝒯 \mathcal{T} and the preceding trajectory Ο„ 1 : i βˆ’ 1 \tau_{1:i-1} . Process-level creativity reflects how effectively the agent sequences and adapts these choices over time. Higher process-level creativity means varying the degree of action-level creativity to benefit the evolving reasoning process, rather than consistently favoring either low- or high-creativity behavior. Definition 3.3 ( Outcome-Level Creativity ) . Let o o denote the final outcome produced for a task 𝒯 \mathcal{T} . Outcome-level creativity captures the creativity of o o itself, based on two complementary properties: its novelty 𝒩 ⁑ ( o ) \mathcal{N}(o) and its usefulness 𝒰 ⁑ ( o ) \mathcal{U}(o) ( Harvey and Berry, 2023 ) . Novelty measures how much o o departs from existing or familiar solutions, while usefulness measures how valuable, appropriate, or effective it is for the task. An outcome should exhibit both in order to be considered creative. Together, this multi-level representation separates how creativity is expressed within an action, when different degrees of creativity are useful across reasoning, and what the process ultimately produces. We represent action-level creativity in natural language to make these choices interpretable and give the user direct semantic control over how and to what extent each action may deviate from its conventional execution. The number and meaning of creativity levels can also vary by action. These action-level choices accumulate into the reasoning trajectory. Intermediate outputs may not appear directly in the final result, but they can shape later decisions and influence when greater or lesser deviation is useful. Greater creativity may help open new directions, while lower-creativity behavior may be better suited for developing or refining promising ones. Making these choices well requires the model to monitor how its reasoning is progressing and adapt accordingly, a key aspect of metacognitive control that is lacking in existing models ( Kargupta et al., 2025b ) . 3.2 Task Formulation: Research Proposal Generation Given a high-level research problem p p , the agent uses the action space π’œ \mathcal{A} in Table 1 to explore and develop a long-horizon research proposal over a multi-step trajectory Ο„ \tau . Search , Debate , and Spark support multiple creativity levels, allowing the agent to vary how broadly it searches, whose perspectives it considers, and how strongly it challenges existing assumptions. Write consolidates the trajectory into proposal text, while Stop ends the process and returns the final proposal o o . We keep Write fixed so that the final proposal primarily reflects the exploration that preceded it. Action π’œ \mathcal{A} Mechanism & Output o ^ \hat{o} Creativity Levels c ∈ π’ž a c\in\mathcal{C}{a} a Search Generate query β†’ \rightarrow retrieve arXiv papers based on embedding similarity L1: Search proposal-specific background using title and core terms. L3: Search tangentially related background for broader context. L5: Search distant domains, alternate perspectives, or broader questions. b Debate Select participants and topic β†’ \rightarrow retrieve relevant papers β†’ \rightarrow simulate discussion L1: Discuss proposal specifics with a close-domain colleague. L3: Debate with a peer from the same field but a different topic. L5: Explore with experts from distant disciplines in an open-ended debate. c Spark Identify assumption ( Bit ) β†’ \rightarrow invert it ( Flip ) β†’ \rightarrow reframe it ( Spark ) L1: Challenge a narrow assumption specific to the current proposal. L3: Challenge a meaningful assumption underlying the approach. L5: Challenge a broad, field-level assumption through radical reframing. d Write Synthesize prior trajectory into a proposal draft or revision Fixed: consolidate prior trajectory; preserve credit assignment e Stop End trajectory & return final proposal Judge final proposal by 𝒫 \mathcal{P} (precedence), β„± \mathcal{F} (feasibility), and β„› \mathcal{R} (relevance) Table 1: Action space for research proposal generation with low, mid and high creativity levels shown. Table 9 in Appendix E.1 includes all levels; full action prompts provided in Appendix F . We focus on proposals with a scope similar to long-term research grants, where ideas are intended to guide work over several years rather than describe an immediately executable experiment. This makes the setting well suited to studying creative reasoning: proposals must remain grounded in existing work, yet many of their central ideas cannot be directly verified at inference time because the required experiments or data may not yet exist. The task therefore rewards reasoning that can move beyond established directions while still producing ideas that are feasible and relevant to p p . 3.3 Reinforcing Adaptive Creativity To teach models when and how different degrees of creativity are useful, we leverage reinforcement learning. RL allows the quality of the final idea to shape the reasoning process without prescribing how that process should unfold. Our training pipeline consists of two phases: (i) a rollout phase , where the agent constructs a reasoning trajectory Ο„ \tau using the creativity-aligned action space above, and (ii) a reward phase , where the resulting proposal o o is evaluated. We use Group Relative Policy Optimization (GRPO) ( Shao et al., 2024 ) , which learns from relative rewards across sampled trajectories without requiring a separate critic. In our primary setting, the reward is applied only to the final proposal, requiring the model to learn which actions to take, how creatively to perform them, and how to sequence them based on their eventual effect on proposal quality. 3.3.1 Rewarding Proposal Quality Scientific proposals contain multiple ideas whose novelty and feasibility may differ substantially, making a single holistic quality judgment difficult to ground. We therefore first decompose the final proposal into its atomic research ideas, separating what each phase proposes from how it plans to execute it. We then evaluate each component against the original research problem and relevant literature along three dimensions: precedence , feasibility , and relevance (Table 2 ). Dimension Definition Positive Example (Reward ↑ \uparrow ) Negative Example (Reward ↓ \downarrow ) Precedence Distance from prior work and existing approaches. Core idea opens genuinely new technical directions. Recombines familiar components without new insight. Feasibility Credibility & specificity of the execution plan. Clear methods with precedent or scoped evaluation strategy. Broad, underspecified plan or implausible scope. Relevance Alignment of the proposal with problem p p . Addresses a critical bottleneck in the target domain. Drifts off-topic or fails to engage with the core challenges of p p . Table 2: Reward dimensions for proposal generation with examples. Reward prompts in Appendix E . We average precedence and relevance across the proposal’s atomic ideas, but take the minimum feasibility across experimental plans, since a single infeasible phase may compromise the executability of the overall proposal. The resulting three proposal-level scores are then weighted equally to form the outcome reward: R out = 1 3 ​ ( R 𝒫 + R β„± + R β„› ) R{\mathrm{out}}=\frac{1}{3}\left(R_{\mathcal{P}}+R_{\mathcal{F}}+R_{\mathcal{R}}\right) . These rewards are designed to provide targeted training signals for proposal quality, with atomic ideas evaluated separately and precedence and feasibility grounded in retrieved literature. Outcome-only training still leaves the model to discover useful creative behaviors through their eventual effect on proposal quality. We therefore explore two additional forms of guidance: exposing the model to occasional serendipitous actions early in training, and directly rewarding properties of the intermediate reasoning process. Neither is required by the framework; we study whether either provides additional benefit beyond the final outcome reward. 3.3.2 Incentivizing Exploration through Serendipitous Actions Human discovery is often shaped by unexpected encounters that expose researchers to directions they may not have deliberately pursued ( Lubart, 2001 ) . Similarly, an LLM early in training may repeatedly select familiar, low-creativity behaviors simply because it has not experienced useful alternatives. We therefore explore a stochastic intervention that occasionally exposes the model to a different action-level choice. With probability Pr ⁑ ( swap ) \Pr(\textit{swap}) , the selected pair ( a i , c i ) (a_{i},c_{i}) is replaced by a sampled alternative ( a i β€² , c i β€² ) (a^{\prime}{i},c^{\prime}{i}) , and a decay factor Ξ³ \gamma gradually reduces this probability over training. The goal is not to prescribe random exploration, but to expose the model early on to creative behaviors whose value it can later learn from the resulting reward. 3.3.3 Exploring Process-Level Rewards Our primary models receive only the outcome reward above, allowing process-level behavior to emerge from credit assigned to the final proposal. We additionally investigate whether directly rewarding intermediate reasoning provides further benefit. The process reward evaluates each intermediate action along two complementary dimensions. Exploration measures whether a i a_{i} introduces a new direction relative to Ο„ 1 : i βˆ’ 1 \tau_{1:i-1} , while contribution measures how much a i a_{i} ultimately contributes to the final proposal o o . We score both dimensions for each intermediate action and average them across the trajectory to obtain R proc R_{\mathrm{proc}} , which then forms the final reward R po = 1 2 ​ ( R proc + R out ) R_{\text{po}}=\frac{1}{2}\left(R_{\mathrm{proc}}+R_{\mathrm{out}}\right) . Full details and prompts are provided in Appendix L . 4 Experiments We train all models using verl ( Sheng et al., 2024 ) with GRPO ( Shao et al., 2024 ) on Qwen3-8B-Base and Qwen3-14B-Base . Each agent trajectory consists of up to 5 actions. For the search action, we index arXiv 1 1 1 https://arxiv.org/ and use embedding-based retrieval with GPT-4.1 as an external judge within the reward pipeline. Full training details are provided in Appendix M . 4.1 Dataset We construct our dataset from the NSF Awards Database 2 2 2 https://huggingface.co/datasets/davidheineman/nsf-awards , focusing on Computer Science, Engineering, and Mathematics (CSE) awards from 2018 onward. We use a 90 : 10 90{:}10 train/test split, yielding 4,414 training and 491 test awards. Our model receives only the award title as p p , which typically describes a broad, long-horizon research problem; we withhold the award abstract, project outcomes, and associated publications so that generation is not anchored to the funded solution. We focus on CSE to ensure broad coverage in open-access arXiv literature, while retaining substantial interdisciplinary breadth: the training set spans 46 research domains, with 23.1% of subfield labels in core AI/ML and over 48% of label occurrences outside AI/ML and core CS and Engineering (Appendix B ). Original NSF grant proposals are not publicly available, so we construct reconstructed reference proposals from evidence surrounding each funded project. For each award, we collect its abstract, project outcomes report, PI information, and award period, then retrieve likely associated arXiv papers by the award PIs. Candidate papers are ranked using PI overlap, topical similarity, publication timing, and explicit funding acknowledgments. We prompt GPT-5.1 to treat these sources as downstream evidence and reconstruct a plausible pre-award research plan that could have led to the observed project and publications. The reference follows the same structured format as generated proposals, including a summary, background, and multi-phase research plan. These reconstructions are not intended to reproduce the original NSF proposals. Rather, they provide standardized, evidence-grounded references for research directions that were actually funded and subsequently pursued. We use them only as matched references for pairwise evaluation, and the proposal-generating model never receives the oracle information used to construct them. Full reconstruction details and prompts are provided in Appendix K . 4.2 Evaluation Metrics Scientific creativity is difficult to capture with any single automatic metric, particularly because the long-term value of a research idea may not be observable for years. We therefore evaluate models along three complementary dimensions: predicted citation impact , literature-grounded originality , and research idea diversity . Together, these capture both the quality of individual proposals and the range of research directions explored across the dataset. Predicted Citation Impact. For each problem, we compare the generated proposal against its matched reconstructed reference in randomized order and report pairwise win rate. We use SciJudge ( Tong et al., 2026 ) , a 30B-parameter model trained to predict relative citation impact from large-scale community signals, as a proxy for the potential usefulness and downstream value of the proposed research (prompts in Appendix G ). It has an average 82.7 % 82.7% citation prediction accuracy. Literature-Grounded Originality. We compare each generated proposal against its matched reference for originality. To ground the judgment in prior work, we retrieve the closest paper to each proposal and ask GPT-5.1 to make a randomized pairwise comparison relative to the retrieved literature (Appendix I ). Research Idea Diversity. Pairwise metrics capture the quality of individual proposals but not whether a model repeatedly produces the same kinds of ideas. Following the annotation format of Chen et al. (2026) , we use GPT-5.4-mini to classify each proposal by its primary research idea paradigm and, analogously, by its primary contribution type using categories adapted from prior taxonomies ( Wobbrock, 2012 ; Miles, 2017 ) . For each taxonomy, we report the effective number of categories represented (Appendix H ), normalized by the seven available categories; higher values indicate a broader and more balanced range of research directions. The research-paradigm judge Chen et al. (2026) achieves ΞΊ = 0.84 \kappa=0.84 agreement with human judgments on 150 samples. Human-LLM Agreement. As an additional check on our pairwise judges, two human annotators evaluated a subset of proposals. Inter-annotator agreement was 80.0% ( ΞΊ = 0.625 \kappa=0.625 ) for impact and 86.7% ( ΞΊ = 0.766 \kappa=0.766 ) for originality, while human-LLM agreement was 76.7% for SciJudge on impact and 72.4% for GPT-5.1 on originality (Appendix J ). 4.3 Baselines We compare against baselines that vary how creativity is introduced into proposal generation. Zero-shot models directly generate a structured proposal from the input problem without retrieval, action selection, or creativity levels. Temperature retains the same five creativity levels, but maps each level only to decoding temperature, with c ∈ { 1 , … , 5 } c\in{1,\ldots,5} corresponding to T ∈ { 0 , 0.25 , 0.5 , 0.75 , 1.0 } T\in{0,0.25,0.5,0.75,1.0} . This tests whether increased stochasticity alone can reproduce the benefits of semantic creativity guidance. ReAct ( Yao et al., 2023 ) uses the same action space π’œ \mathcal{A} with no corresponding π’ž a \mathcal{C}_{a} . Creative baselines apply our full creativity-aligned action space and natural-language creativity levels at inference time, but receive no RL training. We also train several RL variants with the same outcome-level reward as AI Night-Scientist : + Temp and + ReAct apply RL to the corresponding baselines above, while + Search + Write retains semantic creativity levels but restricts π’œ \mathcal{A} to retrieval and writing, similar to retrieval-augmented ideation systems ( Si et al., 2025 ; Wang et al., 2024 ) . Together, these separate the effects of semantic creativity guidance, RL, and the broader action space. Finally, we compare against GIANTS ( He-Yueya et al., 2026 ) , an RL-trained model for scientific insight anticipation. Because GIANTS expects two input papers, we retrieve the two most relevant arXiv papers using the award title and abstract and reformat them with GPT-4.1 . GIANTS t The ai agents story also surfaces in Google’s Gemini AI Accessed Three Outside..., adding another angle. as detailed in the full paper on Arxiv The ai agents story also surfaces in Claude autonomously improved models across 10..., adding another angle.

Comments (0)

No comments yet

Be the first to share your thoughts!