Back to AI Research

AI Research

SciWalker: Synthesizing Scientific Coding Problems... | AI Research

Key Takeaways

  • What the paper is about Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data.
  • Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data.
  • However, such data remain scarce because manually authoring realistic problems is costly and time-consuming, while systematically covering diverse scientific domains and algorithmic combinations remains challenging.
  • To address this, we introduce SciWalker, a framework for synthesizing scientific coding problems through operator-chain sampling and execution feedback.
  • The framework combines scientific library interfaces with operation modes to instantiate operators, organizes them into operator graphs, and samples operator chains as computational workflow cues.
Paper AbstractExpand

Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data. However, such data remain scarce because manually authoring realistic problems is costly and time-consuming, while systematically covering diverse scientific domains and algorithmic combinations remains challenging. To address this, we introduce SciWalker, a framework for synthesizing scientific coding problems through operator-chain sampling and execution feedback. The framework combines scientific library interfaces with operation modes to instantiate operators, organizes them into operator graphs, and samples operator chains as computational workflow cues. Guided by these cues, we adopt LLMs to generate scientifically grounded problem statements, reference solutions, and tests, with failed generations iteratively repaired using execution feedback. By combining structured workflow composition with verification and quality review, SciWalker enables scalable task generation while promoting scientific grounding, computational diversity, and executability. Using this framework, we construct 8,178 high-quality problems spanning 5 scientific domains and 32 subdomains. To evaluate their training utility, we conduct reinforcement learning on Qwen3.5-9B using the GSPO algorithm. This training improves SciCode subproblem accuracy by 9.9 percentage points, from 29.3% to 39.2%, with gains across scientific code generation, code repair, and reasoning benchmarks. The code for SciWalker is available at this https URL .

What the paper is about

Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data. However, such data remain scarce because manually authoring realistic problems is costly and time-consuming, while systematically covering diverse scientific domains and algorithmic combinations remains challenging. To address this, we introduce SciWalker, a framework for synthesizing scientific coding problems through operator-chain sampling and execution feedback. The framework combines scientific library interfaces with operation modes to instantiate operators, organizes them into operator graphs, and samples operator chains as computational workflow cues. Guided by these cues, we adopt LLMs to generate scientifically grounded problem statements, reference solutions, and tests, with failed generations iteratively repaired using execution feedback. By combining structured workflow composition with verification and quality review, SciWalker enables scalable task generation while promoting scientific grounding, computational diversity, and executability. Using this framework, we construct 8,178 high-quality problems spanning 5 scientific domains and 32 subdomains. To evaluate their training utility, we conduct reinforcement learning on Qwen3.5-9B using the GSPO algorithm. This training improves SciCode subproblem accuracy by 9.9 percentage points, from 29.3% to 39.2%, with gains across scientific code generation, code repair, and reasoning benchmarks. The code for SciWalker is available at this https URL . The openai story also surfaces in AI Agents Going Rogue Renew Calls..., adding another angle.

What it covers

SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback Chenxi Li 1,2 Wenxuan Zeng 2,3 Yun Luo 2†, ‡ \ddagger Fangchen Yu 2 Peng Ye 2 Yu Cheng 4 Jun Zhang 1 ‡ \ddagger 1 The Hong Kong University of Science and Technology 2 Shanghai AI Laboratory 3 Tsinghua University 4 Nanyang Technological University † Project Lead, ‡ \ddagger Corresponding authors Abstract Improving the scientific coding capabilities of large language models (LLMs) requires high-quality training data. However, such data remain scarce because manually authoring realistic problems is costly and time-consuming, while systematically covering diverse scientific domains and algorithmic combinations remains challenging. To address this, we introduce SciWalker, a framework for synthesizing scientific coding problems through operator-chain sampling and execution feedback. The framework combines scientific library interfaces with operation modes to instantiate operators, organizes them into operator graphs, and samples operator chains as computational workflow cues. Guided by these cues, we adopt LLMs to generate scientifically grounded problem statements, reference solutions, and tests, with failed generations iteratively repaired using execution feedback. By combining structured workflow composition with verification and quality review, SciWalker enables scalable task generation while promoting scientific grounding, computational diversity, and executability. Using this framework, we construct 8,178 high-quality problems spanning 5 scientific domains and 32 subdomains. To evaluate their training utility, we conduct reinforcement learning on Qwen3.5-9B using the GSPO algorithm. This training improves SciCode subproblem accuracy by 9.9 percentage points, from 29.3% to 39.2%, with gains across scientific code generation, code repair, and reasoning benchmarks. The code for SciWalker is available at https://github.com/lichenx1/SciWalker . 1 Introduction Scientific code plays a crucial role in translating scientific theories into computational practice across a wide range of disciplines ( Virtanen et al., 2020 ) . High-quality, diverse training data have been shown to improve the scientific reasoning and code generation capabilities of large language models (LLMs) ( Zhang et al., 2024 ; Wei et al., 2024 ) . However, such data remain scarce because constructing realistic scientific coding problems typically requires substantial expert effort to formulate scientific contexts, compose computational procedures, implement reference solutions, and design reliable tests ( Tian et al., 2024 ) . Manually carrying out these steps is costly and time-consuming ( Lai et al., 2023 ) , while systematically covering diverse scientific domains, subdomains, and algorithmic combinations remains challenging. To address these challenges, we introduce SciWalker, a framework for scalable synthesis of multistep scientific coding problems. Our central idea is to leverage scientific computing libraries as computational building blocks and compose their interfaces into diverse scientific workflows as illustrated in Figure 1 . Specifically, SciWalker combines library interfaces with operation modes to instantiate operators, organizes them into operator graphs, and samples operator chains as workflow cues. Guided by these cues and corresponding scientific contexts, LLMs generate problem statements, reference solutions, and tests with well-defined computational objectives. Execution feedback is then used to identify and iteratively repair failures in the generated code and tests, while quality review further filters the resulting problems. Based on SciWalker, we construct a large-scale dataset of 8,178 scientific coding problems spanning five major domains, including mathematics, physics, chemistry, biology, and materials science. The problems are generated from operator chains of varying lengths and subsequently refined through execution-guided repair and quality review, resulting in diverse computational workflows grounded in realistic scientific contexts. Collectively, the dataset contains 37,820 computational substeps, with each problem integrating multiple interdependent operations rather than isolated function calls. To assess the training value of the generated data, we conduct reinforcement learning on Qwen3.5-9B ( Qwen Team, 2026 ) using step-level execution rewards and GSPO policy updates ( Zheng et al., 2025 ) . This training improves SciCode subproblem accuracy from 29.3% to 39.2%, a gain of 9.9 percentage points over the baseline. Gains also extend to scientific code generation, code repair, and reasoning benchmarks, with improvements of 17.3 percentage points on DS-1000 and 6.0 percentage points on MATH-500. The main contributions of our work are as follows: Figure 1: Scientific code training data generation workflow. Scientific library APIs are combined with operation modes to form operators, which are used to construct a graph and sample candidate operator chains. Large language models screen these chains and select scientific contexts, generate tasks, reference implementations, and tests, and then refine the problem statements accordingly. Execution validation, feedback-driven repair, and quality review yield high-quality training data.

• We propose SciWalker, which combines scientific operator composition with execution feedback to guide large language models in automatically generating and revising multistep scientific coding problems.

• We construct 8,178 high-quality scientific problems covering 5 scientific domains and 32 subdomains, accompanied by reference solutions and tests.

• We conduct reinforcement learning experiments to assess the training value of the generated data. The trained model shows improvements on both in-domain scientific coding tasks and out-of-domain code repair and reasoning benchmarks. 2 Related Work Scientific and code data construction. SciCode ( Tian et al., 2024 ) provides real research problems, reference solutions, and tests curated by scientists to evaluate multistep scientific coding capabilities. Automated data construction approaches draw on existing problems, scientific literature, and code resources. SciInstruct ( Zhang et al., 2024 ) augments existing scientific problems with model-generated reasoning, refined through self-review and revision. WildSci ( Liu et al., 2026 ) synthesizes scientific multiple-choice questions from peer-reviewed literature, supporting reinforcement learning through unambiguous answer evaluation. UniScientist ( Li et al., 2026 ) combines validated scientific claims with evidence retrieval to generate open-ended research questions, accompanied by rubrics revised and validated by language models and domain experts. Magicoder ( Wei et al., 2024 ) uses open-source code snippets to guide the generation of programming tasks. OpenCodeInterpreter ( Zheng et al., 2024 ) trains code models on multi-turn interactions incorporating execution and human feedback for iterative refinement. Scaling multistep scientific coding data requires diverse combinations of scientific computations, coherent tasks with clear scientific objectives, and reference solutions with tests. Our work addresses these requirements through scientific operator composition and scientific context design, using execution feedback to repair and select the generated problems. 3 Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback Scientific coding tasks often involve interdependent computations, where later steps build on functions implemented earlier. We therefore adopt a multi-turn format that preserves these dependencies and supports stepwise validation. Following SciCode ( Tian et al., 2024 ) , each problem consists of a main problem and an ordered sequence of subproblems, as illustrated in Figure 2 . The main problem defines the overall objective and dependencies. Each subproblem provides a task description, background, and interface specification. At each turn, the model implements the current subproblem using its specification and previously generated code. Figure 2: Multi-turn scientific coding problem format. Our framework combines LLMs with programmatic workflows to construct scientific problems containing problem statements, reference solutions, and tests. The overall workflow consists of four stages as illustrated in Figure 1 : (1) Scientific operator construction; (2) Operator graph sampling; (3) Scientific problem generation; (4) Validation, Repair, and Final refinement. 3.1 Operator Construction We first organize scientific library interfaces by computational purpose into categories such as optimization, dynamics, and matrix decomposition. Each API, defined as a function, class, or method, is combined with operation modes such as single-case execution, parameter sweeps, and result checking. We define each API × \times mode combination as an operator, which serves as a node for graph construction and operator chain sampling. Representing different uses of the same API as distinct operators expands a finite collection of APIs into a larger operator space, enabling more diverse computational workflows for scientific problem generation. 3.2 Operator Graph Sampling Each node in the operator graph represents an API × \times mode combination, as defined in Section 3.1 . Within each subdomain, the framework scores directed connections between operators according to their input/output compatibility and operational workflows, retaining multiple high-scoring successors for each node. Operator chains are then sampled through random walks. Starting from a selected operator, the framework repeatedly chooses among its retained successors until the target chain length is reached. Varying the starting node and walk path produces diverse operator combinations from the same graph. The resulting chains and their node descriptions serve as computational workflow cues for subsequent scientific problem design. 3.3 Scientific Problem Generation Given a sampled operator chain, the framework generates a problem through three stages: (1) Chain assessment , it evaluates whether the chain can support a coherent computational workflow with a natural scientific context, well-defined target quantities, and meaningful multistep reasoning. (2) Context construction , the framework specifies the input conditions, solution objectives, and roles of individual operators, followed by an independent plausibility review. Rejected contexts are regenerated and reviewed within a limited number of rounds. (3) Problem authoring , using the approved context and operator chain, the framework generates the scientific task and substeps, tests, a reference implementation based on basic numerical functions, and an oracle implementation using high-level scientific APIs. The problem statement is then revised for consistency with the generated code and tests. The final problem contains the main problem, subproblems, input/output specifications, reference solutions, and tests. Substeps are organized according to the scientific task and its dependencies, after which the problem proceeds to execution validation and repair. Figure 3: Dataset coverage and composition. (a) Domain distribution of 32 subdomains. (b) Domain distribution of 8,178 high-quality problems. (c) Distribution of substep counts across the same 8,178 problems. Sector areas represent category proportions, with percentages labeled on the sectors and categories identified in the legends. 3.4 Validation, Repair, and Final Refinement After a problem is generated, it undergoes three stages to ensure its executability and overall quality: (1) Execution validation , the framework checks the problem structure, interfaces, and dependencies, executes substep and end-to-end tests, and performs differential testing between the reference and oracle implementations. The tests cover numerical correctness, parameter variations, boundary cases, and scientific invariants. (2) Execution-guided repair , when validation fails, error messages are fed back to the model to repair the problem statement, tests, or implementations. Each revision is revalidated, and candidates that exceed the repair limit are discarded. (3) Quality refinement , validated problems are polished and assessed for scientific naturalness, task difficulty, implementation complexity, and test strength. Candidates below the quality threshold undergo limited rounds of refinement and renewed validation, after which only qualified problems are retained. 3.5 SciWalker Implementation and Dataset Statistics We use DeepSeek-V4-Flash ( DeepSeek-AI, 2026 ) throughout the entire data generation pipeline. With stage-specific prompts, the same model performs chain review, context selection, problem generation, execution-guided repair, polishing, and quality review. We enable its maximum reasoning setting and use a context window of 200,000 tokens. We organize APIs from scientific Python libraries (Appendix E ) into subdomain-specific operator catalogs, which serve as the basis for our data generation pipeline. The resulting problems span five broad scientific domains—mathematics, physics, chemistry, biology, and materials science, covering 32 subdomains. Figure 3 (a) shows the distribution of these subdomains across the five domains. For each subdomain, we construct 560 operators, yielding 17,920 operators and 772,178 graph edges in total. Operator chains of lengths 3-15 are then sampled as workflow cues for problem generation. After execution-based validation and quality filtering, we retain 8,178 high-quality problems, whose domain distribution is shown in Figure 3 (b). Appendix B reports candidate retention across the generation stages. Each problem comprises 3–6 substeps, resulting in 37,820 substeps overall, with a mean of 4.62 and a median of 5 per problem; the corresponding distribution is shown in Figure 3 (c). The number of substeps does not directly correspond to operator-chain length. As discussed in Section 3.3 , operator chains provide high-level computational workflow cues, while the model structures the final substeps according to the scientific objective and computational dependencies of each problem. 4 Reinforcement Learning for Scientific Coding To evaluate the effectiveness of the data generated by SciWalker for training scientific coding models, we perform reinforcement learning with execution-based rewards. We use GSPO for policy optimization with substep execution rewards. 4.1 Policy Optimization with GSPO We optimize the policy with Group Sequence Policy Optimization (GSPO), using the GSPO-token formulation introduced by Zheng et al. (2025) . For each main problem, we sample multiple complete solution trajectories and expand them into substep-level prompt–response samples. Each prompt contains the current task and preceding code from the same trajectory. The response contains the current substep’s reasoning and final code. Only the generated response tokens contribute to the loss. Responses to the same substep of the same main problem form a group, although their preceding code may differ. Within each group, the advantage of trajectory i i at substep t t is A i , t = r i , t − μ t σ t + δ , A_{i,t}=\frac{r_{i,t}-\mu_{t}}{\sigma_{t}+\delta}, (1) where r i , t r_{i,t} is the training reward defined in Equation 5 , μ t \mu_{t} and σ t \sigma_{t} are the group reward mean and sample standard deviation, and δ \delta is a small constant for numerical stability. For policy optimization, we denote a substep sample by ( x i , y i , A i ) (x_{i},y_{i},A_{i}) , where y i y_{i} is the complete response for one substep. GSPO uses the length-normalized sequence probability ratio ρ i ​ ( θ ) = ( π θ ​ ( y i ∣ x i ) π old ​ ( y i ∣ x i ) ) 1 / | y i | , \rho_{i}(\theta)=\left(\frac{\pi_{\theta}(y_{i}\mid x_{i})}{\pi_{\mathrm{old}}(y_{i}\mid x_{i})}\right)^{1/|y_{i}|}, (2) where π old \pi_{\mathrm{old}} is the fixed old policy. GSPO-token uses ρ ^ i , k ​ ( θ ) = sg ⁡ [ ρ i ​ ( θ ) ] ​ π θ ​ ( y i , k ∣ x i , y i , < k ) sg ⁡ [ π θ ​ ( y i , k ∣ x i , y i , < k ) ] , \widehat{\rho}{i,k}(\theta)=\operatorname{sg}!\left[\rho{i}(\theta)\right]\frac{\pi_{\theta}(y_{i,k}\mid x_{i},y_{i,<k})}{\operatorname{sg}!\left[\pi_{\theta}(y_{i,k}\mid x_{i},y_{i,<k})\right]}, (3) where k k indexes response tokens and sg \operatorname{sg} stops gradient propagation. Every token shares the sequence ratio’s forward value, while gradients propagate through its own log probability. The GSPO-token objective is 𝒥 GSPO ​ - ​ token ​ ( θ ) = 1 | ℬ | ​ ∑ i ∈ ℬ 1 | y i | ​ ∑ k = 1 | y i | min ⁡ [ ρ ^ i , k ​ ( θ ) ​ A i , clip ⁡ ( ρ ^ i , k ​ ( θ ) , 1 − ϵ low , 1 + ϵ high ) ​ A i ] , \mathcal{J}{\mathrm{GSPO\text{-}token}}(\theta)=\frac{1}{|\mathcal{B}|}\sum{i\in\mathcal{B}}\frac{1}{|y_{i}|}\sum_{k=1}^{|y_{i}|}\min!\left[\widehat{\rho}{i,k}(\theta)A{i},,\operatorname{clip}!\left(\widehat{\rho}{i,k}(\theta),1-\epsilon{\mathrm{low}},1+\epsilon_{\mathrm{high}}\right)A_{i}\right], (4) where ℬ \mathcal{B} denotes a batch of substep training samples, and ϵ low \epsilon_{\mathrm{low}} and ϵ high \epsilon_{\mathrm{high}} specify the clipping range. All tokens in a substep response share the same advantage A i A_{i} , so this formulation is equivalent to GSPO in objective value, clipping conditions, and gradient. Additional rollout probability correction and numerical safeguards used in training are detailed in Appendix D . 4.2 Constructing Substep Training Samples After complete trajectories are generated and scored, they are expanded into training samples by substep. Let main problem q q have T q T_{q} steps requiring model responses. Each trajectory then yields T q T_{q} prompt–response samples, and eight trajectories yield 8 ​ T q 8T_{q} samples in total. The sample for trajectory i i at step t t is denoted by ( x q , i , t , y q , i , t ) (x_{q,i,t},y_{q,i,t}) , where x q , i , t x_{q,i,t} contains the current task, interface requirements, and preceding code from the same trajectory, while y q , i , t y_{q,i,t} contains the reasoning and final code generated for the current step. Each sample is optimized only over the valid generated tokens in its current response, with both reasoning and the final code contributing to the loss. Preceding code belongs to the prompt and is not counted again in the loss as a prediction target for the current step. Each substep retains its own reward, and rewards from subsequent steps are not accumulated backward. Each substep receives an execution reward of 1 if all its tests pass and 0 otherwise: r i , t = [ all tests for substep ​ t ​ pass ] . r_{i,t}=\mathbf{1}!\left[\text{all tests for substep }t\text{ pass}\right]. (5) Table 1: Shared reinforcement learning and evaluation settings. Item Shared setting Learning rate 1 × 10 − 6 1\times 10^{-6} Optimizer parameters Adam, β 1 = 0.9 \beta_{1}=0.9 , β 2 = 0.98 \beta_{2}=0.98 , weight decay = 0.1 =0.1 GSPO clipping ϵ low = 3 × 10 − 4 \epsilon_{\mathrm{low}}=3\times 10^{-4} , ϵ high = 4 × 10 − 4 \epsilon_{\mathrm{high}}=4\times 10^{-4} Gradient clipping 1.0 KL / entropy terms KL reward, KL loss, and entropy regularization disabled Training budget 150 steps, 32 problems per step, 8 trajectories per problem Training sampling temperature=1.0, top_p=1.0, top_k=-1 Training input / generation limit 40,000 / 20,000 tokens Training subproblem execution timeout 120 seconds Evaluation settings Thinking, generation limit of 32,768 tokens, execution timeout of 300 seconds Evaluation sampling temperature=0.6, top_p=0.95, top_k=20 5 Experiments 5.1 RL Training and Evaluation Settings The shared training and evaluation settings are summarized in Table 1 . As an additional baseline, we apply PPO-based RL with a 50-step warm-up schedule. Evaluation . We adopt 16 different evaluation benchmarks to evaluate the training effectiveness, which comprise five benchmarks for scientific computing and code generation, six for code repair and execution, and five for reasoning and knowledge: 1. Scientific computing and code generation: SciCode ( Tian et al., 2024 ) , DS-1000 ( Lai et al., 2023 ) , HumanEval ( Chen et al., 2021 ) , LiveCodeBench Code Generation v6 ( Jain et al., 2025 ) , and APPS Introductory ( Hendrycks et al., 2021 ) ; 2. Code repair and execution: HumanEvalFix-Python, HumanEvalFix-Rust, and HumanEvalFix-JavaScript ( Muennighoff et al., 2024 ) , QuixBugs-Java and QuixBugs-Python ( Lin et al., 2017 ) , and LiveCodeBench Execution v2 ( Jain et al., 2025 ) ; 3. Reasoning and knowledge: MATH-500 ( Lightman et al., 2024 ) , GSM8K ( Cobbe et al., 2021 ) , BBH Multistep Arithmetic ( Suzgun et al., 2023 ) , ARC-Easy ( Clark et al., 2018 ) , and MMLU-Pro Computer Science ( Wang et al., 2024 ) . Each trained model uses a fixed checkpoint across benchmarks, and the best checkpoint is selected via the evaluation score of SciCode. Across the 16 benchmarks, we calculate the unweighted mean of the benchmark scores, each averaged over three evaluations. 5.2 Results Comparison SciCode Performance. Figure 4 reports model performance on SciCode, combining our local evaluations with externally reported results. Each local evaluation covers 288 test-set subproblems with scientific background. By applying GSPO with SciWalker-generated training data, Qwen3.5-9B improves from 29.3% to 39.2%, an absolute gain of 9.9 points (33.8% relative). It surpasses substantially larger models such as Qwen3-32B (36.0%) and GPT-OSS-120B (34.0%), while matching GPT-5 mini (39.0%) and approaching Qwen3.5-122B (39.7%). These results demonstrate the effectiveness of SciWalker-generated data in substantially improving scientific reasoning capabilities, particularly for compact models. Figure 4: SciCode performance comparison. Local evaluations report mean subproblem accuracies over three evaluations of 288 subproblems with scientific background. External model scores are reported by Artificial Analysis ( Artificial Analysis, 2026 ) . Table 2: Performance of the Qwen3.5-9B baseline, PPO, and GSPO across 16 benchmarks. Scores (%) are means over three evaluations, rounded to one decimal before calculating the parenthesized percentage-point changes relative to the baseline. Increases and decreases are shown in dark blue and dark red, respectively. SciCode reports subproblem accuracy, code benchmarks report pass@1, and reasoning and knowledge benchmarks report accuracy. Bold indicates the highest mean score in each column, including ties. Scientific computing and code generation Model SciCode DS-1000 HumanEval LCB Code Gen. v6 APPS Introductory Baseline 29.3 45.6 94.7 61.9 69.1 PPO 32.6 (+3.3) 49.4 (+3.8) 93.5 (-1.2) 57.6 (-4.3) 72.3 (+3.2) GSPO 39.2 (+9.9) 62.9 (+17.3) 97.8 (+3.1) 74.9 (+13.0) 84.6 (+15.5) Code repair and execution Model HumanEvalFix Python HumanEvalFix Rust HumanEvalFix JavaScript QuixBugs Java QuixBugs Python LCB Execution v2 Baseline 80.5 51.2 43.3 58.3 77.5 97.1 PPO 82.7 (+2.2) 60.6 (+9.4) 66.9 (+23.6) 58.3 (0.0) 63.3 (-14.2) 97.8 (+0.7) GSPO 93.9 (+13.4) 48.4 (-2.8) 91.1 (+47.8) 65.0 (+6.7) 87.5 (+10.0) 97.4 (+0.3) Reasoning and knowledge Model MATH-500 GSM8K BBH Arithmetic ARC-Easy MMLU-Pro CS Baseline 86.1 91.1 92.9 98.9 84.3 PPO 81.4 (-4.7) 90.4 (-0.7) 89.5 (-3.4) 98.8 (-0.1) 85.0 (+0.7) GSPO 92.1 (+6.0) 96.0 (+4.9) 97.1 (+4.2) 99.0 (+0.1) 84.6 (+0.3)

  • LCB denotes LiveCodeBench. BBH Arithmetic denotes BBH Multistep Arithmetic. HumanEvalFix uses the test version. LiveCodeBench Execution contains 479 execution instances from 92 source problems. Cross-Benchmark Performance. Table 2 compares the Qwen3.5-9B baseline with PPO and GSPO across 16 benchmarks. Overall, GSPO achieves substantial improvements over the baseline across a broad range of benchmarks. The gain is particularly pronounced on SciCode, where GSPO improves subproblem accuracy by 9.9 percentage points, demonstrating the effectiveness of the scientific-coding data generated by SciWalker. This benefit generalizes to other coding tasks that GSPO improves DS-1000 by 17.3 points, LiveCodeBench Code Generation v6 by 13.0 points, and APPS Introductory by 15.5 points. The improvements also extend to out-of-domain tasks that are not directly targeted by the SciWalker data. On code repair, GSPO improves HumanEvalFix-Python and HumanEvalFix-JavaScript by 13.4 and 47.8 points, respectively. It also gains 6.7 points on QuixBugs-Java and 10.0 points on QuixBugs-Python. Moreover, GSPO consistently improves all five reasoning and knowledge benchmarks, such as gains of 6.0 points on MATH-500 and 4.9 points on GSM8K. These improvements are notable because these benchmarks are not coding tasks, suggesting that learning from high-quality scientific-coding trajectories can strengthen more general reasoning capabilities. Overall, GSPO improves 15 of the 16 benchmarks and raises the average score from 72.6% to 82.0%, indicating both the quality of the SciWalker-generated data and its broad generalization value. Although PPO is less effective and less consistent than GSPO, it still improves SciCode by 3.3 points and produces clear gains on several coding-related benchmarks, such as DS-1000 (+3.8), APPS Introductory (+3.2). These gains provide additional evidence that the SciWalker-generated data contain useful learning signals. However, PPO improves only eight benchmarks, while degrading seven, and increases the overall mean by only 1.2 points. Its regressions on LiveCodeBench Code Generation and most reasoning benchmarks suggest that PPO does not exploit these signals as reliably as GSPO. One possible explanation is that critic-based credit assignment introduces estimation errors that make optimization less stable and limit generalization, highlighting the importance of the training objective in fully realizing the value of the generated data. 6 Conclusion In this study, we introduced SciWalker, a framework for synthesizing scientific coding problems through operator-chain sampling and execution feedback. By using sampled operator chains as computational workflow cues, SciWalker generates scientifically grounded and diverse problems and iteratively repairs failed generations. Using this framework, we constructed 8,178 high-quality problems spanning 5 scientific domains and 32 subdomains. We conducted reinforcement learning on Qwen3.5-9B using GSPO algorithm and the training improved SciCode subproblem accuracy by 9.9 percentage points, from 29.3% to 39.2%, showing the effectiveness of the generated coding data. We also show that human-authored open-source code, such as Python libraries, is a valuable resource for synthetic coding data generation. Acknowledgments This work was supported by the Shanghai Artificial Intelligence Laboratory. We are grateful to the authors and open-source communities whose work made this project possible. References Artificial Analysis (2026) Artificial Analysis SciCode benchmark leaderboard . Note: Accessed September 23, 2026 External Links: Link Cited by: Figure 4 . Chen et al. (2021) M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba Evaluating Large Language Models Trained on Code . arXiv preprint arXiv:2107.03374 . External Links: 2107.03374 Cited by: item 1 . Clark et al. (2018) P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge . External Links: 1803.05457 , Link Cited by: item 3 . Cobbe et al. (2021) K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman Training Verifiers to Solve Math Word Problems . arXiv preprint arXiv:2110.14168 . Cited by: item 3 . DeepSeek-AI (2026) DeepSeek-AI DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence . External Links: 2606.19348 , Link Cited by: §3.5 . Hendrycks et al. (2021) D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt Measuring Coding Challenge Competence With APPS . In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks , J. Vanschoren and S. Yeung (Eds.) , Vol. 1 . External Links: Link Cited by: item 1 . Jain et al. (2025) N. Jain, K. Han, A. Gu, W. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code . In International Conference on Learning Representations , Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu (Eds.) , Vol. 2025 , pp. 58791–58831 . External Links: Link Cited by: item 1 , item 2 . Lai et al. (2023) Y. Lai, C. Li, Y. Wang, T. Zhang, R. Zhong, L. Zettlemoyer, W. Yih, D. Fried, S. Wang, and T. Yu DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation . In Proceedings of the 40th International Conference on Machine Learning , A. Krause, E. Brunskill, K. Cho, B. Engelhard To see openai in practice, The Ultimate Vibe Coding Guide (2026... walks through a concrete example. as detailed in the full paper on Arxiv The openai story also surfaces in Google opens early access to AI..., adding another angle.

Comments (0)

No comments yet

Be the first to share your thoughts!