A preprint reports synthetic scientific coding problems that boost an LLM’s SciCode accuracy

Original title: SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback

In one sentence

Researchers describe SciWalker, a system that auto-generates scientific coding problems and reports large training gains on a 9B model.

What were the researchers trying to find out?

The researchers wanted to know whether automatically generated, scientifically grounded coding problems could substitute for scarce, costly human-authored training data, and whether training on such data would meaningfully improve a language model's scientific coding and reasoning abilities.

What did they find?

  1. The authors report constructing 8,178 scientific coding problems across 5 domains and 32 subdomains using operator-chain sampling and execution feedback.

    From the paper: we construct 8,178 high-quality problems spanning 5 scientific domains and 32 subdomains · Abstract

  2. The authors report that reinforcement learning with this data improved SciCode subproblem accuracy on Qwen3.5-9B from 29.3% to 39.2%.

    From the paper: improves SciCode subproblem accuracy by 9.9 percentage points, from 29.3% to 39.2% · Abstract

  3. The study finds gains extending to out-of-domain tasks, including 17.3 points on DS-1000 and 6.0 points on MATH-500.

    From the paper: improvements of 17.3 percentage points on DS-1000 and 6.0 percentage points on MATH-500 · 1 Introduction

  4. The authors report that GSPO improved 15 of 16 benchmarks, raising the average score from 72.6% to 82.0%, while PPO was less consistent.

    From the paper: GSPO improves 15 of the 16 benchmarks and raises the average score from 72.6% to 82.0% · 5.2 Results Comparison

Why we're watching this

Scientific coding benchmarks like SciCode are hard for LLMs partly because realistic, multi-step problems are expensive to write by hand. If synthetic problem generation pipelines like SciWalker genuinely produce useful training signal, this could lower the barrier to building domain-specific coding datasets for physics, chemistry, biology and materials science, and help smaller models close gaps with much larger ones. Worth watching whether independent groups reproduce these gains, whether the generated problems hold up to expert scrutiny, and whether the approach generalises beyond the benchmarks tested here.

What should you keep in mind?

  • Reported gains rely on a single trained model (Qwen3.5-9B) and a fixed set of benchmarks, so generalisation to other model families is untested here. (TechiesJournal observation)
  • External comparison scores for other models on SciCode were sourced from a third-party leaderboard rather than the authors' own evaluation. (stated by the authors)

About this source

Format
Preprint
Peer review
Not peer reviewed
Released
24 Sep 2026
Version covered
arXiv v1 · 24 Sep 2026
Added
28 Sep 2026

Authors

Chenxi Li, Wenxuan Zeng, Yun Luo, Fangchen Yu, Peng Ye and 2 others

Show all 7 authors
  1. Chenxi Li
  2. Wenxuan Zeng
  3. Yun Luo
  4. Fangchen Yu
  5. Peng Ye
  6. Yu Cheng
  7. Jun Zhang

Prepared from the original research with automated assistance and reviewed by a TechiesJournal editor before publication.

Report a correction

Corrections go to the editor and are never published automatically. No account needed.