Aha.
正在载入中英对照阅读…

arXiv:2609.22068 · 中英对照阅读

CodeMidas:从代码本身扩展智能体编程强化学习环境

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Bowen Ye, Lei Li, Shicheng Li, Zihao Yue, Linghao Zhang, Hanglong Lv, Yuanxin Liu, Wenhan Ma, Hao Tian, Rang Li, Jinhao Dong, Yikai Zhao, Xiangwei Deng, Hailin Zhang, Liang Zhao, Qi Liu, Lingpeng Kong, Tong Yang, Fuli Luo

中文速览

编码智能体的强化学习缺少大量既多样又能可靠判分的任务,而现有方法主要依赖 issue、提交记录或文档,覆盖有限。CodeMidas 只使用源代码,让多个智能体先找出可观察的已有功能并改写成待实现任务,再依据原程序运行结果生成测试,同时通过容器复测、对抗性尝试和多轮解题筛掉泄漏、误判或质量不稳定的任务。最终它从 3,185 个开源代码库中构建了 5,545 个跨 23 种语言和 15 个领域的任务,用这些任务训练 MiMo-V2.5 后,五个外部基准全都有提升,例如 DeepSWE 提升 11.7%、ProgramBench 提升 17%、Terminal-Bench 提升 8.5%。这说明源代码本身就能规模化变成带可靠验证器的强化学习环境,不仅扩大了训练数据,还能让编码智能体在修复问题、从头写程序和使用终端等不同软件任务上获得可迁移的能力。

摘要

通过强化学习(RL)训练具备能力的编程智能体,需要多样化的任务和可靠的验证器。开源代码库提供了此类任务的丰富来源,而现有方法通常依赖问题单、提交记录等开发产物,从而限制了可提取任务的范围。为更好地扩展强化学习环境,我们提出了 CodeMidas,这是一条智能体流水线,仅使用源代码作为任务特定输入,将现有代码库中已实现的功能转化为可执行的强化学习环境。CodeMidas 将智能体计算资源分配至环境构建的每个阶段:智能体探索已实现的功能,以制定行为规范;构建基于原始代码执行结果的测试;并通过执行检查和反复的解答滚动执行,对候选任务进行验证和过滤。最终得到的数据集包含来自 3,185 个开源代码库的 5,545 个训练任务,覆盖 23 种编程语言和 15 个技术领域。在这些任务上使用群组相对策略优化(GRPO)训练 MiMo-V2.5,可提升其在五个多样化基准上的性能,涵盖问题修复(DeepSWE 提升 11.7%)、完整程序构建(ProgramBench 提升 17%)以及终端操作(Terminal-Bench v2.1 提升 8.5%)。消融实验表明,增加高质量训练任务的数量能够提升性能。轨迹分析显示,经过强化学习训练的智能体表现出更优的行为,例如加强对代码库的探索,并进行更多样化的自验证。这些结果表明,源代码可作为构建强化学习环境的可扩展基础,从而提升编程智能体在多样化软件任务上的能力。

Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input. CodeMidas allocates agentic compute to every stage of environment construction: agents explore implemented functionality to formulate behavioral specifications, construct tests grounded in execution of the original code, and validate and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on these tasks with GRPO improves performance on all five diverse benchmarks, covering issue repair (DeepSWE + 11.7%), whole-program construction (ProgramBench +17%), and terminal work (Terminal-Bench v2.1 +8.5%). Ablations show that increasing the number of high-quality training tasks improves performance. Trajectory analysis shows the RL-trained agent demonstrates better behaviors like increasing codebase exploration and more diverse self-verification. These results establish source code as a scalable foundation for constructing RL environments that improve coding agents across diverse software tasks.

术语表

reinforcement learning (RL)
强化学习(RL)
coding agent
编程智能体
agentic pipeline
智能体流水线
RL environment
强化学习环境
source code
源代码
codebase
代码库
behavioral specification
行为规范
executable test
可执行测试
verifier
验证器
execution-based evaluation
基于执行的评测
post-rollout filtering
滚动执行后过滤
solution rollout
解答滚动执行
self-verification
自验证
task diversity
任务多样性
generalization
泛化
GRPO
群组相对策略优化
MiMo-V2.5
MiMo-V2.5
CodeMidas
CodeMidas
DeepSWE
DeepSWE
ProgramBench
ProgramBench
Terminal-Bench v2.1
Terminal-Bench v2.1
SWE-bench
SWE-bench
SWE-bench Pro
SWE-bench Pro
R2E-Gym
R2E-Gym
SWE-smith
SWE-smith
SWE-Flow
SWE-Flow
SWE-Hub
SWE-Hub
unit test
单元测试
reference patch
参考补丁
reward model
奖励模型