Post-training coding agents on execution-verified trajectories 基于执行验证轨迹的 Coding Agent 后训练
A coding agent proposes software tasks and solves them in a sandbox. Trajectories that compile, pass the tests and reproduce are kept as SFT and RL data, and the updated model runs the agent in the next round. Each round only helps if the filter is reliable, so most of my work is on the filter and the training signal: credit assignment over long trajectories, reward and verifier design, and auditing agent actions with a checker the agent does not control. Coding Agent 提出软件任务并在沙箱中求解;能编译、通过测试且可复现的轨迹保留为 SFT 与 RL 数据;更新后的模型进入下一轮。每一轮是否有效取决于筛选是否可靠,因此我的工作主要在筛选与训练信号上:长轨迹的信用分配、奖励与验证器设计,以及用 Agent 无法控制的检查器审查其行动。
Agent proposes a coding task Agent 生成编码任务 Sandbox execution 沙箱执行 Compile · test · replay 编译 · 测试 · 复现 Keep only verified trajectories 仅保留通过验证的轨迹 SFT · RL post-training SFT · RL 后训练 Updated model runs the next round 更新后的模型进入下一轮
One round = one model update, trained only on trajectories that passed verification. 一轮对应一次模型更新,训练数据只取通过验证的轨迹。
- Long-horizon credit assignment长程任务的信用分配
- Verifiable rewards (RLVR)可验证奖励(RLVR)
- Reward & verifier design奖励与验证器设计
- Independent audit of agent actionsAgent 行动的独立审查
| Work工作 | Status状态 | What it shows说明 |
|---|---|---|
| Zhipu AI (Z.ai) · Foundation-Model Post-Training 智谱(Z.ai)· 基础模型后训练 | current role 在职 | Coding-agent post-training on execution-verified trajectories. 基于执行验证轨迹的 Coding Agent 后训练。 |
| Farsight Farsight | ICLR 2026 · first author, under review ICLR 2026 · 一作在投 | Future-discounted visual credit assignment for VLM reinforcement learning; trained on 100×H100 at ByteDance. 面向视觉语言模型强化学习的未来折扣视觉信用分配;在字节跳动百卡 H100 上训练。 |
| OfficeBuddy OfficeBuddy | open source · 53★ 开源 · 53★ | Word/Excel agent. Each edit is re-rendered in real Office and checked by a separate visual verifier. Word/Excel 编辑 Agent。每步编辑在真实 Office 中重新渲染,由独立的视觉验证器检查。 |
| Med-Banana trajectories Med-Banana 轨迹 | 88K+ released · 150K+ downloads 开源 88K+ · 下载 150K+ | Successful and failed editing trajectories, released on GitHub and Hugging Face for post-training. 成功与失败的编辑轨迹,已在 GitHub 与 Hugging Face 公开,可直接用于后训练。 |