Zhihui Chen 陈致晖

Ph.D. Student in Artificial Intelligence, National University of Singapore 新加坡国立大学人工智能博士生

I am a PhD student in Artificial Intelligence at the National University of Singapore, advised by Prof. Mengling Feng.

I work on post-training for coding agents, focusing on recursive self-improvement (RSI): the agent generates coding tasks and solves them in a sandbox, trajectories that pass compilation and tests are kept, and the next model is trained on them. The questions I care about are credit assignment over long trajectories, reward and verifier design, and auditing agent actions with a checker independent of the agent. I also apply these methods to medical agents (Med-Banana), and work on detecting forged medical images and LLM-generated text (MedForge, DivScore).

I am currently a research intern on the foundation-model post-training team at Zhipu AI (Z.ai), working on coding-agent post-training. Before that I interned at ByteDance (agentic RL for VLMs; Farsight) and StepFun (pre-training data for StepAudio2.5). I have received a research internship offer from the Qwen team at Alibaba Tongyi Lab, and was selected for Tencent’s Qingyun Program.

新加坡国立大学人工智能博士生,导师为 冯梦凌教授。

研究方向是 Coding Agent + RSI(递归自进化)的后训练:Agent 生成编码任务并在沙箱中求解,通过编译和测试的轨迹被保留,下一版模型在这些轨迹上训练。我关心的问题是长轨迹的信用分配、奖励与验证器设计,以及用独立于 Agent 的检查器审查其行动。同样的方法也用于医疗 Agent(Med-Banana);另一部分工作是医学伪造图像与大模型生成文本的检测(MedForge、DivScore)。

现于智谱(Z.ai)基础模型后训练团队实习,做 Coding Agent 后训练。此前在字节跳动(视觉语言模型的 Agent 强化学习,Farsight)和阶跃星辰(StepAudio2.5 预训练数据)实习。已获阿里通义实验室 Qwen 团队研究实习录用,并入选腾讯「青云计划」。

Zhihui Chen

Research 研究方向

Direction 01 · Coding Agent + RSI 方向一 · Coding Agent + RSI

Post-training coding agents on execution-verified trajectories 基于执行验证轨迹的 Coding Agent 后训练

A coding agent proposes software tasks and solves them in a sandbox. Trajectories that compile, pass the tests and reproduce are kept as SFT and RL data, and the updated model runs the agent in the next round. Each round only helps if the filter is reliable, so most of my work is on the filter and the training signal: credit assignment over long trajectories, reward and verifier design, and auditing agent actions with a checker the agent does not control. Coding Agent 提出软件任务并在沙箱中求解;能编译、通过测试且可复现的轨迹保留为 SFT 与 RL 数据;更新后的模型进入下一轮。每一轮是否有效取决于筛选是否可靠,因此我的工作主要在筛选与训练信号上:长轨迹的信用分配、奖励与验证器设计,以及用 Agent 无法控制的检查器审查其行动。

Agent proposes a coding task Agent 生成编码任务 Sandbox execution 沙箱执行 Compile · test · replay 编译 · 测试 · 复现 Keep only verified trajectories 仅保留通过验证的轨迹 SFT · RL post-training SFT · RL 后训练 Updated model runs the next round 更新后的模型进入下一轮

One round = one model update, trained only on trajectories that passed verification. 一轮对应一次模型更新,训练数据只取通过验证的轨迹。

Open problems 研究问题
  • Long-horizon credit assignment长程任务的信用分配
  • Verifiable rewards (RLVR)可验证奖励(RLVR)
  • Reward & verifier design奖励与验证器设计
  • Independent audit of agent actionsAgent 行动的独立审查
Evidence 证据与代表工作
Work工作 Status状态 What it shows说明
Zhipu AI (Z.ai) · Foundation-Model Post-Training 智谱(Z.ai)· 基础模型后训练 current role 在职 Coding-agent post-training on execution-verified trajectories. 基于执行验证轨迹的 Coding Agent 后训练。
Farsight Farsight ICLR 2026 · first author, under review ICLR 2026 · 一作在投 Future-discounted visual credit assignment for VLM reinforcement learning; trained on 100×H100 at ByteDance. 面向视觉语言模型强化学习的未来折扣视觉信用分配;在字节跳动百卡 H100 上训练。
OfficeBuddy OfficeBuddy open source · 53★ 开源 · 53★ Word/Excel agent. Each edit is re-rendered in real Office and checked by a separate visual verifier. Word/Excel 编辑 Agent。每步编辑在真实 Office 中重新渲染,由独立的视觉验证器检查。
Med-Banana trajectories Med-Banana 轨迹 88K+ released · 150K+ downloads 开源 88K+ · 下载 150K+ Successful and failed editing trajectories, released on GitHub and Hugging Face for post-training. 成功与失败的编辑轨迹,已在 GitHub 与 Hugging Face 公开,可直接用于后训练。
Direction 02 · Medical Agents & Medical AI Safety 方向二 · 医疗 Agent 与医疗 AI 安全

Medical agents and detection of forged medical content 医疗 Agent 与医疗伪造内容检测

In medicine, checking an output often needs a clinician, so feedback is scarce and expensive. Med-Banana applies RSI to medical image editing: a verifier scores each edit, its feedback refines the prompt over rounds, and the successful and failed trajectories are used for post-training. On the safety side, MedForge detects forged medical images and explains where the forgery is, and DivScore detects LLM-generated medical and legal text without training a detector. 医疗场景中检查一个输出往往需要医生,反馈少且贵。Med-Banana 把 RSI 用于医学图像编辑:验证器为每次编辑打分,反馈逐轮修正 Prompt,成功与失败的轨迹用于后训练。安全方面,MedForge 检测伪造的医学影像并指出伪造位置,DivScore 无需训练检测器即可识别医疗与法律领域的大模型生成文本。

Medical task (image / workflow) 医疗任务(影像 / 工作流) Agent acts: multimodal reasoning + tools Agent 执行:多模态推理 + 工具调用 Verifier checks against clinical criteria 验证器按临床标准检查 Feedback → recursive prompt refinement 反馈驱动 Prompt 递归优化 Trajectory-based post-training 成败轨迹后训练 Updated medical agent 更新后的医疗 Agent

Verifier feedback refines the prompt over rounds; successful and failed trajectories are then used for post-training. 验证器反馈逐轮修正 Prompt;成功与失败的轨迹随后用于后训练。

Open problems 研究问题
  • Verifiability of clinical feedback临床反馈的可验证性
  • Multimodal medical reasoning & tool use多模态医学推理与工具调用
  • Medical image integrity & forgery detection医学影像真实性与伪造检测
  • Evidence integrity & rigorous output evaluation证据可信性与输出严格评估
Evidence 证据与代表工作
Work工作 Status状态 What it shows说明
Med-Banana Med-Banana EMNLP 2026 EMNLP 2026 RSI through medical agent self-improvement: verifier-guided prompt refinement and trajectory-based post-training for medical image editing. RSI through Medical Agent Self-Improvement:医学图像编辑中由验证器引导的 Prompt 修正与基于轨迹的后训练。
MedForge MedForge ACL 2026 Main ACL 2026 主会议 Medical deepfake detection that first localizes the forged region, then explains it; 19 lesion types, 10 forgery generators. 医学伪造检测:先定位伪造区域再给出解释;覆盖 19 类病灶、10 种伪造模型。
DivScore DivScore EMNLP 2025 EMNLP 2025 Zero-shot detection of LLM-generated text in medical and legal domains. 医疗与法律领域大模型生成文本的零样本检测。
MiniMax Cowork Team Fellowship MiniMax Cowork Team Fellowship USD 4,500 compute grant 4500 美金算力支持 Compute for medical foundation-model and agent experiments. 用于医疗大模型与 Agent 实验的算力。

News 近期动态

Sep 2026
[Zhipu AI] Joined the foundation-model post-training team at Zhipu AI (Z.ai) as a research intern, working on coding-agent post-training with RSI.
Aug 2026
[Qwen · Tencent] Selected for the Qwen foundation-model research internship at Alibaba Tongyi Lab, and for Tencent's Qingyun (青云计划) Elite Research Talent Program.
Aug 2026
[Tencent Cloud] Received the Tencent Cloud Quant Infrastructure Grant (USD 1,000 cloud credits) for quant-trading agent experiments.
Aug 2026
[EMNLP 2026] Med-Banana is accepted to EMNLP 2026: RSI through medical agent self-improvement, with verifier-guided prompt refinement and trajectory-based post-training.
Aug 2026
[Tencent] Selected for Tencent's Qingyun Program (青云计划), in the game AI × LLM agent track.

Selected Publications 精选论文 All publications → 全部论文 →

2026

Looking Ahead to Stay Grounded: Future-Discounted Visual Credit Assignment for VLM Reinforcement Learning

Zhihui Chen, Yike Yun, et al.

ICLR 2026 submission Under Review First Author

Background 教育与经历

Experience 工作经历
Sep. 2026 - present
Zhipu AI (Z.ai, 智谱)
Foundation Model Post-Training Research Intern
Coding-agent post-training with RSI: the agent generates and executes coding tasks (including quant/finance code), and the model is trained on trajectories that pass execution checks
2026
Alibaba Tongyi Lab · Qwen
Qwen Foundation Model Research Intern — selected
Qwen foundation model team · agentic mid-training & post-training
2026
Tencent · Qingyun Program
Qingyun Program (青云计划) — Elite Research Talent, selected
Elite research talent track · game-AI × LLM agent research
May 2026 - Sep 2026
ByteDance · Data Foundation Model (Data 业务基模)
Foundation Model Research Intern — Reasoning Post-Training
Agentic RL for multimodal models on the Data foundation-model team: tool use, long-horizon RLVR, reward design and evaluation, on 100×H100 · Farsight (first author, ICLR 2026 submission)
Mar. 2025 - Jul. 2025
StepFun (阶跃星辰)
Speech Foundation Model Research Intern — StepAudio2.5 Pre-Training Data
Pre-training data governance for the StepAudio2.5 speech foundation model — audio segmentation & transcription alignment, dialect cleanup (Hakka / Teochew / Cantonese), emotion & scene tagging schemas, and speech–text corpus construction for ASR / TTS / real-time voice
Jan. 2025 - Mar. 2025
AQUMON (Hong Kong)
Quant Trading Agent Intern
LangChain orchestration over Futu OpenAPI real-time data & execution
Feb. 2024 - Apr. 2024
NUS Business School
Teaching Assistant — "Generative AI and LLM"
Curriculum design · tutorials on pre-training, SFT and RLHF
2026年9月 - 至今
智谱(Z.ai)
基础模型后训练 研发实习生
基于 RSI 的 Coding Agent 后训练:Agent 生成并执行编码任务(含量化/金融代码),模型在通过执行检查的轨迹上训练
2026
阿里通义实验室 · Qwen
Qwen 基座模型 研发实习生 —— 已获录用
Qwen 基座模型团队 · Agentic 中期训练与后训练
2026
腾讯 · 青云计划
「青云计划」精英科研人才 入选
精英科研人才计划 · 游戏 AI × 大模型 Agent 方向
2026年5月 - 2026年9月
字节跳动 · Data 业务基模
大模型推理后训练 研究实习生
Data 业务基模团队的多模态 Agent 强化学习:工具调用、长链路 RLVR、奖励设计与评测,百卡 H100 · Farsight(一作,ICLR 2026 在投)
2025年3月 - 2025年7月
阶跃星辰(StepFun)
语音基础模型 研发实习生(StepAudio2.5 预训练数据)
StepAudio2.5 语音基模预训练数据治理——音频切分与转写对齐、方言清洗(客家 / 潮汕 / 粤语)、情绪与场景标签体系、面向 ASR / TTS / 实时语音交互的语音-文本配对语料建设
2025年1月 - 2025年3月
AQUMON(香港)
量化交易 Agent 研发实习生
LangChain Agent 编排 · 对接 Futu OpenAPI 实时行情与自动下单
2024年2月 - 2024年4月
NUS 商学院
教学助理 —《Generative AI and LLM》
课程设计 · 预训练、SFT 与 RLHF 教学
Education 教育经历
Jan. 2025 - present
National University of Singapore
Ph.D. in Artificial Intelligence
Coding Agent + RSI · AI for healthcare
Sep. 2022 - Jul. 2024
The University of Hong Kong
M.Sc. in Artificial Intelligence
Sep. 2018 - May. 2022
The Chinese University of Hong Kong, Shenzhen
B.Sc. in Statistics, Data Science Stream
2025年1月 - 至今
新加坡国立大学
人工智能博士
Coding Agent + RSI(递归自进化)· AI for Healthcare
2022年9月 - 2024年7月
香港大学
人工智能理学硕士
2018年9月 - 2022年5月
香港中文大学(深圳)
统计学理学学士(数据科学方向)
Honors & Awards 荣誉与奖励
  • Qwen Foundation Model Research Internship (Alibaba Tongyi Lab) — selected2026
  • Qingyun Program (青云计划) — Tencent, elite research talent2026
  • Tencent Cloud Quant Infrastructure Grant — USD 1,000 cloud credits2026
  • MiniMax Cowork Team Fellowship, USD 4,500 compute grant2026
  • Full Ph.D. Scholarship, National University of Singapore2025
  • Outstanding College Graduate (Top 5%), CUHK-Shenzhen2022
  • Undergraduate Research Excellence Award, CUHK-Shenzhen2021
  • 阿里通义实验室 Qwen 基座模型研发实习 已获录用2026
  • 腾讯「青云计划」精英科研人才 入选2026
  • 腾讯云量化基础设施资助 —— 1000 美金云算力额度2026
  • MiniMax Cowork Team Fellowship,4500 美金算力支持2026
  • 新加坡国立大学全额博士奖学金2025
  • 香港中文大学(深圳)优秀毕业生(前 5%)2022
  • 香港中文大学(深圳)本科生科研卓越奖2021
Academic Service 学术服务
  • 40+ reviews — ACL · NeurIPS · IJCAI · IEEE TAFFC · ACM TIST2024 - 2026
  • Program Committee, LASS 2026 Workshop @ ACM CIKM 2026 (Rome)2026
  • Session Chair, ACL 2026 Main2026
  • 累计 40+ 篇审稿 —— ACL · NeurIPS · IJCAI · IEEE TAFFC · ACM TIST2024 - 2026
  • ACM CIKM 2026 · LASS 2026 研讨会 程序委员会(PC)委员(罗马)2026
  • ACL 2026 主会议 Session Chair2026

Open Source 开源项目 github.com/richardChenzhihui