Zilong Wang

I am a Research Scientist at Google DeepMind, where I work on agentic coding and reinforcement learning for Gemini. I was an early member of the Gemini Cyber effort and have contributed to Gemini Flash Cyber, Gemini's frontier models for cybersecurity.

I received my Ph.D. in Computer Science from UC San Diego in 2025, advised by Professor Jingbo Shang, and my B.S. in Computer Science from Peking University in 2020, where I was fortunate to be advised by Professor Xiaojun Wan.

I am broadly interested in agentic reinforcement learning, recursive self-improvement, and AI for cybersecurity. If you'd like to discuss research, or just chat, feel free to reach out at .

X /  GitHub  /  Scholar /  LinkedIn

profile photo

Education

Ph.D. Oct. 2020 – Mar. 2025
University of California, San Diego, La Jolla, California
Ph.D. in Computer Science
B.S. Sep. 2016 – Jun. 2020
Peking University, Beijing, China
B.S. in Computer Science

Selected Publications  [full list]

COLM 2026 CocoaBench: Evaluating Unified Digital Agents in the Wild
Shibo Hao, Zhining Zhang, Zhiqi Liang, Tianyang Liu, Yuheng Zha, Qiyue Gao, Jixuan Chen, Zilong Wang, …, Zhiting Hu  /  arXiv / code
ICLR 2026 Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini, …, Zilong Wang, …, Ludwig Schmidt  /  arXiv
TACL 2026 Learning to Optimize Multi-objective Alignment through Dynamic Reward Weighting
Yining Lu, Zilong Wang**, Shiyang Li, Xin Liu, Changlong Yu, Qingyu Yin, Zhan Shi, Zixuan Zhang, Meng Jiang (** corresponding author)  /  arXiv / code
NeurIPS 2025 Training Language Models to Generate Quality Code with Program Analysis Feedback
Feng Yao*, Zilong Wang*, Liyuan Liu, Junxia Cui, Li Zhong, Xiaohan Fu, Haohui Mai, Vish Krishnan, Jianfeng Gao, Jingbo Shang (* equal contribution)  /  arXiv / code
COLM 2025 RRO: LLM Agent Optimization Through Rising Reward Trajectories
Zilong Wang, Jingfeng Yang, Sreyashi Nag, Samarth Varshney, Xianfeng Tang, Haoming Jiang, Jingbo Shang, Sheikh Muhammad Sarwar  /  arXiv
arXiv OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
Zilong Wang, Yuedong Cui, Li Zhong, Zimin Zhang, Da Yin, Bill Yuchen Lin, Jingbo Shang  /  arXiv / code
ACL Findings 2024 Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step
Li Zhong, Zilong Wang, Jingbo Shang  /  arXiv / code / featured: MarkTechPost / talk: BAAI / sota: HumanEval 98.2%
ICLR 2024 Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding
Zilong Wang, Hao Zhang, Chun-Liang Li, Julian Martin Eisenschlos, Vincent Perot, Zifeng Wang, Lesly Miculicich, Yasuhisa Fujii, Jingbo Shang, Chen-Yu Lee, Tomas Pfister  /  arXiv / code / featured: Google Research Blog
AAAI 2024 Can LLM Replace Stack Overflow? A Study on Robustness and Reliability of Large Language Model Code Generation
Li Zhong, Zilong Wang  /  arXiv / code / featured: TheRegister

Experiences

2025 – Present, Google DeepMind, Mountain View, California
Research Scientist. Early member of the Gemini Cyber effort, working on agentic coding, reinforcement learning, and cybersecurity. See Gemini Flash Cyber and CodeMender.
2024 – 2025, Amazon, Palo Alto, California
Research Intern, then Applied Scientist. Worked on agentic RL for long-horizon tasks and multi-objective alignment for LLMs. See RRO and Dynamic Reward Weighting.
2022 – 2024, Google Research, Mountain View & Sunnyvale, California
Student Researcher. Worked on table understanding agents and multimodal LMs for document AI. See Chain-of-Table and LMDX.
2021, Adobe Research, San Jose, California
Research Intern. Worked on multimodal LMs for document image understanding. See MGDoc.
2020 – 2021, Microsoft Research Asia, Beijing, China
Research Intern. Worked on pre-training for reading order detection in document understanding. See LayoutReader.

Last updated: September 2026