I am a Research Scientist at AWS AI in Santa Clara, working on agent post-training, continual learning, and harness optimization. I lead long-horizon agentic reinforcement learning for large-scale coding agent, benchmark and RL environment curation, and design continual learning and recursive self-improvement algorithms for efficient online adaptation at scale.
I completed my Ph.D. in Statistics at UC San Diego, advised by Professor Danna Zhang and Professor Lily Weng. Before that, I received my Master degree in Statistics from the University of Pennsylvania, supervised by Professor Weijie Su.
My research aims to make AI agents more capable and reliable in the real world. I approach this from several angles: learning from experience at scale, optimizing the agent's context and harness, and making models robust to distribution shift and adversarial inputs.
2018 – 2022
Ph.D. in Statistics
University of California San Diego
2016 – 2018
M.S. in Statistics
University of Pennsylvania
2012 – 2016
B.S. in Applied Mathematics
Tongji University
Can an agent keep learning after deployment, directly from real production-style interactions where each interaction usually gives only one trajectory?
We train Qwen3-Coder-30B-A3B using AgentCore Runtime integration to outperform Claude 4.5 Haiku on long-horizon Java repo migrations.
Tuning the agent harness itself — system prompts, tool descriptions, skills, and advisor models — from collected rollout trajectories.
MigrationBench is a large-scale benchmark for repository-level code migration from Java 8 to long-term support versions such as Java 17/21. It also provides an automated and robust framework for evaluating code migration success at the repository level.
To appear in the Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining.
Can an agent keep learning after deployment, directly from real production-style interactions where each interaction usually gives only one trajectory?
We train Qwen3-Coder-30B-A3B using AgentCore Runtime integration to outperform Claude 4.5 Haiku on long-horizon Java repo migrations.
Tuning the agent harness itself — system prompts, tool descriptions, skills, and advisor models — from collected rollout trajectories.
Announcing MigrationBench and Poly-MigrationBench, which extends repository-level migration beyond Java to .NET Framework, Node.js, and Python.
A large-scale, repository-level code migration benchmark from Java 8 to long-term support versions (Java 17/21), with a comprehensive evaluation framework. Ships an RL training recipe and environment on AgentCore for scalable agentic RL training.
Turns Bedrock AgentCore into a scalable rollout engine for agentic reinforcement learning, integrating with major RL frameworks such as veRL, Tinker, and Slime.
A framework for flexible harness optimization via customizable objectives, supporting system prompt optimization, tool description optimization, skill creation, and advisor model optimization.
See also Google Scholar
CLEAR: Context Augmentation from Contrastive Learning of Experience via Agentic Reflection
arXiv preprint arXiv:2604.07487, 2026
LLM Agents Already Know When to Call Tools – Even Without Reasoning
arXiv preprint arXiv:2605.09252, 2026
The Cold-Start Safety Gap in LLM Agents
arXiv preprint arXiv:2606.07867, 2026
QualityFlow: An Agentic Workflow for Program Synthesis Controlled by LLM Quality Checks
arXiv preprint arXiv:2501.17167, 2025
High-dimensional Simultaneous Inference on Non-Gaussian VAR Model via De-biased Estimator
arXiv preprint arXiv:2111.01382, 2021
MigrationBench: Repository-Level Code Migration Benchmark from Java 8
Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2026
Enhancing Language Model Agents using Diversity of Thoughts
The Thirteenth International Conference on Learning Representations (ICLR), 2025
Reasoning and Planning with Large Language Models in Code Development
Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2024
CodeFort: Robust Training for Code Generation Models
Findings of the Association for Computational Linguistics: EMNLP 2024
Promoting Robustness of Randomized Smoothing: Two Cost-Effective Approaches
23rd IEEE International Conference on Data Mining (ICDM), 2023
Robust Multivariate Time-Series Forecasting: Adversarial Attacks and Defense Mechanisms
The Eleventh International Conference on Learning Representations (ICLR), 2023
Statistica Sinica, 35 (2025), 151–170