News
- Aug 2026 NewsOur paper Distilling Photon-Counting CT into Routine Chest CT through Clinically Validated Degradation Modeling was selected as a MICCAI 2026 Spotlight.
Projects
Benchmark-as-Teacher
August 2026
Benchmarks as training infrastructure
- A self-evolving post-training framework that turns stage-level benchmark feedback into targeted data, policy updates, and measurable agent improvement.
AutoMedBench
UC Santa Cruz × NVIDIA · 2026
Towards medical AutoResearch
- A workflow-aware benchmark that puts a base LLM in the driver's seat across the full automated research pipeline: planning, setup, validation, inference, submission.
- Five research categories — segmentation, enhancement, VQA, report, detection. 7 models; 5,400+ runs; two-container isolation.
- Uses LLM-as-judge scoring and curated skill sets to evaluate agentic ability across two scaffolding tiers — Lite (step-by-step guidance) and Standard (open-ended, minimal hints).
- Key finding: over 99% of error codes triggered are engineering-based; a single error code causes an average 48% score drop.
OpenVAE
UCSC × UCSF × Johns Hopkins × NVIDIA · 2025–2026
Scaling latent backbones with worldwide data
- Pretrained 2D and 3D KL-VAE / VQ-VAE backbones for CT and MRI volumes.
- Trained on 400,000+ CTs collected from 145 hospitals worldwide; used as drop-in priors for downstream medical generative work.
Publications
- BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics.
- AutoMedBench: Towards Medical AutoResearch with Agentic AI Models.
- Distilling Photon-Counting CT into Routine Chest CT through Clinically Validated Degradation Modeling.
- See More, Change Less: Anatomy-Aware Diffusion for Contrast Enhancement.
- ShapeKit: Shape-Aware Postprocessing for Organ Segmentation.
- AI-Powered Translation Tool to Enable Contrast-Aware CT Synthesis.
- Towards Robust Out-of-Distribution Generalization via Bayesian Optimization.
Experience
Agent Algorithm Engineer — Alibaba, Qwen
Jul – Sep 2026
Lead projects: BenchAttributer and QwenPaw-RSI
- Built BenchAttributer to separate model errors from agent harness errors. Used its feedback to improve QwenPaw.
- Helped train QwenPaw with multi-turn RL. Used error data to guide sandbox task creation and improve Qwen CoWork-Bench results.
Research Scholar — NVIDIA Research
2026
Lead projects: AutoMedBench, Benchmark-as-Teacher, BiCuRL
- Led AutoMedBench, a benchmark for long medical research tasks. Ran over 6,000 agent tests of planning, reasoning, and recovery.
- Led Benchmark-as-Teacher. Turned long task traces into step-level feedback for agent training.
- Built sandbox skills and tools for AutoResearch. Studied how tool use and memory affect long tasks. Used multi-turn RL to improve research agents.
Research Scholar — Johns Hopkins University, CCVL
Aug 2025 – Mar 2026

Advised by Prof. Alan Yuille and Prof. Zongwei Zhou
- ShapeKit: improved organ segmentation DSC by over 10% on average. Licensed by Johns Hopkins.
- SMILE: used diffusion for CT contrast translation. Improved FID by about 50% and early tumor detection F1 by about 10%.
- OpenVAE: trained on over 400,000 CT scans. Improved cancer detection AUC by 10%.
LLM Pretrain Researcher — Shanghai AI Lab
Oct 2022 – Oct 2023

Advised by Prof. Nanyang Ye, SJTU
- Built the AI4S data pipeline for InternLM-7B. Curated training data and ran pretraining experiments. First-author publication.