Measuring what AI agents can do,
and what they actually do.

I'm an applied research scientist working on AI and agent evaluation. A benchmark score is the model's raw ability; the real performance involves model, harness, tools, environment and policy, which is a system. My work is on building evaluations that survive contact with the systems.

Current threads: agentic evaluations for long-horizon tasks, agent evaluation for self-evolving systems, and eval infrastructure that stays valid as capabilities move. Previously Ph.D. under Jijun Tang and Yan Tong at University of South Carolina. Master's at Tsinghua.

Working paper · 2026 · independent evaluation research · code

Stress-tests activation-based multi-agent collusion monitoring when suspected colluders are opaque. Finds that the original white-box signal is local, trusted honest agents recover a distinct sentinel signal (AUROC 0.806→0.917), and strong ranking still fails to yield a stable low-false-intervention threshold across held-out domains.

Open-source system · 2026 · evaluation infrastructure · code

A diagnosis-and-decision layer above existing LLM and agent evaluation stacks. EvalMedic uses versioned measurement manifests, controlled counterfactual replay, evaluator-drift checks, and uncertainty-aware abstention to distinguish model regressions from prompt, tool/runtime, data, harness, or grader changes.

Physics in Medicine & Biology 2025 · with Xuzhou Wu, Guangxin Li, Xing Wang, Kehong Yuan, et al.

Small segmentation models pull tumor and vessel features out of the scan; a large model reasons over them with chain-of-thought prompting and retrieval-augmented generation, so answers stay anchored to trusted cases instead of hallucinated ones.

IPMV 2024 · with Haonan Hu, Desheng Sun, Qiongyu Ye, et al.

Pairs explainable attribution with uncertainty quantification, so the classifier reports not just a label but how far that label should be trusted.

arXiv 2023 · with Dong Xing, Guojun Liao, Yongpei Zhu, et al.

A control-function regularizer for deformable registration, measured against gradient and Laplacian smoothness losses on ADNI and IBSR with VoxelMorph.

reading
Reading Notes, W. Somerset Maugham
learning cooking
exploring cooking

I cherish small things in life.

New York City, 2025
nyc, 2025
Kīlauea, Hawaii, 2025
kīlauea, hawaii, 2025
The first meal home, China, 2026
the first meal home, China, 2026
My parents' yard, Jiangyou, China, 2026
my parents' yard, jiangyou, China, 2026
changelog
v2026.08 my website online. -.-