Measuring what AI agents can do,
and what they actually do.
I'm an applied research scientist working on AI and agent evaluation. A benchmark score is the model's raw ability; the real performance involves model, harness, tools, environment and policy, which is a system. My work is on building evaluations that survive contact with the systems.
Current threads: agentic evaluations for long-horizon tasks, agent evaluation for self-evolving systems, and eval infrastructure that stays valid as capabilities move. Previously Ph.D. under Jijun Tang and Yan Tong at University of South Carolina. Master's at Tsinghua.
Selected research
Stress-tests activation-based multi-agent collusion monitoring when suspected colluders are opaque. Finds that the original white-box signal is local, trusted honest agents recover a distinct sentinel signal (AUROC 0.806→0.917), and strong ranking still fails to yield a stable low-false-intervention threshold across held-out domains.
A diagnosis-and-decision layer above existing LLM and agent evaluation stacks. EvalMedic uses versioned measurement manifests, controlled counterfactual replay, evaluator-drift checks, and uncertainty-aware abstention to distinguish model regressions from prompt, tool/runtime, data, harness, or grader changes.
Small segmentation models pull tumor and vessel features out of the scan; a large model reasons over them with chain-of-thought prompting and retrieval-augmented generation, so answers stay anchored to trusted cases instead of hallucinated ones.
Pairs explainable attribution with uncertainty quantification, so the classifier reports not just a label but how far that label should be trusted.
A control-function regularizer for deformable registration, measured against gradient and Laplacian smoothness losses on ADNI and IBSR with VoxelMorph.
Writing
Off-hours
I cherish small things in life.



