Skip to main content
Xiang Li, PhD Image

Xiang Li, PhD, is a Postdoctoral Researcher in the Division of Biostatistics within the Department of Biostatistics, Epidemiology and Informatics at the University of Pennsylvania. He works under the joint supervision of Professors Qi Long and Weijie J. Su.

His research lies at the intersection of statistics, optimization, and machine learning. He develops statistical and algorithmic foundations for reliable generative AI, with a focus on language watermarking, robust detection of AI-generated content, estimation and localization in mixed human-AI text, and rigorous evaluation of large language model capabilities. He also studies uncertainty quantification for learning algorithms, privacy-preserving and communication-efficient federated learning, and online decision-making. His research has appeared in top journals such as the Annals of Statistics, Journal of the Royal Statistical Society Series B, Journal of the American Statistical Association, Operations Research, Journal of Machine Learning Research, and leading machine learning conferences including NeurIPS, ICLR, ICML, and COLT. He received his PhD and BS in Statistics and BA in Economics from Peking University. He will join Rutgers University as an Assistant Professor of Statistics in January 2027.

Education

  • PhD in Statistics, Peking University 2023
  • BS in Statistics, Peking University 2018
  • BA in Economics, Peking University 2018

Contact

Email: lx10077@upenn.edu

🌐 Personal Website
🎓 Google Scholar
💻 GitHub

Research Highlights

My current research broadly studies the mathematical and statistical foundations of artificial intelligence, with a particular focus on generative AI and large language models.

Current projects include:

AI Watermarking and Provenance

Overview: Developing principled methods for watermark design, detection, estimation, and localization, with an emphasis on robustness to human editing and reliable attribution of AI-generated content.

LLM Reasoning

Overview: Investigating the mechanisms and theoretical principles underlying reasoning in large language models, including generalization beyond memorized patterns and the reliability of multi-step reasoning.

AI Evaluation and Measurement

Overview: Developing statistically principled approaches for evaluating model knowledge, reasoning abilities, uncertainty, and capabilities that are not captured by standard benchmarks.

Optimization and Learning Dynamics

Overview: Studying the convergence, efficiency, and statistical behavior of optimization and learning algorithms in stochastic, distributed, and large-scale settings.

Statistical Inference for Learning Algorithms

Overview: Developing methods for uncertainty quantification and valid statistical inference based on the dynamics and randomness of learning algorithms, including stochastic approximation, federated learning, and online decision-making.

Broader Foundations of Reliable AI

Overview: Exploring how data, algorithms, and model behavior interact, with the goal of understanding generalization, robustness, privacy, and the fundamental limits of modern AI systems.

 

More information:
View Research on GitHub

Contact