Welcome to my personal homepage!
I am a second-year Ph.D. student in Computer Science at the University of Illinois Chicago (UIC), advised by Prof. Philip S. Yu. Before that, I received my M.S. degree in Artificial Intelligence from the Institute of Automation, Chinese Academy of Sciences (CASIA) in July 2024, where I was advised by Prof. Qiang Liu, Prof. Shu Wu and Prof. Liang Wang.
I also interned at BAAI, where I contributed to Aquila-VL-2B, a 2B-parameter vision-language model.
🔍 Current Research Interests
My research develops evaluation methods and personalization systems for LLMs and AI agents, with an emphasis on understanding model behavior and building trustworthy, user-aligned systems. I currently focus on:
- LLM & Agent Evaluation: reliable multi-turn evaluation, user simulators, and benchmark design for real-world deployment.
- Model Behavior & Trustworthy AI: behavioral probes, robustness, explainability, and evaluation beyond self-report metrics.
- Personalization & User Modeling: long-term preferences, memory, user-intention reasoning, and adaptive interaction.
I am open to part-time AI research internships during the academic year and full-time research internships for Summer 2027.
📄 Curriculum Vitae (PDF, updated August 2026)
🔥 News
- 2026.08: 🎉🎉 Our paper “Sycophancy Suppression Can Impair Rational Updating” has been accepted to EMNLP 2026 Findings! [arXiv]
- 2026.06: 📢 Our survey, “Scaling LLM Agent Learning with Data Synthesis,” is available as a preprint. [Paper]
- 2025.09: 📢 Release a novel personality traits evaluation tool, CSI, for assessing LLMs. [GitHub].
- 2024.07: 🎉🎉 A CIKM short paper has been accepted.
- 2024.05: 📢 Our paper “EX-FEVER: A Dataset for Multi-hop Explainable Fact Verification” has been accepted to ACL 2024 Findings!
- 2023.12: 🎉 Our paper “Interpretable Multimodal Out-of-Context Detection with Soft Logic Regularization” has been accepted as an oral presentation at ICASSP 2024!
- 2023.10: 🎉 Our paper “MenatQA: A New Dataset for Testing the Temporal Comprehension and Reasoning Abilities of Large Language Models” has been accepted to EMNLP 2023 Findings!
📝 Publications

Huanhuan Ma, Henry Peng Zou, Chengze Li, Enze Ma, Yunyue Su, Philip S. Yu
Findings of the Association for Computational Linguistics: EMNLP 2026
We separate unsupported yielding (sycophancy) from rational updating (evidence-driven correction), and show that DPO, SFT, and activation steering all suppress sycophancy at a measurable cost to the ability to update — a trade-off we localize to a shared internal substrate.

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits
Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu
Preprint, 2025
We propose Core Sentiment Inventory (CSI), an implicit-association-test-inspired behavioral evaluation framework for more reliable assessment of LLM traits beyond self-report metrics.

Scaling LLM Agent Learning with Data Synthesis: A Comprehensive Survey
Hanrong Zhang, Yankai Chen, Shicheng Fan, Dehai Min, Shaowen Chen, Huanhuan Ma, et al.
Preprint, 2026
We survey how task specifications, trajectories, feedback signals, and environments can be synthesized to support reliable and scalable learning for LLM agents.

InterruptBench: Evaluating LLM Agents Under User Interruptions in Long-Horizon Web Tasks
Henry Peng Zou, Chunyu Miao, Wei-Chieh Huang, Yankai Chen, Yue Zhou, Hanrong Zhang, Yaozu Wu, Liancheng Fang, Zhengyao Gu, Zhen Zhang, Kening Zheng, Fangxin Wang, Yi Nian, Shanghao Li, Wenzhe Fan, Langzhou He, Shicheng Fan, Huanhuan Ma, Dehai Min, Weizhi Zhang, Xue Liu, Philip S. Yu
Preprint, 2026
We formalize three realistic interruption types — addition, revision, and retraction — and build a benchmark from WebArena-Lite to test whether agents can adapt to mid-task intent changes in long-horizon web navigation.

Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
Shuhao Gu, Jialing Zhang, et al., Huanhuan Ma, et al.
arXiv preprint, revised 2025
We introduce a large-scale, high-quality multimodal instruction dataset and a targeted synthetic-data pipeline for scaling vision-language models.

EX-FEVER: A Dataset for Multi-hop Explainable Fact Verification
Huanhuan Ma, Weizhi Xu, Yifan Wei, Liuji Chen, Liang Wang, Qiang Liu, Shu Wu, Liang Wang
Findings of the Association for Computational Linguistics ACL 2024
We introduce a large scale Multi-hop fact checking dataset with textual explanations, which can be used to evaluate the explainability of fact verification models.

Interpretable Multimodal Out-of-Context Detection with Soft Logic Regularization
Huanhuan Ma*, Jinghao Zhang*, Qiang Liu, Shu Wu, Liang Wang
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024
We introduce a novel multimodal out-of-context detection framework with soft logic regularization, which can effectively detect out-of-context information with interpretability.

Yifan Wei, Xiaoyan Yu, Yixuan Weng, Huanhuan Ma, Yuanzhe Zhang, Jun Zhao, Kang Liu
Proceedings of the 33rd ACM International Conference on Information and Knowledge Management: CIKM 2024:
This study investigates the differences between entity and relational knowledge through knowledge editing. Our findings reveal that entity and relational knowledge cannot be directly transferred or mapped to each other.

Yifan Wei, Yisong Su, Huanhuan Ma, Xiaoyan Yu, Fangyu Lei, Yuanzhe Zhang, Jun Zhao, Kang Liu
Findings of the Association for Computational Linguistics: EMNLP 2023
We construct Multiple Sensitive Factors Time QA (MenatQA), which encompasses three temporal factors (scope factor, order factor, counterfactual factor) with total 2,853 samples for evaluating the time comprehension and reasoning abilities of LLMs.

Assessing knowledge editing in language models via relation perspective
Yifan Wei, Xiaoyan Yu, Huanhuan Ma, Fangyu Lei, Yixuan Weng, Ran Song, Kang Liu

Multi-Cause Learning for Diagnosis Prediction
Liping Wang, Qiang Liu, Huanhuan Ma, Shu Wu, Liang Wang
Data Mining and Big Data (DMBD) 2022
🚀 Projects
- CSI: Core Sentiment Inventory: Behavioral evaluation toolkit for probing LLM traits beyond self-report.
- EX-FEVER: Dataset and code for multi-hop explainable fact verification (ACL 2024 Findings).
- Awesome-LLM-based-Evaluators: A curated list of LLM-based evaluators for various NLP tasks.
📖 Educations
-
2025.08 - Present, Ph.D. in Computer Science, University of Illinois Chicago (UIC), expected May 2029. Advisor: Prof. Philip S. Yu.
-
2021.09 - 2024.07, M.S. in Artificial Intelligence, Institute of Automation, Chinese Academy of Sciences. Advisors: Prof. Liang Wang and Prof. Qiang Liu.
-
2016.09 - 2020.07, B.E. in Software Engineering, Zhengzhou University.
💻 Internships
- 2024.07 - 2025.04: Research Intern, BAAI, Beijing, China. Contributed to Aquila-VL-2B, a 2B-parameter vision-language model.
📅 Academic Services
📖 Reviewers
- Conference on Neural Information Processing Systems (NeurIPS), Reviewer (2026)
- International Conference on Learning Representations (ICLR), Reviewer (2025, 2026)
- ACL Rolling Review (ACL ARR), Reviewer (2025, 2026)
- AAAI 2026 Workshop (PerFM), Reviewer
- ACM International Conference on Information and Knowledge Management (CIKM), Reviewer (2024, 2025)