I am a third-year Ph.D. student in Computer Science at the DAMI Lab, Rensselaer Polytechnic Institute, where I am fortunate to be advised by Prof. Yao Ma. I study reliable multimodal reasoning, particularly how VLMs ground visual evidence, integrate external and internal knowledge, and decide when the evidence is sufficient to answer. I also have expertise in graph learning, graph foundation models, graph self-supervised learning, spatio-temporal forecasting, time-series modeling, and urban computing.
Ph.D. Student in Computer Science
Jan. 2024
Jan. 2028 (expected)
Rensselaer Polytechnic Institute
M.Sc. in Multimedia Information Technology, Distinction
2022
2024
City University of Hong Kong
B.Eng. in Software Engineering - Systems and Technology
2018
2022
University of Electronic Science and Technology of China
I study reliable multimodal reasoning, particularly how VLMs ground visual evidence, integrate external and internal knowledge, and decide when the evidence is sufficient to answer.
I am interested in how foundation models ground non-textual structure, visual entities, and external knowledge. Before moving toward multimodal RAG and KB-VQA, I worked on graph learning, graph self-supervised learning, graph foundation models, spatio-temporal forecasting, time-series modeling, and urban computing.
See my Google Scholar profile for updates.
Conference photos and research-travel snapshots, organized as visual field notes.
Latest: ACL 2026 · San Diego →
Department of Computer Science
Rensselaer Polytechnic Institute
Troy, NY, United States