I am an incoming PhD student at The Chinese University of Hong Kong, advised by Prof. Weiyang Liu, and currently a research intern at ByteDance. I am also fortunate to work closely with Prof. Yandong Wen at Westlake University as a visiting PhD student. Before joining CUHK, I received my B.S. from Peking University in 2025. Previously, I was fortunate to work with Prof. Chaowei Xiao at UW–Madison.
My long-term goal is to build world models that understand physical dynamics well enough to be trusted for embodied decision-making. I work on embodied video understanding, visual representation learning, and evaluation frameworks that make physical reasoning measurable. Previously I worked on multimodal LLM safety.
I am always happy to discuss research. Feel free to reach out by email.
News
- 2026.08 🎉 Our work Your Cursor is Not Secure was accepted to EMNLP 2026 Findings.
- 2026.08 📌 Will be joining CUHK as a PhD student.
- 2026.07 🎉 I will serve as an organizer for the CoRL 2026 workshop/tutorial The Action Gap: Large Generative Models (Video, Image, 3D) Learn Rich Knowledge.
- 2026.07 🎓 Serving as a reviewer for EMNLP 2026.
- 2026.06 🎓 Joined ByteDance as a research intern, working on embodied video understanding.
- 2026.06 🎉 Our work MME-CoF-Pro was accepted to ECCV 2026.
- 2026.05 🎉 Our work BEAR was accepted to ICML 2026.
- 2025.07 🚀 Our work JailBreakV-28K has been integrated into NVIDIA's garak, an LLM vulnerability scanner, enhancing multimodal AI security assessment capabilities.
- 2025.07 🎓 Graduated from Peking University.
- 2025.04 🎉 Our work JailBreakV-28K wins $20,000 SafeBench Prize for Advancing MultiModal Large Language Model Security Benchmarking from the Center for AI Safety.
- 2025.01 🎉 Our work on Vision-Language Model Unlearning was accepted to ICLR 2025.
- 2024.07 🎉 Our work JailBreakV-28K was accepted to COLM 2024.
Publications
* equal contribution. See Google Scholar for the full list.
-
MME-CoF-Pro: Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
Y. Qi*, X. Xu*, Z. Guo*, S. Ma*, R. Zhang, X. Chen, R. An, R. Xing, J. Zhang, et al., L. L. S. Wong.
European Conference on Computer Vision (ECCV), 2026. [arXiv] -
BEAR: Benchmarking and Enhancing Multimodal Language Models for Atomic Embodied Capabilities
Y. Qi*, H. Zhao*, Z. Guo*, S. Ma, Z. Chen, Y. Han, R. Zhang, Z. Lin, S. Xin, et al., L. L. S. Wong.
International Conference on Machine Learning (ICML), 2026. [arXiv] -
Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity Dataset
Y. Ma, J. Wang, F. Wang, S. Ma, J. Li, J. Pan, X. Li, F. Huang, L. Sun, B. Li, et al., C. Xiao.
International Conference on Learning Representations (ICLR), 2025. -
Your Cursor is Not Secure: Command Line Interface Agent Can Expose Realistic Risks Through Tactics, Techniques, and Procedures
W. Luo, Q. Zhang, T. Lu, X. Liu, B. Ma, H.-C. Chiu, S. Ma, et al.
Findings of the Association for Computational Linguistics: EMNLP, 2026. -
JailBreakV-28K: A Benchmark for Assessing the Robustness of Multimodal Large Language Models against Jailbreak Attacks
W. Luo*, S. Ma*, X. Liu*, X. Guo, C. Xiao.
Conference on Language Modeling (COLM), 2024. [Paper] [Project Page] [Code] [Dataset]🏆 $20,000 SafeBench Award from Center for AI Safety
🛡 Integrated into NVIDIA/garak · Microsoft's PyRIT
Preprints
-
Visual-RolePlay: Universal Jailbreak Attack on Multimodal Large Language Models via Role-playing Image Character
S. Ma, W. Luo, Y. Wang, X. Liu.
arXiv:2405.20773, 2024. [arXiv]
Service
- Organizer: CoRL 2026 Workshop/Tutorial, The Action Gap: Large Generative Models (Video, Image, 3D) Learn Rich Knowledge
- Reviewer: EMNLP 2026