Dr Ruizhe Li BEng, PhD, FHEA

Dr Ruizhe Li

School of Computer Science
Assistant Professor

Contact details

Address
University of Birmingham
Edgbaston
Birmingham
B15 2TT
UK

Dr Ruizhe Li is an Assistant Professor in the School of Computer Science at University of Birmingham. His research focuses on AI safety, interpretability of LLMs and Multimodal LLMs.

For more information, please visit Ruizhe’s personal website.

Qualifications

  • Fellow of the Higher Education Academy, 2025
  • PhD in Computer Science, University of Sheffield, 2021

Biography

Ruizhe Li is an Assistant Professor who is associated with the PLA group (Perception, Language, Action), together with leading the PRISM Lab (Process & Reliability in Interpretability and Safety of Multimodality) within the School of Computer Science, University of Birmingham. He is also associated with the Edinburgh clinical NLP group. He was previously a lecturer at Department of Computing Science, University of Aberdeen. He has also been a postdoctoral research fellow in the Web Intelligence Group, affiliated with the Centre for Artificial Intelligence at University College London (UCL). He received his PhD from the Sheffield NLP group, the Department of Computer Science at University of Sheffield. He completed his bachelor degree BEng Electronic Information & Engineering at Shanghai University, China.

Ruizhe’s research area focuses on AI safety, interpretability of LLMs and Multimodal LLMs. He has consistently published on different top-tier AI, machine learning and NLP conferences, such as ACL, IJCAI, ICML, ICLR, EMNLP, AAAI. He serves as an Area Chair of NeurIPS, ICLR, ACL, EMNLP, IJCAI, NAACL, CIKM, etc. He was awarded INLG 2019 Best paper runner-up, Area Chair Award of ACL 2023, Best Paper Award of COLM 2025 XLLM-Reason-Plan Workshop, Best reviewers of EMNLP 2020 and 2024 and Gold Reviewer of ICML 2026. His research is supported by Google, Cohere, Thinking Machines Lab and The Royal Society.

Postgraduate supervision

  • Natural Language Processing
  • AI Safety
  • Interpretability of LLMs
  • Multimodality LLMs

Research

Ruizhe’s research interests focus on AI safety, interpretability of LLMs and Multimodality LLMs by understanding LLMs internal working mechanism, improving LLMs safety internally across different domains, and applying to multimodality LLMs as well.

Publications

Recent publications

Conference contribution

Li, R, Chen, C, Hu, Y, Gao, Y, Wang, X & Yilmaz, E 2026, Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation. in The Fourteenth International Conference on Learning Representations: ICLR 2026. International Conference on Learning Representations, ICLR, Fourteenth International Conference on Learning Representations, Rio de Janeiro, Brazil, 23/04/26.

Ma, Y, Lu, X, Sang, J, Jiang, X & Li, R 2026, Behind the Scenes: Mechanistic Interpretability of Lora-Adapted Whisper for Speech Emotion Recognition. in ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Proceedings of the ... IEEE International Conference on Acoustics, Speech, and Signal Processing, Institute of Electrical and Electronics Engineers (IEEE), pp. 5286-5290, 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 3/05/26. https://doi.org/10.1109/ICASSP55912.2026.11464049

Yang, C, Xu, R, Li, R, Cao, B & Fan, J 2026, Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs. in M Liakata, VP Moreira, J Zhang & D Jurgens (eds), Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, ACL, pp. 35198-35220, The 64th Annual Meeting of the Association for Computational Linguistics, San Diego, California, United States, 2/07/26. https://doi.org/10.18653/v1/2026.acl-long.1625

Wu, J, Wu, Z, Li, R, Chen, T, Hasan, A, Kim, Y, Cheung, JP-Y, Zhang, T & Wu, H 2026, Error Correction in Radiology Reports: A Knowledge Distillation-Based Multi-Stage Framework. in S Koenig, C Jenkins & ME Taylor (eds), Proceedings of the 40th Annual AAAI Conference on Artificial Intelligence. Proceedings of the AAAI Conference on Artificial Intelligence, no. 46: AAAI-26 Special Track AI for Social Impact II, vol. 40, Association for the Advancement of Artificial Intelligence, pp. 39451-39459, 40th AAAI Conference on Artificial Intelligence, Singapore, Singapore, 20/01/26. https://doi.org/10.1609/aaai.v40i46.41295

Cao, B, Lu, H, Ma, C, Wang, T, Li, R & Fan, J 2026, Orthogonal Hierarchical Decomposition for Structure-Aware Table Understanding with Large Language Models. in Proceedings of the 43rd International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 306, PMLR, 43rd International Conference on Machine Learning, Seoul, Korea, Republic of, 6/07/26.

Yan, L, Li, R, Chen, G, Li, Q, Geng, J, Li, W, Wang, L & Lyu, C 2026, Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs. in Proceedings of the 43rd International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 306, PMLR, 43rd International Conference on Machine Learning, Seoul, Korea, Republic of, 6/07/26.

Wang, X, Sen, P, Li, R & Yilmaz, E 2025, Adaptive Retrieval-Augmented Generation for Conversational Systems. in L Chiruzzo, A Ritter & L Wang (eds), Findings of the Association for Computational Linguistics: NAACL 2025. Association for Computational Linguistics, ACL, pp. 491-503, 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics, Albuquerque, New Mexico, United States, 29/04/25. https://doi.org/10.18653/v1/2025.findings-naacl.30

Li, R & Gao, Y 2025, Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions. in W Che, J Nabende, E Shutova & MT Pilehvar (eds), Findings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics, ACL, pp. 2439-2465, 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, 27/07/25. https://doi.org/10.18653/v1/2025.findings-acl.124

Preprint

Hu, H, Hou, C, Cao, B & Li, R 2026 'Benchmarking Text-to-Python against Text-to-SQL: The Impact of Explicit Logic and Ambiguity' arXiv. https://doi.org/10.48550/arXiv.2601.15728

Sun, C, Cao, B, Li, T, Hou, C, Li, R & Fan, J 2026 'FGTR: Fine-Grained Multi-Table Retrieval via Hierarchical LLM Reasoning' arXiv. https://doi.org/10.48550/arXiv.2603.12702

Pucci, G, Hemendinger, E, Li, R, Abercrombie, G, Dinkar, T & Sinclair, A 2026 'Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback' arXiv. https://doi.org/10.48550/arXiv.2606.02444

Ma, Y, Sang, J, Li, R, Yu, J & Li, A 2026 'From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding' arXiv. https://doi.org/10.48550/arXiv.2607.25355

Jackson, P, Li, R & Edelstein, E 2026 'Nationality encoding in language model hidden states: Probing culturally differentiated representations in persona-conditioned academic text' arXiv. https://doi.org/10.48550/arXiv.2604.10151

Hu, S, Li, R & Gao, Y 2026 'Race, Ethnicity and Their Implication on Bias in Large Language Models' arXiv. https://doi.org/10.48550/arXiv.2601.12868

Yan, L, Li, R, Han, X, Li, W, Wang, B, Wang, L, Lyu, C & Chen, G 2026 'Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback' arXiv. https://doi.org/10.48550/arXiv.2605.17453

View all publications in research portal