AI for Science
Since 2023, we have explored multimodal approaches to AI for Science. Our early work focused on modelling complex multimodal material systems, including representation learning for crystalline materials and generative modelling, with work published at ICLR 2025, Digital Discovery 2025, and ICLR 2026.
As large language models have advanced, our focus has expanded to joint modelling of proteins, materials, and text, including LLM-based material generation (KDD 2026), protein–text multimodal fusion, and LLM-agent-based protein directed evolution.
MultiModal Learning
Our early research investigated image captioning from the perspectives of interpretability, efficiency, diversity, and knowledge enrichment, including work at EMNLP 2022, the NLPCC 2023 Best Paper Award, ACM MM 2023, and Findings of ACL 2025.
Building on advances in LLMs, we now focus on: (1) Large Multimodal Models (LMMs)—foundation models that understand vision and language, and support generation and reasoning (EMNLP 2024, NAACL 2025, and Findings of CVPR 2026); and (2) Digital Agents / Computer-Use Agents—autonomous agents that perform complex tasks on digital devices (ACL 2024, NeurIPS 2025, and COLM 2026). This line of work has received more than 1,000 citations and over 1 million downloads.