Large Multimodal & Language Models大型多模态与语言模型
We explore unified representations across text, vision, audio, motion, and 3D modalities, with emphasis on cross-modal alignment, complex reasoning, context learning, and interpretable model behavior.探索文本、视觉、音频、运动与三维等模态的统一表征,重点研究跨模态对齐、复杂推理、上下文学习及模型行为解释。