Research研究方向

Building intelligent systems that understand and act面向理解、推理与行动的智能系统

The center studies the complete path from multimodal representation to cognition, generation, interaction, and real-world application.中心关注从多模态表征到认知理解、内容生成、交互决策与真实场景应用的完整研究链条。

MM

Large Multimodal & Language Models大型多模态与语言模型

We explore unified representations across text, vision, audio, motion, and 3D modalities, with emphasis on cross-modal alignment, complex reasoning, context learning, and interpretable model behavior.探索文本、视觉、音频、运动与三维等模态的统一表征,重点研究跨模态对齐、复杂推理、上下文学习及模型行为解释。

EP

Embodied & Physical Intelligence具身与物理智能

We investigate perception and interaction in physical environments, including human and robotic representation learning, sensorimotor control, motion understanding, physical commonsense, and recovery of 3D behavior under real-world conditions.研究智能体在物理环境中的感知与交互,包括人类及机器人表征学习、感知运动控制、动作理解、物理常识,以及真实复杂环境中的三维行为恢复。

MA

Medical Artificial Intelligence医疗人工智能

We study multimodal representation and reasoning for healthcare, including medical image retrieval and segmentation, cross-modal evidence integration, and cognitive decision support for clinically meaningful applications.研究面向医疗健康的多模态表征与推理,包括医学图像检索与分割、跨模态证据融合,以及服务临床应用的认知决策支持。

3D

Understanding & Generation理解与生成

We develop generative intelligence for high-fidelity and controllable content, covering image editing, dynamic video, 3D scene generation and editing, human motion, and interactive virtual humans.面向高保真、可控制的生成智能,研究图像编辑、动态视频、三维场景生成与编辑、人类动作及交互式虚拟人。

Embodied Intelligence具身智能

Human-centered cross-space cognitive intelligence以人为中心的跨空间认知智能研究全景

Slides converted from the embodied intelligence portfolio, showing the center's work across digital humans, virtual humans, embodied agents, and cross-space cognitive applications.由具身智能 PPT 转换生成,展示中心围绕数字人、虚拟人、具身人以及跨空间认知应用形成的研究布局。

Research Pipeline研究链条

From data to cognitive intelligence从多模态数据走向认知智能

Representation统一表征Text, vision, audio, motion, and 3D signals.融合文本、视觉、音频、运动与三维信号。
Alignment跨模态对齐Connect semantics, geometry, time, and interaction.连接语义、几何、时间与交互关系。
Reasoning推理与生成Understand intention, causality, and physical constraints.理解意图、因果关系与物理约束。
Application场景应用Human behavior, healthcare, embodied agents, and AIGC.服务人类行为、医疗、具身智能与AIGC。
Technical Foundations技术积累

Three subprojects from research to application从研究到应用的三个子工作

A concise overview distilled from the center's technical portfolio in efficient vision, infrared remote sensing, and multimodal understanding.从中心在高效视觉与多模态理解方向的技术积累中,提炼形成清晰可展示的研究与应用链条。

Efficient computing backbone高效计算 Backbone

CAS-ViT introduces convolutional additive self-attention for lightweight, high-throughput perception. The backbone has been applied in multiple SenseTime embedded terminals.以 CAS-ViT 为代表的高效视觉骨干,通过卷积加性自注意力降低端侧推理成本,已应用于商汤科技多个嵌入式终端。

Video-image-text multimodal understanding视频图像文本多模态理解

The center focuses on foundational research in pose-language pretraining, unified embeddings, long-video compression, and multimodal evaluation, serving advanced international industry needs.围绕姿态—语言预训练、统一嵌入、长视频关键帧压缩与多模态评测开展基础研究,面向海外企业对复杂视频图文理解的应用需求。