Efficient computing backbones represented by CAS-ViT support lightweight recognition, detection, and segmentation pipelines under edge-device constraints.以 CAS-ViT 为代表的高效计算骨干,支撑端侧算力约束下的识别、检测与分割能力,已落地多个嵌入式端应用。
Research results with practical reach面向真实场景的研究成果
The center's recent portfolio includes three subprojects: efficient embedded perception, infrared remote-sensing intelligence, and video-image-text multimodal understanding.中心近期成果包括高效嵌入式感知、视频图像文本多模态理解两个子工作。
Video, image, and text understanding remains a foundational research line, with methods relevant to overseas industry use cases such as long-video reasoning and behavior analysis.视频图像文本多模态理解以基础研究为主,面向海外企业在长视频推理、跨模态对齐与行为分析中的应用需求。
From representation and interaction to evaluation and recovery连接表征、交互、评测与恢复
The video presents a coherent research line across native motion-video-text representation, social dependencies in multi-person motion, intention explanation, long-form multimodal evaluation, and continuous 3D motion recovery under occlusion.视频展示了一条连续的研究路径:原生动作—视频—文本表征、多人动态与社会依赖、意图解释、长时序多模态评测,以及遮挡环境下的连续三维运动恢复。
Native integration of motion, video, and text.原生融合动作、视频与文本表征。
Multi-person motion and social dependencies.多人动态及其社会依赖关系。
Intention explanation and multi-agent analysis.意图解释与多智能体行为分析。
Continuous 3D motion recovery under occlusion.遮挡环境下的连续三维运动恢复。
Research question核心问题
How can intelligent systems move beyond recognizing isolated actions and instead understand the meaning, intention, interaction, and physical context behind human behavior?智能系统如何从识别孤立动作,进一步理解行为背后的意义、意图、互动关系与物理环境?
Technical path技术路径
The project combines multimodal representation, temporal modeling, social interaction analysis, benchmark evaluation, and robust 3D recovery to form a verifiable research chain.项目结合多模态表征、时序建模、社会交互分析、基准评测与鲁棒三维恢复,形成可验证的研究链条。