Home
I am a postdoctoral researcher at MMLab, The Chinese University of Hong Kong, working with Prof. Wanli Ouyang. My research focuses on embodied AI, 3D scene intelligence, multimodal foundation models, and AI for Science. I previously completed my Ph.D. at the University of Science and Technology of China under Prof. Yanyong Zhang, and I also spent time working with Tong He and Prof. Wanli Ouyang at Shanghai AI Laboratory.
I develop perception, reasoning, planning, and evaluation methods for intelligent agents operating in complex scientific and physical environments, with an emphasis on transparent-object perception, long-horizon manipulation, and multi-agent coordination.
You can reach me at zhangsha2048@gmail.com.
News
- Aug 2026: LabDex is officially released as an open benchmark for scientific dexterous manipulation. [Paper] [Dataset]
- 2026: Trans2Occ, LabUtopia, LabBuilder, and several scientific embodied AI projects appeared across arXiv, NeurIPS, and ICML venues.
- 2025: Joined The Chinese University of Hong Kong as a postdoctoral researcher.
- 2024: Presented Agent3D-Zero at ECCV and published HVDistill in IJCV.
Featured Project
- LabDex: an open benchmark for scientific dexterous manipulation, covering real-robot experiments, simulation environments, hierarchical task suites, and an open dataset. Resources: [Paper] [Dataset].
Research Topics
- Scientific embodied intelligence: simulation, evaluation, perception, manipulation, and recovery for long-horizon laboratory procedures.
- Transparent-object perception and manipulation: 3D occupancy estimation, grasping, and collision-aware interaction for challenging lab objects.
- 3D multimodal reasoning and planning: connecting vision-language models with spatial grounding, affordance reasoning, and robot motion planning.
- 3D representation learning: unsupervised hybrid-view distillation and universal pre-training for point clouds, scene understanding, and autonomous driving.
Publications
Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI
Sha Zhang, Suorong Yang, Tong Xie, Xiangyuan Xue, Zixuan Hu, Rui Li, Wenxi Qu, Zhenfei Yin, Tianfan Fu, Di Hu, et al. (2025). "Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI." arXiv preprint arXiv:2506.19613.
Trans2Occ: Voxel Occupancy Estimation and Grasp for Transparent Objects from Simulation to Reality
Yixuan Yang*, Sha Zhang*, Rui Li, Zhenfei Yin, Xinzhu Ma, Yiran Qin, Lei Bai, Xudong Xu, Shilin Shan, Wangmeng Zuo, et al. (2026). "Trans2Occ: Voxel Occupancy Estimation and Grasp for Transparent Objects from Simulation to Reality." arXiv preprint arXiv:2606.01777.
Agent3D-Zero: An Agent for Zero-Shot 3D Understanding
Sha Zhang, Di Huang, Jiajun Deng, Shixiang Tang, Wanli Ouyang, Tong He, and Yanyong Zhang. (2024). "Agent3D-Zero: An Agent for Zero-Shot 3D Understanding." European Conference on Computer Vision, pp. 186-202.
HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation
Sha Zhang, Jiajun Deng, Lei Bai, Houqiang Li, Wanli Ouyang, and Yanyong Zhang. (2024). "HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation." International Journal of Computer Vision, 132(7):2585-2599.
PoiFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest
Jiajun Deng*, Sha Zhang*, Feras Dayoub, Wanli Ouyang, Yanyong Zhang, and Ian Reid. (2024). "PoiFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest." arXiv preprint arXiv:2403.09212.
VLMPlanner: Integrating Visual Language Models with Motion Planning
Zhipeng Tang, Sha Zhang (corresponding author), Jiajun Deng, Chenjie Wang, Guoliang You, Yuting Huang, Xinrui Lin, and Yanyong Zhang. (2025). "VLMPlanner: Integrating Visual Language Models with Motion Planning." ACM International Conference on Multimedia.
GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction
Yedong Shen, Shiqi Zhang, Sha Zhang (corresponding author), Yifan Duan, Xinran Zhang, Wenhao Yu, Lu Zhang, Jiajun Deng, and Yanyong Zhang. (2026). "GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction." arXiv preprint arXiv:2604.04331.
LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories
Zhipeng Tang, Sihang Chen, Sha Zhang, Peihao Yang, Yan Liu, Wentao Zhao, Xinrui Liu, Rui Huang, Wensheng Du, Yuting Huang, Jiajun Deng, Lidian Wang, Yuan Zhang, and Yanyong Zhang. (2026). "LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories." arXiv preprint arXiv:2608.18618.
LabFix: Closed-Loop Visual Refinement for Long-Horizon Scientific Manipulation
Jiancheng He, Sha Zhang (corresponding author), and Wanli Ouyang. (2026). "LabFix: Closed-Loop Visual Refinement for Long-Horizon Scientific Manipulation." Under review.
ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes
Xinrui Lin, Sha Zhang (corresponding author), Shuming Wang, Jiajun Deng, and Yanyong Zhang. (2026). "ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes." Under review.
Multi-Modality Fusion Perception and Computing in Autonomous Driving
Yanyong Zhang, Sha Zhang, Yu Zhang, Jianmin Ji, Yifan Duan, Y. Huang, Jie Peng, and Y. Zhang. (2020). "Multi-Modality Fusion Perception and Computing in Autonomous Driving." Journal of Computer Research and Development, 57:1781-1799.
UniPad: A Universal Pre-Training Paradigm for Autonomous Driving
Honghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu, Haoyi Zhu, Tong He, Shixiang Tang, Hengshuang Zhao, Qibo Qiu, Binbin Lin, et al. (2024). "UniPad: A Universal Pre-Training Paradigm for Autonomous Driving." IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15238-15250.
PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum
Shiqi Zhang, Sha Zhang, Jiajun Deng, Yedong Shen, Mingxiao Ma, and Yanyong Zhang. (2025). "PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum." ACM International Conference on Multimedia.
TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs through Trajectory Matching
Haojie Ren, Sha Zhang, Sugang Li, Yao Li, Xinchen Li, Jianmin Ji, Yu Zhang, and Yanyong Zhang. (2023). "TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs through Trajectory Matching." IEEE Transactions on Intelligent Transportation Systems, 24(11):12549-12559.
PonderV2: Pave the Way for 3D Foundation Model with a Universal Pre-Training Paradigm
Haoyi Zhu, Honghui Yang, Xiaoyang Wu, Di Huang, Sha Zhang, Xianglong He, Hengshuang Zhao, Chunhua Shen, Yu Qiao, Tong He, et al. (2025). "PonderV2: Pave the Way for 3D Foundation Model with a Universal Pre-Training Paradigm." IEEE Transactions on Pattern Analysis and Machine Intelligence.
LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents
Rui Li, Zixuan Hu, Wenxi Qu, Jinouwen Zhang, Zhenfei Yin, Sha Zhang, Xuantuo Huang, Hanqing Wang, Tai Wang, Jiangmiao Pang, et al. (2026). "LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents." Advances in Neural Information Processing Systems.
LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory
Jianbao Cao, Bohan Feng, Zixuan Hu, Rui Li, Haiyuan Wan, Chenxi Li, Wenzhe Cai, Lei Bai, Wanli Ouyang, Lingyu Duan, Di Huang, Minting Pan, Sha Zhang, Xinzhu Ma, Shixiang Tang, and Dongzhan Zhou. (2026). "LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory." International Conference on Machine Learning.
Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures
Guoliang You, Xiaomeng Chu, Yifan Duan, Wenyu Zhang, Xingchen Li, Sha Zhang, Yao Li, Jianmin Ji, and Yanyong Zhang. (2024). "Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures." IEEE Robotics and Automation Letters.
LFP: Efficient and Accurate End-to-End Lane-Level Planning via Camera-LiDAR Fusion
Guoliang You, Xiaomeng Chu, Yifan Duan, Xingchen Li, Sha Zhang, Jianmin Ji, and Yanyong Zhang. (2024). "LFP: Efficient and Accurate End-to-End Lane-Level Planning via Camera-LiDAR Fusion." arXiv preprint arXiv:2409.14170.
Self-Supervised Pre-Training with Combined Datasets for 3D Perception in Autonomous Driving
Shumin Wang, Zhuoran Yang, Lidian Wang, Zhipeng Tang, Heng Li, Lehan Pan, Sha Zhang, Jie Peng, Jianmin Ji, and Yanyong Zhang. (2025). "Self-Supervised Pre-Training with Combined Datasets for 3D Perception in Autonomous Driving." arXiv preprint arXiv:2504.12709.
Service and leadership
- Organizer, Workshop on Scientific Multimodal Agents, World Artificial Intelligence Conference (WAIC) 2026.
- Conference reviewer for CVPR, ECCV, ACM Multimedia (ACM MM), IROS, and ICRA.
- Journal reviewer for IEEE Robotics and Automation Letters (RA-L), Neurocomputing, IEEE Transactions on Intelligent Transportation Systems (T-ITS), and IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI).
