Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
Future Blog Post
Published:
This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.
Blog Post number 4
Published:
This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.
publications
Multi-Modality Fusion Perception and Computing in Autonomous Driving
Published in Journal of Computer Research and Development, 2020
Multi-modality fusion perception and computing in autonomous driving.
Recommended citation: Yanyong Zhang, Sha Zhang, Yu Zhang, Jianmin Ji, Yifan Duan, Y. Huang, Jie Peng, and Y. Zhang. (2020). "Multi-Modality Fusion Perception and Computing in Autonomous Driving." Journal of Computer Research and Development, 57:1781-1799.
TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs through Trajectory Matching
Published in IEEE Transactions on Intelligent Transportation Systems, 2023
Automatic spatio-temporal calibration for roadside LiDARs through trajectory matching.
Recommended citation: Haojie Ren, Sha Zhang, Sugang Li, Yao Li, Xinchen Li, Jianmin Ji, Yu Zhang, and Yanyong Zhang. (2023). "TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs through Trajectory Matching." IEEE Transactions on Intelligent Transportation Systems, 24(11):12549-12559.
HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation
Published in International Journal of Computer Vision, 2024
Unsupervised hybrid-view distillation for point-cloud representation learning.
Recommended citation: Sha Zhang, Jiajun Deng, Lei Bai, Houqiang Li, Wanli Ouyang, and Yanyong Zhang. (2024). "HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation." International Journal of Computer Vision, 132(7):2585-2599.
Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures
Published in IEEE Robotics and Automation Letters, 2024
Multi-stage lane-level integration via double-edge structures.
Recommended citation: Guoliang You, Xiaomeng Chu, Yifan Duan, Wenyu Zhang, Xingchen Li, Sha Zhang, Yao Li, Jianmin Ji, and Yanyong Zhang. (2024). "Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures." IEEE Robotics and Automation Letters.
UniPad: A Universal Pre-Training Paradigm for Autonomous Driving
Published in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024
A universal pre-training paradigm for autonomous driving.
Recommended citation: Honghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu, Haoyi Zhu, Tong He, Shixiang Tang, Hengshuang Zhao, Qibo Qiu, Binbin Lin, et al. (2024). "UniPad: A Universal Pre-Training Paradigm for Autonomous Driving." IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15238-15250.
Agent3D-Zero: An Agent for Zero-Shot 3D Understanding
Published in European Conference on Computer Vision, 2024
An agent-based framework for zero-shot 3D understanding.
Recommended citation: Sha Zhang, Di Huang, Jiajun Deng, Shixiang Tang, Wanli Ouyang, Tong He, and Yanyong Zhang. (2024). "Agent3D-Zero: An Agent for Zero-Shot 3D Understanding." European Conference on Computer Vision, pp. 186-202.
PoiFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest
Published in arXiv preprint arXiv:2403.09212, 2024
Multi-modal 3D object detection via fusion at points of interest.
Recommended citation: Jiajun Deng*, Sha Zhang*, Feras Dayoub, Wanli Ouyang, Yanyong Zhang, and Ian Reid. (2024). "PoiFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest." arXiv preprint arXiv:2403.09212.
LFP: Efficient and Accurate End-to-End Lane-Level Planning via Camera-LiDAR Fusion
Published in arXiv preprint arXiv:2409.14170, 2024
Efficient end-to-end lane-level planning with camera-LiDAR fusion.
Recommended citation: Guoliang You, Xiaomeng Chu, Yifan Duan, Xingchen Li, Sha Zhang, Jianmin Ji, and Yanyong Zhang. (2024). "LFP: Efficient and Accurate End-to-End Lane-Level Planning via Camera-LiDAR Fusion." arXiv preprint arXiv:2409.14170.
PonderV2: Pave the Way for 3D Foundation Model with a Universal Pre-Training Paradigm
Published in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025
A universal pre-training paradigm for 3D foundation models.
Recommended citation: Haoyi Zhu, Honghui Yang, Xiaoyang Wu, Di Huang, Sha Zhang, Xianglong He, Hengshuang Zhao, Chunhua Shen, Yu Qiao, Tong He, et al. (2025). "PonderV2: Pave the Way for 3D Foundation Model with a Universal Pre-Training Paradigm." IEEE Transactions on Pattern Analysis and Machine Intelligence.
Self-Supervised Pre-Training with Combined Datasets for 3D Perception in Autonomous Driving
Published in arXiv preprint arXiv:2504.12709, 2025
Self-supervised pre-training with combined datasets for 3D perception in autonomous driving.
Recommended citation: Shumin Wang, Zhuoran Yang, Lidian Wang, Zhipeng Tang, Heng Li, Lehan Pan, Sha Zhang, Jie Peng, Jianmin Ji, and Yanyong Zhang. (2025). "Self-Supervised Pre-Training with Combined Datasets for 3D Perception in Autonomous Driving." arXiv preprint arXiv:2504.12709.
Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI
Published in arXiv preprint arXiv:2506.19613, 2025
A position paper on integrating cognitive and embodied AI for intelligent science laboratories.
Recommended citation: Sha Zhang, Suorong Yang, Tong Xie, Xiangyuan Xue, Zixuan Hu, Rui Li, Wenxi Qu, Zhenfei Yin, Tianfan Fu, Di Hu, et al. (2025). "Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI." arXiv preprint arXiv:2506.19613.
VLMPlanner: Integrating Visual Language Models with Motion Planning
Published in ACM International Conference on Multimedia, 2025
Integrating visual language models with motion planning.
Recommended citation: Zhipeng Tang, Sha Zhang (corresponding author), Jiajun Deng, Chenjie Wang, Guoliang You, Yuting Huang, Xinrui Lin, and Yanyong Zhang. (2025). "VLMPlanner: Integrating Visual Language Models with Motion Planning." ACM International Conference on Multimedia.
PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum
Published in ACM International Conference on Multimedia, 2025
Open-vocabulary 3D semantic segmentation with partial-to-global curriculum.
Recommended citation: Shiqi Zhang, Sha Zhang, Jiajun Deng, Yedong Shen, Mingxiao Ma, and Yanyong Zhang. (2025). "PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum." ACM International Conference on Multimedia.
ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes
Published in Under review, 2026
Affordance-centric reasoning for fine-grained 3D grounding in cluttered scenes.
Recommended citation: Xinrui Lin, Sha Zhang (corresponding author), Shuming Wang, Jiajun Deng, and Yanyong Zhang. (2026). "ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes." Under review.
LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory
Published in International Conference on Machine Learning, 2026
Protocol-grounded 3D layout generation for interactable and safe laboratories.
Recommended citation: Jianbao Cao, Bohan Feng, Zixuan Hu, Rui Li, Haiyuan Wan, Chenxi Li, Wenzhe Cai, Lei Bai, Wanli Ouyang, Lingyu Duan, Di Huang, Minting Pan, Sha Zhang, Xinzhu Ma, Shixiang Tang, and Dongzhan Zhou. (2026). "LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory." International Conference on Machine Learning.
LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents
Published in Advances in Neural Information Processing Systems, 2026
High-fidelity simulation and hierarchical benchmarking for scientific embodied agents.
Recommended citation: Rui Li, Zixuan Hu, Wenxi Qu, Jinouwen Zhang, Zhenfei Yin, Sha Zhang, Xuantuo Huang, Hanqing Wang, Tai Wang, Jiangmiao Pang, et al. (2026). "LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents." Advances in Neural Information Processing Systems.
LabFix: Closed-Loop Visual Refinement for Long-Horizon Scientific Manipulation
Published in Under review, 2026
Closed-loop visual refinement for long-horizon scientific manipulation.
Recommended citation: Jiancheng He, Sha Zhang (corresponding author), and Wanli Ouyang. (2026). "LabFix: Closed-Loop Visual Refinement for Long-Horizon Scientific Manipulation." Under review.
SciDex: Hierarchical Benchmarking of Scientific Dexterous Manipulation
Published in Under review, 2026
Hierarchical benchmarking for scientific dexterous manipulation.
Recommended citation: Zhipeng Tang, Sihang Chen, Sha Zhang (corresponding author), and Yanyong Zhang. (2026). "SciDex: Hierarchical Benchmarking of Scientific Dexterous Manipulation." Under review.
GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction
Published in arXiv preprint arXiv:2604.04331, 2026
Generation-assisted Gaussian splatting for static scene reconstruction.
Recommended citation: Yedong Shen, Shiqi Zhang, Sha Zhang (corresponding author), Yifan Duan, Xinran Zhang, Wenhao Yu, Lu Zhang, Jiajun Deng, and Yanyong Zhang. (2026). "GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction." arXiv preprint arXiv:2604.04331.
Trans2Occ: Voxel Occupancy Estimation and Grasp for Transparent Objects from Simulation to Reality
Published in arXiv preprint arXiv:2606.01777, 2026
Voxel occupancy estimation and grasping for transparent objects from a single RGB image.
Recommended citation: Yixuan Yang*, Sha Zhang*, Rui Li, Zhenfei Yin, Xinzhu Ma, Yiran Qin, Lei Bai, Xudong Xu, Shilin Shan, Wangmeng Zuo, et al. (2026). "Trans2Occ: Voxel Occupancy Estimation and Grasp for Transparent Objects from Simulation to Reality." arXiv preprint arXiv:2606.01777.
