Sitemap

A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.

Pages

Posts

Future Blog Post

less than 1 minute read

Published:

This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.

Blog Post number 4

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

publications

Multi-Modality Fusion Perception and Computing in Autonomous Driving

Published in Journal of Computer Research and Development, 2020

Multi-modality fusion perception and computing in autonomous driving.

Recommended citation: Yanyong Zhang, Sha Zhang, Yu Zhang, Jianmin Ji, Yifan Duan, Y. Huang, Jie Peng, and Y. Zhang. (2020). "Multi-Modality Fusion Perception and Computing in Autonomous Driving." Journal of Computer Research and Development, 57:1781-1799.

TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs through Trajectory Matching

Published in IEEE Transactions on Intelligent Transportation Systems, 2023

Automatic spatio-temporal calibration for roadside LiDARs through trajectory matching.

Recommended citation: Haojie Ren, Sha Zhang, Sugang Li, Yao Li, Xinchen Li, Jianmin Ji, Yu Zhang, and Yanyong Zhang. (2023). "TrajMatch: Toward Automatic Spatio-Temporal Calibration for Roadside LiDARs through Trajectory Matching." IEEE Transactions on Intelligent Transportation Systems, 24(11):12549-12559.

HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation

Published in International Journal of Computer Vision, 2024

Unsupervised hybrid-view distillation for point-cloud representation learning.

Recommended citation: Sha Zhang, Jiajun Deng, Lei Bai, Houqiang Li, Wanli Ouyang, and Yanyong Zhang. (2024). "HVDistill: Transferring Knowledge from Images to Point Clouds via Unsupervised Hybrid-View Distillation." International Journal of Computer Vision, 132(7):2585-2599.

Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures

Published in IEEE Robotics and Automation Letters, 2024

Multi-stage lane-level integration via double-edge structures.

Recommended citation: Guoliang You, Xiaomeng Chu, Yifan Duan, Wenyu Zhang, Xingchen Li, Sha Zhang, Yao Li, Jianmin Ji, and Yanyong Zhang. (2024). "Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures." IEEE Robotics and Automation Letters.

UniPad: A Universal Pre-Training Paradigm for Autonomous Driving

Published in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

A universal pre-training paradigm for autonomous driving.

Recommended citation: Honghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu, Haoyi Zhu, Tong He, Shixiang Tang, Hengshuang Zhao, Qibo Qiu, Binbin Lin, et al. (2024). "UniPad: A Universal Pre-Training Paradigm for Autonomous Driving." IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 15238-15250.

Agent3D-Zero: An Agent for Zero-Shot 3D Understanding

Published in European Conference on Computer Vision, 2024

An agent-based framework for zero-shot 3D understanding.

Recommended citation: Sha Zhang, Di Huang, Jiajun Deng, Shixiang Tang, Wanli Ouyang, Tong He, and Yanyong Zhang. (2024). "Agent3D-Zero: An Agent for Zero-Shot 3D Understanding." European Conference on Computer Vision, pp. 186-202.

PoiFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest

Published in arXiv preprint arXiv:2403.09212, 2024

Multi-modal 3D object detection via fusion at points of interest.

Recommended citation: Jiajun Deng*, Sha Zhang*, Feras Dayoub, Wanli Ouyang, Yanyong Zhang, and Ian Reid. (2024). "PoiFusion: Multi-Modal 3D Object Detection via Fusion at Points of Interest." arXiv preprint arXiv:2403.09212.

LFP: Efficient and Accurate End-to-End Lane-Level Planning via Camera-LiDAR Fusion

Published in arXiv preprint arXiv:2409.14170, 2024

Efficient end-to-end lane-level planning with camera-LiDAR fusion.

Recommended citation: Guoliang You, Xiaomeng Chu, Yifan Duan, Xingchen Li, Sha Zhang, Jianmin Ji, and Yanyong Zhang. (2024). "LFP: Efficient and Accurate End-to-End Lane-Level Planning via Camera-LiDAR Fusion." arXiv preprint arXiv:2409.14170.

PonderV2: Pave the Way for 3D Foundation Model with a Universal Pre-Training Paradigm

Published in IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

A universal pre-training paradigm for 3D foundation models.

Recommended citation: Haoyi Zhu, Honghui Yang, Xiaoyang Wu, Di Huang, Sha Zhang, Xianglong He, Hengshuang Zhao, Chunhua Shen, Yu Qiao, Tong He, et al. (2025). "PonderV2: Pave the Way for 3D Foundation Model with a Universal Pre-Training Paradigm." IEEE Transactions on Pattern Analysis and Machine Intelligence.

Self-Supervised Pre-Training with Combined Datasets for 3D Perception in Autonomous Driving

Published in arXiv preprint arXiv:2504.12709, 2025

Self-supervised pre-training with combined datasets for 3D perception in autonomous driving.

Recommended citation: Shumin Wang, Zhuoran Yang, Lidian Wang, Zhipeng Tang, Heng Li, Lehan Pan, Sha Zhang, Jie Peng, Jianmin Ji, and Yanyong Zhang. (2025). "Self-Supervised Pre-Training with Combined Datasets for 3D Perception in Autonomous Driving." arXiv preprint arXiv:2504.12709.

Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI

Published in arXiv preprint arXiv:2506.19613, 2025

A position paper on integrating cognitive and embodied AI for intelligent science laboratories.

Recommended citation: Sha Zhang, Suorong Yang, Tong Xie, Xiangyuan Xue, Zixuan Hu, Rui Li, Wenxi Qu, Zhenfei Yin, Tianfan Fu, Di Hu, et al. (2025). "Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI." arXiv preprint arXiv:2506.19613.

VLMPlanner: Integrating Visual Language Models with Motion Planning

Published in ACM International Conference on Multimedia, 2025

Integrating visual language models with motion planning.

Recommended citation: Zhipeng Tang, Sha Zhang (corresponding author), Jiajun Deng, Chenjie Wang, Guoliang You, Yuting Huang, Xinrui Lin, and Yanyong Zhang. (2025). "VLMPlanner: Integrating Visual Language Models with Motion Planning." ACM International Conference on Multimedia.

PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum

Published in ACM International Conference on Multimedia, 2025

Open-vocabulary 3D semantic segmentation with partial-to-global curriculum.

Recommended citation: Shiqi Zhang, Sha Zhang, Jiajun Deng, Yedong Shen, Mingxiao Ma, and Yanyong Zhang. (2025). "PGOV3D: Open-Vocabulary 3D Semantic Segmentation with Partial-to-Global Curriculum." ACM International Conference on Multimedia.

LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory

Published in International Conference on Machine Learning, 2026

Protocol-grounded 3D layout generation for interactable and safe laboratories.

Recommended citation: Jianbao Cao, Bohan Feng, Zixuan Hu, Rui Li, Haiyuan Wan, Chenxi Li, Wenzhe Cai, Lei Bai, Wanli Ouyang, Lingyu Duan, Di Huang, Minting Pan, Sha Zhang, Xinzhu Ma, Shixiang Tang, and Dongzhan Zhou. (2026). "LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory." International Conference on Machine Learning.

LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents

Published in Advances in Neural Information Processing Systems, 2026

High-fidelity simulation and hierarchical benchmarking for scientific embodied agents.

Recommended citation: Rui Li, Zixuan Hu, Wenxi Qu, Jinouwen Zhang, Zhenfei Yin, Sha Zhang, Xuantuo Huang, Hanqing Wang, Tai Wang, Jiangmiao Pang, et al. (2026). "LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents." Advances in Neural Information Processing Systems.

SciDex: Hierarchical Benchmarking of Scientific Dexterous Manipulation

Published in Under review, 2026

Hierarchical benchmarking for scientific dexterous manipulation.

Recommended citation: Zhipeng Tang, Sihang Chen, Sha Zhang (corresponding author), and Yanyong Zhang. (2026). "SciDex: Hierarchical Benchmarking of Scientific Dexterous Manipulation." Under review.

GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction

Published in arXiv preprint arXiv:2604.04331, 2026

Generation-assisted Gaussian splatting for static scene reconstruction.

Recommended citation: Yedong Shen, Shiqi Zhang, Sha Zhang (corresponding author), Yifan Duan, Xinran Zhang, Wenhao Yu, Lu Zhang, Jiajun Deng, and Yanyong Zhang. (2026). "GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction." arXiv preprint arXiv:2604.04331.

Trans2Occ: Voxel Occupancy Estimation and Grasp for Transparent Objects from Simulation to Reality

Published in arXiv preprint arXiv:2606.01777, 2026

Voxel occupancy estimation and grasping for transparent objects from a single RGB image.

Recommended citation: Yixuan Yang*, Sha Zhang*, Rui Li, Zhenfei Yin, Xinzhu Ma, Yiran Qin, Lei Bai, Xudong Xu, Shilin Shan, Wangmeng Zuo, et al. (2026). "Trans2Occ: Voxel Occupancy Estimation and Grasp for Transparent Objects from Simulation to Reality." arXiv preprint arXiv:2606.01777.