| Date | Title | Authors | Code | Comments | |
|---|---|---|---|---|---|
| 2026-6-26 | PLOT: Pseudo-Labeling via Object Tracking for Monocular 3D Object Detection | Seokyeong Lee et.al | paper | code | <summary>detail</summary>ECCV 2026 |
| 2026-5-14 | MonoPRIO: Adaptive Prior Conditioning for Unified Monocular 3D Object Detection | Leon Davies et.al | paper | code | - |
| 2026-4-5 | MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label | Junyoung Jung et.al | paper | code | <summary>detail</summary>CVPR 2026 |
| 2026-3-28 | Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection | Zhihao Zhang et.al | paper | - | <summary>detail</summary>Journal ref:CVPR 2026 |
| 2026-3-27 | Towards Intrinsic-Aware Monocular 3D Object Detection | Zhihao Zhang et.al | paper | - | <summary>detail</summary>This paper is accepted by CVPR 2026 |
| 2026-3-10 | SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection | Yifan Wang et.al | paper | - | <summary>detail</summary>Accepted by CVPR 2026 |
| 2026-3-10 | SpikeSMOKE: Spiking Neural Networks for Monocular 3D Object Detection with Cross-Scale Gated Coding | Xuemei Chen et.al | paper | - | - |
| 2026-3-8 | Selective Transfer Learning of Cross-Modality Distillation for Monocular 3D Object Detection | Rui Ding et.al | paper | - | - |
| 2026-2-24 | Object-Scene-Camera Decomposition and Recomposition for Data-Efficient Monocular 3D Object Detection | Zhaonian Kuang et.al | paper | - | <summary>detail</summary>IJCV |
| 2026-1-2 | Mono3DV: Monocular 3D Object Detection with 3D-Aware Bipartite Matching and Variational Query DeNoising | Kiet Dang Vu et.al | paper | - | - |
| 2025-11-25 | Open Vocabulary Monocular 3D Object Detection | Jin Yao et.al | paper | code | <summary>detail</summary>3DV 2026 |
| 2025-11-17 | Difficulty-Aware Label-Guided Denoising for Monocular 3D Object Detection | Soyul Lee et.al | paper | - | <summary>detail</summary>AAAI 2026 accepted |
| 2025-11-14 | Efficient Feature Aggregation and Scale-Aware Regression for Monocular 3D Object Detection | Yifan Wang et.al | paper | - | - |
| 2025-11-11 | MonoCLUE : Object-Aware Clustering Enhances Monocular 3D Object Detection | Sunghun Yang et.al | paper | - | <summary>detail</summary>AAAI 2026 |
| 2025-11-8 | RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cues for 3D Object Detection | Xiaokai Bai et.al | paper | - | - |
| 2025-9-7 | S-LAM3D: Segmentation-Guided Monocular 3D Object Detection via Feature Space Fusion | Diana-Alexandra Sas et.al | paper | - | - |
| 2025-9-5 | 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection | Yung-Hsu Yang et.al | paper | - | <summary>detail</summary>ICCV 2025 |
| 2025-8-28 | Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time Shifts | Zixuan Hu et.al | paper | - | <summary>detail</summary>Accepted by ICCV 2025 (Highlight) |
| 2025-8-27 | Generalizing Monocular 3D Object Detection | Abhinav Kumar et.al | paper | - | <summary>detail</summary>PhD Thesis submitted to MSU |
| 2025-6-14 | MonoVQD: Monocular 3D Object Detection with Variational Query Denoising and Self-Distillation | Kiet Dang Vu et.al | paper | - | - |
| Date | Title | Authors | Code | Comments | |
|---|---|---|---|---|---|
| 2025-11-10 | Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding | Yuzhen Li et.al | paper | - | - |
| 2025-8-26 | Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding | Yuzhen Li et.al | paper | - | - |
| 2023-12-13 | Mono3DVG: 3D Visual Grounding in Monocular Images | Yang Zhan et.al | paper | code | <summary>detail</summary>Accepted by the Thirty-Eighth AAAI Conference on Artificial Intelligence (AAAI 2024) |
| Date | Title | Authors | Code | Comments | |
|---|---|---|---|---|---|
| 2026-7-14 | VersaQ-3D: Architecture Support for Visual Geometry Grounded Transformers via Versatile Quantization | Yipu Zhang et.al | paper | - | - |
| 2026-7-7 | OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding | Wenyuan Huang et.al | paper | code | <summary>detail</summary>ECCV2026 |
| 2026-7-1 | PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding | Seongmin Jung et.al | paper | - | <summary>detail</summary>ECCV 2026 |
| 2026-6-30 | PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding | Duc Cao Dinh et.al | paper | code | <summary>detail</summary>Preprint |
| 2026-6-29 | UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer | Tianchen Deng et.al | paper | code | <summary>detail</summary>Accepted by ECCV 2026 |
| 2026-6-26 | Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds | Jiaming Bian et.al | paper | - | - |
| 2026-6-18 | Scaling Diverse Language Generation for 3D Visual Grounding | Austin T. Wang et.al | paper | code | - |
| 2026-5-28 | JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments | Zhan Liu et.al | paper | code | <summary>detail</summary>ICML 2026 |
| 2026-5-25 | AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models | Cuong Huynh et.al | paper | code | <summary>detail</summary>Code: https://github |
| 2026-5-20 | SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching | Xuefei Sun et.al | paper | - | - |
| 2026-4-28 | Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding | Yufei Yin et.al | paper | - | - |
| 2026-4-28 | DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework | Yani Zhang et.al | paper | - | <summary>detail</summary>1st place on EmbodiedScan visual grounding |
| 2026-4-2 | Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding | Haibo Wang et.al | paper | - | - |
| 2026-3-31 | MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation | Changli Wu et.al | paper | code | <summary>detail</summary>CVPR 2026 |
| 2026-3-18 | OmniVLN: Omnidirectional 3D Perception and Token-Efficient LLM Reasoning for Visual-Language Navigation across Air and Ground Platforms | Zhongyuang Liu et.al | paper | - | - |
| 2026-3-9 | UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing | Jiaxi Zhang et.al | paper | - | - |
| 2026-2-3 | Z3D: Zero-Shot 3D Visual Grounding from Images | Nikita Drozdov et.al | paper | code | - |
| 2026-1-30 | Learning Geometrically-Grounded 3D Visual Representations for View-Generalizable Robotic Manipulation | Di Zhang et.al | paper | - | - |
| 2026-1-13 | Reasoning Matters for 3D Visual Grounding | Hsiang-Wei Huang et.al | paper | - | <summary>detail</summary>2025 CVPR Workshop on 3D-LLM/VLA: Bridging Language |
| 2025-12-30 | MoniRefer: A Real-world Large-scale Multi-modal Dataset based on Roadside Infrastructure for 3D Visual Grounding | Panquan Yang et.al | paper | - | - |