Towards Driving Agents: Evolution and Outlook of Cognitive Architectures for End-to-end Autonomous Driving in the Era of Large Models
-
摘要: 自动驾驶作为具身智能的关键应用方向, 正经历从模块化结构向端到端框架的深刻变革, 而大模型的爆发为其从“感知智能”向“认知智能”的跃迁提供了关键驱动力. 首先回顾自动驾驶技术的演进历程, 剖析传统模块化结构的局限与端到端范式的兴起, 指出现有的基于监督学习的端到端模型在可解释性与因果推理上的不足. 接着引出大模型尤其是多模态大模型赋能自动驾驶的前沿机制, 系统梳理大语言模型、视觉−语言模型、视觉−语言−动作模型的贡献. 与此同时, 也梳理了新兴的世界模型的发展脉络及其为自动驾驶带来的环境推演能力. 进一步地, 探讨自动驾驶的训练范式与验证体系变革, 分析数据集从几何感知向认知理解的演变, 以及开环评估与闭环验证的方法. 最后, 总结当前技术面临的挑战, 并展望构建数据驱动、认知增强与虚实闭环的一体化自动驾驶系统的未来方向.Abstract: As a pivotal application of embodied intelligence, autonomous driving is shifting from modular to end-to-end frameworks. The surge of large models provides a key driving force for its leap from perceptual intelligence to cognitive intelligence. The paper reviews the evolution of autonomous driving technology, dissects limitations of traditional modular structure and the rise of end-to-end paradigms, and points out the deficiencies in interpretability and causal reasoning of supervised learning-based end-to-end models. Then it introduces how large models, especially multimodal large language models, empower autonomous driving, systematically categorizing contributions of large language models, vision-language models, and vision-language-action models. It also sorts out the development of emerging world models and their capability for environmental reasoning in autonomous driving. Furthermore, it explores changes in training paradigms and validation systems for autonomous driving, analyzing dataset evolution from geometric perception to cognitive understanding, and evaluation methods from open-loop to closed-loop verification. Finally, it summarizes the challenges currently facing the technology and envisions future directions for integrated autonomous driving systems that are data-driven, cognitively enhanced, and based on a virtual-real closed loop.
-
表 1 自动驾驶领域大模型与世界模型应用的对比与分析
Table 1 Comparison and analysis of large models and world model for autonomous driving applications
模型类型 核心优势 主要局限 代表性工作 LLM 核心优势在于其通过海量文本训练获得的、高度抽象的世界知识与常识推理能力, 这在早期为自动驾驶系统注入了“通用智能”的雏形 模型过于依赖文本描述作为环境输入, 对视觉信息理解较弱, 存在“语言世界”与“物理世界”的割裂, 无法直接、精准、安全地操控物理实体 Drive like a human[114]
GPT-Driver[115]
DiLu[117]
Talk2Drive[120]
LanguageMPC[118]VLM 实现了视觉与文本信息的对齐, 使自动驾驶系统拥有了强大的“感知−理解”能力, 实现了对环境的直观、多模态、语义级理解, 是连接感知与认知的关键桥梁 虽然它实现了视觉与语义的跨模态对齐, 但它主要是一个观察者和评论者, 而非一个决策者与执行者. 它缺乏直接与物理环境进行具身交互和闭环控制的能力 DriveGPT4[84]
DriveVLM[86]
DriveMLM[131]
LMDrive[132]
Talk2BEV[125]VLA模型 通过构建从多模态感知到车辆控制的端到端映射, 打破了高层语义认知与底层物理控制之间的异构壁垒, 推动自动驾驶系统向具备感知−决策−执行一体化闭环能力的具身智能体跃迁 是迈向通用具身智能的关键一步, 但在数据、安全、可解释性、算力等方面依旧面临着巨大的挑战, 距离大规模、高可靠性的商用部署仍有距离, 需在系统架构、验证方法与基础设施上取得关键突破 LangCoop[139]
OpenEMMA[141]
AutoVLA[146]
OpenDriveVLA[134]
OmniDrive[154]
CoVLA-Agent[138]WM 被认为是实现L4级及以上全自动驾驶的潜在关键技术路径之一, 其重要性不仅源于技术架构的革新, 更在于它直接指向并试图复现人类驾驶行为的深层认知本质——对复杂交通环境未来演变的动态预判以及在此基础上的多步推理与想象性规划 目前主要处于前沿研究和大规模数据训练阶段, 仍存在高质量多模态数据稀缺与鲁棒性较低等问题, 并在可靠性、实时响应能力以及安全等方面尚未通过严格的车规级验证, 仍需在理论框架、算法设计、系统集成与标准化评价体系等方面完成关键突破 GAIA-2[155]
SimGen[156]
DrivingDiffusion[157]
DriveWM[158]
DriveDreamer[159]
DriveDreamer4D[87]表 2 大模型时代面向感知、认知与推理任务的自动驾驶数据集汇总
Table 2 Summary of autonomous driving datasets for perception, cognition, and reasoning tasks in the era of large models
模型输入 数据集名称 数据来源 数据集任务定位 规模 单模态 CityScapes[176] 真实世界 像素级语义与实例分割 5K个细粒度与20K个粗粒度图像 MetaAD 真实世界 闭环驾驶评测 110K个训练片段和10K个验证片段 SimLingo[142] CARLA 驾驶动作指令跟随与视觉问答 基于1M帧片段的28M个问答对 多模态 nuScenes-QA[189] nuScenes[27] 多模态场景问答 34K个场景下的460K个问答对 NuPrompt[194] nuScenes 目标场景描述与轨迹预测 35.4K个目标−提示词对 DriveGPT4[84] BDD-X[129] 感知与推理问答 56K个视频−文本指令样本 DriveLM-nuScenes[190] nuScenes 驾驶任务图结构问答 4.8K个帧片段下的440K个问答对 DriveLM-CARLA CARLA[71] 驾驶任务图结构问答 180K个帧片段下的3.75M个问答对 Reason2Drive[195] nuScenes、Waymo[181]等 感知与推理问答 600K个问答对 LingoQA[196] 真实世界 驾驶行为问答与解释推理 419K个问答对 OmniReason-nuScenes[197] nuScenes 时空因果推理与问答 1K个约20s的驾驶场景 OmniReason-Bench2Drive Bench2Drive[76] 闭环驾驶因果推理 1K个约150m的驾驶片段 WOMD[198] Waymo 交互式轨迹预测 约104K个真实城市和郊区驾驶场景 LMDrive[132] CARLA 驾驶指令跟随 64K个指令跟随数据片段 Talk2Car[183] nuScenes 自然语言指令的视觉定位与指称 12K条指令与9.2K张图像 CoVLA-Dataset[138] 真实世界 轨迹预测与帧级场景推理描述 10K个视频片段/6M帧片段对应字幕 DATAD[199] CARLA 驾驶员接管注意力与视线预测 在12种接管场景下的600K帧数据 W3DA[200] BDD-A[201]、DADA-2000[202]等 可解释性注意力预测 3.5K个场景中的70K个关键样本 OmniDrive[154] nuScenes 规划、推理与反事实问答 通过对抗性推理生成大规模问答对 Rank2Tell[203] 真实世界 重要性分级与因果推理描述 116个密集推理片段 DriveBench[204] DriveLM VLM驾驶可靠性评测 20.5K个问答对与19.2K帧片段 ImpromptuVLA[145] nuScenes、Waymo等 驾驶VLA任务 80K余个长尾场景片段 Chat-B2D[144] CARLA 闭环驾驶VQA与因果推理 1.1M个问答对 NuGrounding[205] nuScenes 多视角3D视觉定位 2.2M文本提示与34K关键帧 RefAV[206] Argoverse 2[207] 基于语言的时空场景挖掘 10K条自然语言查询 -
[1] 国家发展改革委. 智能汽车创新发展战略[Online], available: https://www.ndrc.gov.cn/xxgk/zcfb/tz/202002/t20200224_1221077.html, 2020-02-24National Development and Reform Commission. Strategy for innovation and development of intelligent vehicles [Online], available: https://www.ndrc.gov.cn/xxgk/zcfb/tz/202002/t20200224_1221077.html, February 24, 2020 [2] 工业和信息化部. 国家车联网产业标准体系建设指南(智能网联汽车)(2023版)[Online], available: https://wap.miit.gov.cn/jgsj/kjs/wjfb/art/2023/art_28a7501f51ae4b408f32f3fb2c49e271.html, 2023-07-26Ministry of Industry and Information Technology. Guidelines for the construction of the national internet of vehicles industry standard system (intelligent connected vehicles) (2023 Version)[Online], available: https://wap.miit.gov.cn/jgsj/kjs/wjfb/art/2023/art_28a7501f51ae4b408f32f3fb2c49e271.html, July 26, 2023 [3] 工业和信息化部. 关于开展智能网联汽车准入和上路通行试点工作的通知[Online], available: https://www.gov.cn/zhengce/zhengceku/202311/content_6915788.htm, 2023-11-17Ministry of Industry and Information Technology. Notice on carrying out pilot work for the admission and road traffic of intelligent connected vehicles[Online], available: https://www.gov.cn/zhengce/zhengceku/202311/content_6915788.htm, November 17, 2023 [4] 工业和信息化部. 关于公布智能网联汽车“车路云一体化”应用试点城市名单的通知[Online], available: https://www.gov.cn/zhengce/zhengceku/202407/content_6965771.htm, 2024-07-01Ministry of Industry and Information Technology. Notice on Announcing the List of Pilot Cities for "Vehicle-Road-Cloud Integration" Application of Intelligent Connected Vehicles [Online], available: https://www.gov.cn/zhengce/zhengceku/202407/content_6965771.htm, July 1, 2024 [5] Li X S, Guo Z Z, Dai X Y, Lin Y L, Jin J C, Zhu F H, et al. Deep imitation learning for traffic signal control and operations based on graph convolutional neural networks. In: Proceedings of the 23rd International Conference on Intelligent Transportation Systems (ITSC). Rhodes, Greece: IEEE, 2020. 1-6 [6] Han S S, Wang X, Zhang J J, Cao D P, Wang F Y. Parallel vehicular networks: A CPSS-based approach via multimodal big data in IoV. IEEE Internet of Things Journal, 2019, 6(1): 1079−1089 doi: 10.1109/JIOT.2018.2867039 [7] 中华人民共和国全国人民代表大会. 中华人民共和国国民经济和社会发展第十四个五年规划和2035年远景目标纲要[Online], available: http://www.gov.cn/xinwen/2021-03/13/content_5592681.htm, 2021-03-13The National People's Congress of the People's Republic of China. Outline of the 14th five-year plan for national economic and social development of the people's republic of china and the long-range objectives through the year 2035[Online], available: http://www.gov.cn/xinwen/2021-03/13/content_5592681.htm, March 13, 2021 [8] Hu Y H, Yang J Z, Chen L, Li K Y, Sima C H, Zhu X Z, et al. Planning-oriented autonomous driving. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, BC, Canada: IEEE, 2023. 17853-17862 [9] Bojarski M, Del Testa D, Dworakowski D, Firner B, Flepp B, Goyal P, et al. End-to-end learning for self-driving cars. arXiv preprint arXiv: 1604.07316, 2016 [10] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, et al. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS). Long Beach, CA, USA: Curran Associates, Inc., 2017. 5998-6008 [11] Brown T, Mann B, Ryder N, Subbiah M, Kaplan J D, Dhariwal P, et al. Language models are few-shot learners. In: Proceedings of the 34th International Conference on Neural Information Processing Systems (NeurIPS). Virtual: Curran Associates, Inc., 2020. 1877-1901 [12] Yang Y X, Han C R, Mao R H, Wang H S, Chen Z W, Yang Y T, et al. Survey of general end-to-end autonomous driving: A unified perspective. Authorea Preprints, 2025 [13] Chib P S, Singh P. Recent advancements in end-to-end autonomous driving using deep learning: A survey. IEEE Transactions on Intelligent Vehicles, 2023, 9(1): 103−118 doi: 10.1109/tiv.2023.3318070 [14] Chen Y Y, Tian D X, Lin C M, Yin H B. Survey of end-to-end autonomous driving systems. Journal of Image and Graphics, 2024, 29(11): 3216−3237 [15] Zhu Y X, Wang S Y, Zhong W Q, Shen N C, Li Y Q, Wang S Q, et al. A survey on large language model-powered autonomous driving. Engineering, 2025 [16] Jiang S C, Huang Z L, Qian K A, Luo Z A, Zhu T Z, Zhong Y, et al. A survey on vision-language-action models for autonomous driving. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Tucson, AZ, USA: IEEE, 2025. 4524-4536 [17] Janai J, Güney F, Behl A, Geiger A. Computer vision for autonomous vehicles: Problems, datasets and state-of-the-art. Foundations and Trends in Computer Graphics and Vision, 2020, 12(1-3): 1−308 doi: 10.1561/0600000079 [18] Li L, Wang F Y. Advanced motion control and sensing for intelligent vehicles. Boston, MA, USA: Springer, 2007 [19] Yurtsever E, Lambert J, Carballo A, Takeda K. A survey of autonomous driving: Common practices and emerging technologies. IEEE Access, 2020, 8: 58443−58469 doi: 10.1109/ACCESS.2020.2983149 [20] Yeong D J, Velasco-Hernandez G, Barry J, Walsh J. Sensor and sensor fusion technology in autonomous vehicles: A review. Sensors, 2021, 21(6): Article No. 2140 doi: 10.3390/s21062140 [21] Dai X R. HybridNet: A fast vehicle detection system for autonomous driving. Signal Processing: Image Communication, 2019, 70: 79−88 doi: 10.1016/j.image.2018.09.002 [22] Huang K L, Shi B T, Li X, Li X, Huang S Y, Li Y K. Multi-modal sensor fusion for auto driving perception: A survey. arXiv preprint arXiv: 2202.02703, 2022 [23] Philion J, Fidler S. Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3D. In: Proceedings of the European Conference on Computer Vision (ECCV). Glasgow, UK: Springer, 2020. 194-210 [24] Li H Y, Sima C H, Dai J F, Wang W H, Lu L W, Wang H J, et al. Delving into the devils of bird’s-eye-view perception: A review, evaluation and recipe. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(4): 2151−2170 doi: 10.1109/TPAMI.2023.3333838 [25] Li Z Q, Wang W H, Li H Y, Xie E, Sima C H, Lu T, et al. BEVFormer: Learning bird’s-eye-view representation from LiDAR-camera via spatiotemporal transformers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(3): 2020−2036 doi: 10.1109/TPAMI.2024.3515454 [26] Liu Z J, Tang H T, Amini A, Yang X Y, Mao H Z, Rus D, et al. BEVFusion: Multi-task multi-sensor fusion with unified bird's-eye-view representation. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). London, UK: IEEE, 2023. 2774-2781, [27] Caesar H, Bankiti V, Lang A H, Vora S, Liong V E, Xu Q, et al. nuScenes: A multimodal dataset for autonomous driving. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, WA, USA: IEEE, 2020. 11621-11631 [28] 朱冰, 张培兴, 赵健, 陈虹, 徐志刚, 赵祥模, 等. 基于场景的自动驾驶汽车虚拟测试研究进展. 中国公路学报, 2019, 32(6): 1−19 doi: 10.19721/j.cnki.1001-7372.2019.06.001Zhu Bing, Zhang Pei-Xing, Zhao Jian, Chen Hong, Xu Zhi-Gang, Zhao Xiang-Mo, et al. Review of scenario-based virtualvalidation methods for automated vehicles. China Journal of Highway and Transport, 2019, 32(6): 1−19 doi: 10.19721/j.cnki.1001-7372.2019.06.001 [29] Mozaffari S, Al-Jarrah O Y, Dianati M, Jennings P, Mouzakitis A. Deep learning-based vehicle behaviour prediction for autonomous driving applications: A review. IEEE Transactions on Intelligent Transportation Systems, 2022, 23(1): 33−47 doi: 10.1109/TITS.2020.3012034 [30] Huang Y J, Du J T, Yang Z Y, Zhou Z W, Zhang L, Chen H. A survey on trajectory-prediction methods for autonomous driving. IEEE Transactions on Intelligent Vehicles, 2022, 7(3): 652−674 doi: 10.1109/TIV.2022.3167103 [31] 池荣虎, 侯忠生, 黄彪. 间歇过程最优迭代学习控制的发展:从基于模型到数据驱动. 自动化学报, 2017, 43(6): 917−932 doi: 10.16383/j.aas.2017.c170086CHI Rong-Hu, HOU Zhong-Sheng, HUANG Biao. Optimal Iterative Learning Control of Batch Processes: From Model-based to Data-driven. ACTA AUTOMATICA SINICA, 2017, 43(6): 917−932 doi: 10.16383/j.aas.2017.c170086 [32] Min K, Kim D, Park J, Huh K. RNN-based path prediction of obstacle vehicles with deep ensemble. IEEE Transactions on Vehicular Technology, 2019, 68(10): 10252−10256 doi: 10.1109/TVT.2019.2933232 [33] Nikhil N, Morris B T. Convolutional neural network for trajectory prediction. In: Proceedings of the European Conference on Computer Vision (ECCV) Workshops. Munich, Germany: Springer, 2018 [34] Zhang K P, Feng X L, Wu L, He Z B. Trajectory prediction for autonomous driving using spatial-temporal graph attention transformer. IEEE Transactions on Intelligent Transportation Systems, 2022, 23(11): 22343−22353 doi: 10.1109/TITS.2022.3164450 [35] Gao J Y, Sun C, Zhao H, Shen Y, Anguelov D, Li C C, et al. VectorNet: Encoding HD maps and agent dynamics from vectorized representation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, WA, USA: IEEE, 2020. 11525-11533 [36] Zhao H, Gao J Y, Lan T, Sun C, Sapp B, Varadarajan B, et al. TNT: Target-driven trajectory prediction. In: Proceedings of the Conference on Robot Learning (CoRL). Cambridge, MA, USA: PMLR, 2021. 895-904 [37] Zu C Y, Yang C, Wang J, Gao W B, Cao D P, Wang F Y. Simulation and field testing of multiple vehicles collision avoidance algorithms. IEEE/CAA Journal of Automatica Sinica, 2020, 7(4): 1045−1063 doi: 10.1109/jas.2020.1003246 [38] Guo Y Q, Yao D Y, Li B, He Z M, Gao H C, Li L. Trajectory planning for an autonomous vehicle in spatially constrained environments. IEEE Transactions on Intelligent Transportation Systems, 2022, 23(10): 18326−18336 doi: 10.1109/TITS.2022.3164548 [39] Guo Y Q, Guo Z L, Wang Y Z, Yao D Y, Li B, Li L. A survey of trajectory planning methods for autonomous driving-Part I: Unstructured scenarios. IEEE Transactions on Intelligent Vehicles, 2024, 9(9): 5407−5434 doi: 10.1109/TIV.2023.3337318 [40] Paden B, Cap M, Yong S Z, Yershov D, Frazzoli E. A survey of motion planning and control techniques for self-driving urban vehicles. IEEE Transactions on Intelligent Vehicles, 2016, 1(1): 33−55 doi: 10.1109/TIV.2016.2578706 [41] Soltani A R, Tawfik H, Goulermas J Y, Fernando T. Path planning in construction sites: Performance evaluation of the Dijkstra, A*, and GA search algorithms. Advanced Engineering Informatics, 2002, 16(4): 291−303 doi: 10.1016/S1474-0346(03)00018-1 [42] Wang X, Tang K, Dai X Y, Xu J T, Xi J H, Ai R, et al. Safety-balanced driving-style aware trajectory planning in intersection scenarios with uncertain environment. IEEE Transactions on Intelligent Vehicles, 2023, 8(4): 2888−2898 doi: 10.1109/TIV.2023.3239903 [43] LaValle S. Rapidly-exploring random trees: A new tool for path planning. Research Report 9811. Ames, IA, USA: Computer Science Department, Iowa State University, 1998 [44] Teng S Y, Hu X M, Deng P, Li B, Li Y C, Ai Y F, et al. Motion planning for autonomous driving: The state of the art and future perspectives. IEEE Transactions on Intelligent Vehicles, 2023, 8(6): 3692−3711 doi: 10.1109/TIV.2023.3274536 [45] Bernhard J, Gieselmann R, Esterle K, Knoll A. Experience-based heuristic search: Robust motion planning with deep Q-learning. In: Proceedings of the 2018 IEEE 21st International Conference on Intelligent Transportation Systems (ITSC). Maui, HI, USA: IEEE, 2018. 3175-3182 [46] Puterman M L. Markov decision processes. Handbooks in Operations Research and Management Science, 1990, 2: 331−434 [47] Codevilla F, Müller M, López A, Koltun V, Dosovitskiy A. End-to-end driving via conditional imitation learning. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). Brisbane, Australia: IEEE, 2018. 4693-4700 [48] 白天翔, 王帅, 沈震, 曹东璞, 郑南宁, 王飞跃. 平行机器人与平行无人系统:框架、结构、过程、平台及其应用. 自动化学报, 2017, 43(2): 161−175 doi: 10.16383/j.aas.2017.y000002Bai Tian-Xiang, Wang Shuai, Shen Zhen, Cao Dong-Pu, Zheng Nan-Ning, Wang Fei-Yue. Parallel robotics and parallel unmanned systems: Framework, structure, process, platform and applications. Acta Automatica Sinica, 2017, 43(2): 161−175 doi: 10.16383/j.aas.2017.y000002 [49] Wang J K, Chi W Z, Li C M, Wang C Q, Meng M Q H. Neural RRT*: Learning-based optimal path planning. IEEE Transactions on Automation Science and Engineering, 2020, 17(4): 1748−1758 doi: 10.1109/TASE.2020.2976560 [50] Ma N, Wang J K, Liu J B, Meng M Q H. Conditional generative adversarial networks for optimal path planning. IEEE Transactions on Cognitive and Developmental Systems, 2022, 14(2): 662−671 doi: 10.1109/TCDS.2021.3063273 [51] Qureshi A H, Miao Y L, Simeonov A, Yip M C. Motion planning networks: Bridging the gap between learning-based and classical motion planners. IEEE Transactions on Robotics, 2021, 37(1): 48−66 doi: 10.1109/TRO.2020.3006716 [52] 陈虹, 宫洵, 胡云峰, 刘奇芳, 高炳钊, 郭洪艳. 汽车控制的研究现状与展望. 自动化学报, 2013, 39(4): 322−346 doi: 10.16383/j.aas.c180136Chen Hong, Gong Xun, Hu Yun-Feng, Liu Qi-Fang, Gao Bing-Zhao, Guo Hong-Yan. Automotive control: The state of the art and perspective. Acta Automatica Sinica, 2013, 39(4): 322−346 doi: 10.16383/j.aas.c180136 [53] 段艳杰, 吕宜生, 张杰, 赵学亮, 王飞跃. 深度学习在控制领域的研究现状与展望. 自动化学报, 2016, 42(5): 643−654 doi: 10.16383/j.aas.2016.c160019Duan Yan-Jie, Lv Yi-Sheng, Zhang Jie, Zhao Xue-Liang, Wang Fei-Yue. Deep learning for control: The state of the art and prospects. Acta Automatica Sinica, 2016, 42(5): 643−654 doi: 10.16383/j.aas.2016.c160019 [54] Liu T, Tian B, Ai Y F, Wang F Y. Parallel reinforcement learning-based energy efficiency improvement for a cyber-physical system. IEEE/CAA Journal of Automatica Sinica, 2020, 7(2): 617−626 doi: 10.1109/jas.2020.1003072 [55] Chen L, Li Y C, Huang C, Xing Y, Tian D X, Li L, et al. Milestones in autonomous driving and intelligent vehicles-Part I: Control, computing system design, communication, HD map, testing, and human behaviors. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2023, 53(9): 5831−5847 doi: 10.1109/TSMC.2023.3276218 [56] Ang K H, Chong G, Li Y. PID control system analysis, design, and technology. IEEE Transactions on Control Systems Technology, 2005, 13(4): 559−576 doi: 10.1109/TCST.2005.847331 [57] Thrun S, Montemerlo M, Dahlkamp H, Stavens D, Aron A, Diebel J, et al. Stanley: The robot that won the DARPA Grand Challenge. Journal of Field Robotics, 2006, 23(9): 661−692 doi: 10.1002/rob.20147 [58] Falcone P, Borrelli F, Asgari J, Tseng H E, Hrovat D. Predictive active steering control for autonomous vehicle systems. IEEE Transactions on Control Systems Technology, 2007, 15(3): 566−580 doi: 10.1109/TCST.2007.894653 [59] Yu X K, Wang H, Teng C L, Sun X Q, Chen L, Cai Y F. DGPR-MPC: Learning-based model predictive controller for autonomous vehicle path following. IET Intelligent Transport Systems, 2023, 17(10): 1992−2003 doi: 10.1049/itr2.12391 [60] Kabzan J, Hewing L, Liniger A, Zeilinger M N. Learning-based model predictive control for autonomous racing. IEEE Robotics and Automation Letters, 2019, 4(4): 3363−3370 doi: 10.1109/LRA.2019.2926677 [61] Zhu Y X, Li Z H, Wang F Y, Li L. Control sequences generation for testing vehicle extreme operating conditions based on latent feature space sampling. IEEE Transactions on Intelligent Vehicles, 2023, 8(4): 2712−2722 doi: 10.1109/TIV.2023.3235732 [62] Tampuu A, Matiisen T, Semikin M, Fishman D, Muhammad N. A survey of end-to-end driving: Architectures and training methods. IEEE Transactions on Neural Networks and Learning Systems, 2020, 33(4): 1364−1384 [63] Singh A. End-to-end autonomous driving using deep learning: A systematic review. arXiv preprint arXiv: 2311.18636, 2023 [64] Chen L, Wu P H, Chitta K, Jaeger B, Geiger A, Li H Y. End-to-end autonomous driving: Challenges and frontiers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 10164−10183 doi: 10.1109/TPAMI.2024.3435937 [65] Xiao Y, Codevilla F, Gurram A, Urfalioglu O, López A M. Multimodal end-to-end autonomous driving. IEEE Transactions on Intelligent Transportation Systems, 2022, 23(1): 537−547 doi: 10.1109/TITS.2020.3013234 [66] Bansal M, Krizhevsky A, Ogale A. ChauffeurNet: Learning to drive by imitating the best and synthesizing the worst. In: Proceedings of Robotics: Science and System. Freiburg im Breisgau, Germany: RSS Foundation, 2019 [67] Wu P H, Jia X S, Chen L, Yan J C, Li H Y, Qiao Y. Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong baseline. In: Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS). New Orleans, LA, USA: Curran Associates, Inc., 2022. 6119-6132 [68] Pomerleau D A. ALVINN: An autonomous land vehicle in a neural network. In: Proceedings of the 3rd International Conference on Neural Information Processing Systems (NeurIPS). Denver, CO, USA: Morgan Kaufmann, 1989. 305-313 [69] Chitta K, Prakash A, Geiger A. NEAT: Neural attention fields for end-to-end autonomous driving. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, QC, Canada: IEEE, 2021. 15793-15803 [70] Chitta K, Prakash A, Jaeger B, Yu Z H, Renz K, Geiger A. TransFuser: Imitation with transformer-based sensor fusion for autonomous driving. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(11): 12878−12895 doi: 10.1109/TPAMI.2022.3200245 [71] Dosovitskiy A, Ros G, Codevilla F, Lopez A, Koltun V. CARLA: An open urban driving simulator. In: Proceedings of the Conference on Robot Learning (CoRL). Mountain View, CA, USA: PMLR, 2017. 78: 1-16 [72] Ye T J, Jing W, Hu C Y, Huang S K, Gao L P, Li F Z, et al. FusionAD: Multi-modality fusion for prediction and planning tasks of autonomous driving. arXiv preprint arXiv: 2308.01006, 2023 [73] Hu S C, Chen L, Wu P H, Li H Y, Yan J C, Tao D C. ST-P3: End-to-end vision-based autonomous driving via spatial-temporal feature learning. In: Proceedings of the European Conference on Computer Vision (ECCV). Tel Aviv, Israel: Springer, 2022. 533-549 [74] Jiang B, Chen S Y, Xu Q, Liao B C, Chen J J, Zhou H L, et al. VAD: Vectorized scene representation for efficient autonomous driving. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 8340-8350 [75] Sun W C, Lin X W, Shi Y N, Zhang C, Wu H R, Zheng S F. SparseDrive: End-to-end autonomous driving via sparse scene representation. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). Atlanta, GA, USA: IEEE, 2025. 8795-8801 [76] Jia X S, Yang Z H, Li Q F, Zhang Z Y, Yan J C. Bench2Drive: Towards multi-ability benchmarking of closed-loop end-to-end autonomous driving. In: Proceedings of the 38th International Conference on Neural Information Processing Systems (NeurIPS). Vancouver, Canada: Curran Associates, Inc., 2024. 819-844 [77] Zhang B Z, Song N, Jin X, Zhang L. Bridging past and future: End-to-end autonomous driving with historical prediction and planning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2025. 6854-6863 [78] Renz K, Chen L, Marcu A M, Hünermann J, Hanotte B, Karnsund A, et al. CarLLaVA: Vision language models for camera-only closed-loop driving. arXiv preprint arXiv: 2406.10165, 2024 [79] Lu Y, Tu J, Ma Y, Zhu X. ReAL-AD: Towards human-like reasoning in end-to-end autonomous driving. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 27783-27793 [80] Wei J, Wang X Z, Schuurmans D, Bosma M, Xia F, Chi E, et al. Chain-of-thought prompting elicits reasoning in large language models. In: Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS). New Orleans, LA, USA: Curran Associates, Inc., 2022. 24824-24837 [81] Jiang B, Chen S Y, Zhang Q, Liu W Y, Wang X G. AlphaDrive: Unleashing the power of VLMs in autonomous driving via reinforcement learning and reasoning. arXiv preprint arXiv: 2503.07608, 2025 [82] Kahneman D. Thinking, Fast and Slow. New York: Farrar, Straus and Giroux, 2011. [83] Jia X S, Wu P H, Chen L, Xie J W, He C H, Yan J C, et al. Think twice before driving: Towards scalable decoders for end-to-end autonomous driving. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, BC, Canada: IEEE, 2023. 21983-21994 [84] Xu Z, Zhang Y, Xie E, Zhao Z, Guo Y, Wong K K, et al. DriveGPT4: Interpretable end-to-end autonomous driving via large language model. IEEE Robotics and Automation Letters, 2024, 9(10): 8186−8193 doi: 10.1109/LRA.2024.3440097 [85] Wei J L, Yuan S S, Li P F, Hu Q D, Gan Z X, Ding W C. OccLLaMA: An occupancy-language-action generative world model for autonomous driving. arXiv preprint arXiv: 2409.03272, 2024 [86] Tian X, Gu J, Li B, Liu Y, Wang Y, Zhao Z, et al. DriveVLM: The convergence of autonomous driving and large vision-language models. In: Proceedings of the 8th Conference on Robot Learning (CoRL). Munich, Germany, PMLR, 2025, 270: 4698-4726 [87] Zhao G S, Ni C J, Wang X F, Zhu Z, Zhang X Y, Wang Y D, et al. DriveDreamer4D: World models are effective data machines for 4D driving scene representation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2025. 12015-12026 [88] Graves A, Wayne G, Reynolds M, Harley T, Danihelka I, Grabska-Barwinska A, et al. Hybrid computing using a neural network with dynamic external memory. Nature, 2016, 538(7626): 471−476 doi: 10.1038/nature20101 [89] 王晓, 张翔宇, 周锐, 田永林, 王建功, 陈龙, 等. 基于平行测试的认知自动驾驶智能架构研究. 自动化学报, 2024, 50(2): 356−371 doi: 10.16383/j.aas.c220820Wang Xiao, Zhang Xiang-Yu, Zhou Rui, Tian Yong-Lin, Wang Jian-Gong, Chen Long, et al. An intelligent architecture for cognitive autonomous driving based on parallel testing. Acta Automatica Sinica, 2024, 50(2): 356−371 doi: 10.16383/j.aas.c220820 [90] Hinton G E, Salakhutdinov R R. Reducing the dimensionality of data with neural networks. Science, 2006, 313(5786): 504−507 doi: 10.1126/science.1127647 [91] Deng J, Dong W, Socher R, Li L J, Li K, Fei-Fei L. ImageNet: A large-scale hierarchical image database. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Miami, FL, USA: IEEE, 2009. 248-255 [92] Radford A, Narasimhan K, Salimans T, Sutskever I. Improving Language Understanding by Generative Pre-Training, Technical Report, OpenAI, USA, 2018 [93] Devlin J, Chang M W, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL). Minneapolis, MN, USA: Association for Computational Linguistics, 2019. 4171-4186 [94] Kaplan J, McCandlish S, Henighan T, Brown T B, Chess B, Child R, et al. Scaling laws for neural language models. arXiv preprint arXiv: 2001.08361, 2020 [95] Touvron H, Lavril T, Izacard G, Martinet X, Lachaux M A, Lacroix T, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv: 2302.13971, 2023 [96] Taylor R, Kardas M, Cucurull G, Scialom T, Hartshorn A, Saravia E, et al. Galactica: A large language model for science. arXiv preprint arXiv: 2211.09085, 2022 [97] Chowdhery A, Narang S, Devlin J, Bosma M, Mishra G, Roberts A, et al. PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research, 2023, 24(240): 1−113 [98] Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, Aleman F L, et al. GPT-4 technical report. arXiv preprint arXiv: 2303.08774, 2023 [99] Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C, Mishkin P, et al. Training language models to follow instructions with human feedback. In: Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS). New Orleans, LA, USA: Curran Associates, Inc., 2022. 27730-27744 [100] Radford A, Kim J W, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning transferable visual models from natural language supervision. In: Proceedings of the 38th International Conference on Machine Learning (ICML). Virtual: PMLR, 2021. 8748-8763 [101] Alayrac J B, Donahue J, Luc P, Miech A, Barr I, Hasson Y, et al. Flamingo: A visual language model for few-shot learning. In: Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS). New Orleans, LA, USA: Curran Associates, Inc., 2022. 23716-23736 [102] Li J N, Li D X, Savarese S, Hoi S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In: Proceedings of the 40th International Conference on Machine Learning (ICML). Honolulu, HI, USA: PMLR, 2023. 202: 19730-19742 [103] Yang Z Y, Li L J, Lin K, Wang J F, Lin C C, Liu Z C, et al. The dawn of LMMs: Preliminary explorations with GPT-4V(ision). arXiv preprint arXiv: 2309.17421, 2023 [104] Gemini Team, Anil R, Borgeaud S, Alayrac J B, Yu J, Soricut R, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv: 2312.11805, 2023 [105] Liu H T, Li C Y, Wu Q Y, Lee Y J. Visual instruction tuning. In: Proceedings of the 37th International Conference on Neural Information Processing Systems (NeurIPS). New Orleans, LA, USA: Curran Associates, Inc., 2023. 34892-34916 [106] Zitkovich B, Yu T H, Xu S C, Xu P, Xiao T, Xia F, et al. RT-2: Vision-language-action models transfer web knowledge to robotic control. In: Proceedings of the 7th Conference on Robot Learning (CoRL). Atlanta, GA, USA: PMLR, 2023. 2165-2183 [107] Reed S, Zolna K, Parisotto E, Colmenarejo S G, Novikov A, Barth-Maron G, et al. A generalist agent. Transactions on Machine Learning Research, 20221−42 [108] Kim M J, Pertsch K, Karamcheti S, Xiao T, Balakrishna A, Nair S, et al. OpenVLA: An open-source vision-language-action model. arXiv preprint arXiv: 2406.09246, 2024 [109] Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv: 2307.09288, 2023 [110] Hu E J, Shen Y L, Wallis P, Allen-Zhu Z, Li Y Z, Wang S A, et al. LoRA: Low-rank adaptation of large language models. In: Proceedings of the International Conference on Learning Representations (ICLR). Virtual Event: OpenReview.net, 2022 [111] Dettmers T, Pagnoni A, Holtzman A, Zettlemoyer L. QLoRA: Efficient finetuning of quantized LLMs. In: Proceedings of the 37th International Conference on Neural Information Processing Systems (NeurIPS). New Orleans, LA, USA: Curran Associates, Inc., 2023. 10088-10115 [112] Huang D w, Yan C, Li Q, Peng X J. From large language models to large multimodal models: A literature review. Applied Sciences, 2024, 14(12): 5068 doi: 10.3390/app14125068 [113] Gao H X, Wang Z R, Li Y Q, Long K W, Yang M, Shen Y Q. A survey for foundation models in autonomous driving. arXiv preprint arXiv: 2402.01105, 2024 [114] Fu D C, Li X, Wen L C, Dou M, Cai P L, Shi B T, et al. Drive like a human: Rethinking autonomous driving with large language models. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Waikoloa, HI, USA: IEEE, 2024. 910-919 [115] Mao J G, Qian Y X, Ye J J, Zhao H, Wang Y. GPT-Driver: Learning to drive with GPT. arXiv preprint arXiv: 2310.01415, 2023 [116] Yao S Y, Zhao J, Yu D, Du N, Shafran I, Narasimhan K, et al. ReAct: Synergizing reasoning and acting in language models. In: Proceedings of the International Conference on Learning Representations (ICLR). Kigali, Rwanda: OpenReview.net, 2023 [117] Wen L C, Fu D C, Li X, Cai X Y, Ma T, Cai P L, et al. DiLu: A knowledge-driven approach to autonomous driving with large language models. In: Proceedings of the International Conference on Learning Representations (ICLR). Vienna, Austria, 2024 [118] Sha H, Mu Y, Jiang Y X, Zhan G J, Chen L, Xu C F, et al. LanguageMPC: Large language models as decision makers for autonomous driving. arXiv preprint arXiv: 2310.03026, 2023 [119] Liao H C, Shen H M, Li Z N, Wang C Y, Li G F, Bie Y M, et al. GPT-4 enhanced multimodal grounding for autonomous driving: Leveraging cross-modal attention with large language models. Communications in Transportation Research, 2024, 4: Article No. 100116 doi: 10.1016/j.commtr.2023.100116 [120] Cui C, Yang Z C, Zhou Y P, Ma Y S, Lu J W, Li L X, et al. Personalized autonomous driving with large language models: Field experiments. In: Proceedings of the 2024 IEEE 27th International Conference on Intelligent Transportation Systems (ITSC). Edmonton, Canada: IEEE, 2024. 20-27 [121] Zhou X C, Liu M Y, Yurtsever E, Zagar B L, Zimmer W, Cao H, et al. Vision language models in autonomous driving: A survey and outlook. IEEE Transactions on Intelligent Vehicles, 2024, 9(1): 2516−2534 doi: 10.1109/tiv.2024.3402136 [122] Kirillov A, Mintun E, Ravi N, Mao H Z, Rolland C, Gustafson L, et al. Segment anything. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 4015-4026 [123] Liu S L, Zeng Z Y, Ren T H, Li F, Zhang H, Yang J, et al. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. In: Proceedings of the European Conference on Computer Vision (ECCV). Milan, Italy: Springer, 2024. 38-55 [124] Ding X, Han J H, Xu H, Zhang W, Li X M. HiLM-D: Enhancing MLLMs with multi-scale high-resolution details for autonomous driving. International Journal of Computer Vision, 2025, 133(8): 5379−5395 doi: 10.1007/s11263-025-02433-3 [125] Choudhary T, Dewangan V, Chandhok S, Priyadarshan S, Jain A, Singh A K, et al. Talk2BEV: Language-enhanced bird’s-eye view maps for autonomous driving. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). Yokohama, Japan: IEEE, 2024. 16345-16352 [126] Yang S Q, Liu J M, Zhang R R, Pan M J, Guo Z Y, Li X Q, et al. LiDAR-LLM: Exploring the potential of large language models for 3D LiDAR understanding. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). Philadelphia, PA, USA: AAAI Press, 2025. 39(9): 9247-9255 [127] Lübberstedt J, Rivera Guerrero E, Uhlemann N, Lienkamp M. V3lma: Visual 3D-enhanced language model for autonomous driving. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2025. 4769-4778 [128] Atakishiyev S, Salameh M, Yao H, Goebel R. Explainable artificial intelligence for autonomous driving: A comprehensive overview and field guide for future research directions. IEEE Access, 2024, 12: 101603−101625 doi: 10.1109/ACCESS.2024.3431437 [129] Kim J, Rohrbach A, Darrell T, Canny J, Akata Z. Textual explanations for self-driving vehicles. In: Proceedings of the European Conference on Computer Vision (ECCV). Munich, Germany: Springer, 2018. 563-578 [130] Shalev-Shwartz Shai, Shammah Shaked, Shashua Amnon. On a formal model of safe and scalable self-driving cars. arXiv preprint arXiv: 1708.06374, 2017 [131] Cui E, Wang W, Li Z, Xie J, Zou H, Deng H, et al. DriveMLM: Aligning multi-modal large language models with behavioral planning states for autonomous driving. Visual Intelligence, Springer, 2025, 3(1): 22 [132] Shao H, Hu Y X, Wang L T, Song G L, Waslander S L, Liu Y, et al. LMDrive: Closed-loop end-to-end driving with large language models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, WA, USA: IEEE, 2024. 15120-15130 [133] Jiang Y F, Gupta A, Zhang Z C, Wang G Z, Dou Y Q, Chen Y J, et al. VIMA: Robot manipulation with multimodal prompts. In: Proceedings of the 40th International Conference on Machine Learning (ICML). Honolulu, HI, USA: PMLR, 2023. 14975-15022 [134] Zhou X C, Han X Y, Yang F, Ma Y P, Tresp V, Knoll A. OpenDriveVLA: Towards end-to-end autonomous driving with large vision language action model. arXiv preprint arXiv: 2503.23463, 2025 [135] Yang Z J, Chai Y L, Jia X S, Li Q F, Shao Y Q, Zhu X K, et al. DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving. arXiv preprint arXiv: 2505.16278, 2025 [136] Zhang J W, Yang X, Wang T Q, Yao Y, Petiushko A, Li B. SafeAuto: Knowledge-enhanced safe autonomous driving with multimodal foundation models. arXiv preprint arXiv: 2503.00211, 2025 [137] Yuan J H, Sun S Y, Omeiza D, Zhao B, Newman P, Kunze L, et al. RAG-Driver: Generalisable driving explanations with retrieval-augmented in-context learning in multi-modal large language model. In: Proceedings of Robotics: Science and Systems (RSS). Delft, Netherlands: RSS Foundation, 2024 [138] Arai H, Miwa K, Sasaki K, Watanabe K, Yamaguchi Y, Aoki S, et al. CoVLA: Comprehensive vision-language-action dataset for autonomous driving. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Tucson, AZ, USA: IEEE, 2025. 1933-1943 [139] Gao X B, Wu Y H, Wang R J, Liu C X, Zhou Y, Tu Z Z. LangCoop: Collaborative driving with language. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2025. 4226-4237 [140] Xu Z H, Bai Y, Zhang Y J, Li Z L, Xia F, Wong K K, et al. DriveGPT4-V2: Harnessing large language model capabilities for enhanced closed-loop autonomous driving. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2025. 17261-17270 [141] Xing S, Qian C Y, Wang Y P, Hua H Y, Tian K X, Zhou Y, et al. OpenEMMA: Open-source multimodal model for end-to-end autonomous driving. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Tucson, AZ, USA: IEEE, 2025. 1001-1009 [142] Renz K, Chen L, Arani E, Sinavski O. SimLingo: Vision-only closed-loop autonomous driving with language-action alignment. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2025. 11993-12003 [143] Wang T H, Maalouf A, Xiao W, Ban Y T, Amini A, Rosman G, et al. Drive Anywhere: Generalizable end-to-end autonomous driving with multi-modal foundation models. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). Yokohama, Japan: IEEE, 2024. 6687-6694 [144] Fu H Y, Zhang D K, Zhao Z C, Cui J F, Liang D K, Zhang C, et al. ORION: A holistic end-to-end autonomous driving framework by vision-language instructed action generation. arXiv preprint arXiv: 2503.19755, 2025 [145] Chi H H, Gao H A, Liu Z M, Liu J N, Liu C Y, Li J W, et al. Impromptu VLA: Open weights and open data for driving vision-language-action models. arXiv preprint arXiv: 2505.23757, 2025 [146] Zhou Z W, Cai T H, Zhao S Z, Zhang Y, Huang Z Y, Zhou B L, et al. AutoVLA: A vision-language-action model for end-to-end autonomous driving with adaptive reasoning and reinforcement fine-tuning. arXiv preprint arXiv: 2506.13757, 2025 [147] Li Y, Tian M, Zhu D C, Zhu J T, Lin Z Y, Xiong Z W, et al. Drive-R1: Bridging reasoning and planning in VLMs for autonomous driving with reinforcement learning. arXiv preprint arXiv: 2506.18234, 2025 [148] Zeng S, Chang X Y, Xie M W, Liu X R, Bai Y F, Pan Z, et al. FutureSightDrive: Thinking visually with spatio-temporal CoT for autonomous driving. arXiv preprint arXiv: 2505.17685, 2025 [149] Jiang T T, Jiang X F, Ma Y, Wen X, Li B L, Zhan K, et al. The better you learn, the smarter you prune: Towards efficient vision-language-action models via differentiable token pruning. arXiv preprint arXiv: 2509.12594, 2025 [150] Chen L H, Hassani H, Nikan S. TS-VLM: Text-guided SoftSort pooling for vision-language models in multi-view driving reasoning. arXiv preprint arXiv: 2505.12670, 2025 [151] Zhou X R, Shan L L, Gui X L. DynRsl-VLM: Enhancing autonomous driving perception with dynamic resolution vision-language models. arXiv preprint arXiv: 2503.11265, 2025 [152] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models. In: Proceedings of the 34th International Conference on Neural Information Processing Systems (NeurIPS). Virtual: Curran Associates, Inc., 2020. 6840-6851 [153] Ha D, Schmidhuber J. Recurrent world models facilitate policy evolution. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems (NeurIPS). Montréal, QC, Canada: Curran Associates, Inc., 2018. 2450-2462 [154] Wang S H, Yu Z D, Jiang X H, Lan S Y, Shi M, Chang N, et al. Omnidrive: A holistic vision-language dataset for autonomous driving with counterfactual reasoning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, TN, USA: IEEE, 2025. 22442-22452 [155] Russell L, Hu A, Bertoni L, Fedoseev G, Shotton J, Arani E, et al. GAIA-2: A controllable multi-view generative world model for autonomous driving. arXiv preprint arXiv: 2503.20523, 2025 [156] Zhou Y S, Simon M, Peng Z H, Mo S C, Zhu H Z, Guo M Y, et al. SimGen: Simulator-conditioned driving scene generation. In: Proceedings of the 38th International Conference on Neural Information Processing Systems (NeurIPS). Vancouver, Canada: Curran Associates, Inc., 2024. 48838-48874 [157] Li X F, Zhang Y F, Ye X Q. DrivingDiffusion: Layout-guided multi-view driving scenarios video generation with latent diffusion model. In: Proceedings of the European Conference on Computer Vision (ECCV). Milan, Italy: Springer, 2024. 469-485 [158] Wang Y Q, He J W, Fan L, Li H X, Chen Y T, Zhang Z X. Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, WA, USA: IEEE, 2024. 14749-14759 [159] Wang X F, Zhu Z, Huang G, Chen X Z, Zhu J G, Lu J W. DriveDreamer: Towards real-world-drive world models for autonomous driving. In: Proceedings of the European Conference on Computer Vision (ECCV). Milan, Italy: Springer, 2024. 55-72 [160] Ding J T, Zhang Y K, Shang Y, Zhang Y H, Zong Z F, Feng J, et al. Understanding world or predicting future? A comprehensive survey of world models. ACM Computing Surveys, 2025, 58(3): 1−38 doi: 10.1145/3746449 [161] Bruce J, Dennis M D, Edwards A, Parker-Holder J, Shi Y G, Hughes E, et al. Genie: Generative interactive environments. In: Proceedings of the 41st International Conference on Machine Learning (ICML). Vienna, Austria: PMLR, 2024 [162] Bardes A, Garrido Q, Ponce J, Chen X L, Rabbat M, LeCun Y, et al. Revisiting feature prediction for learning visual representations from video. arXiv preprint arXiv: 2404.08471, 2024 [163] Liu Y X, Zhang K, Li Y, Yan Z L, Gao C J, Chen R X, et al. Sora: A review on background, technology, limitations, and opportunities of large vision models. arXiv preprint arXiv: 2402.17177, 2024 [164] Hu A, Russell L, Yeo H, Murez Z, Fedoseev G, Kendall A, et al. GAIA-1: A generative world model for autonomous driving. arXiv preprint arXiv: 2309.17080, 2023 [165] Liang D K, Zhang D Y, Zhou X, Tu S F, Feng T R, Li X F, et al. Seeing the future, perceiving the future: A unified driving world model for future generation and perception. arXiv preprint arXiv: 2503.13587, 2025 [166] Hu A, Corrado G, Griffiths N, Murez Z, Gurau C, Yeo H, et al. Model-based imitation learning for urban driving. In: Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS). New Orleans, LA, USA, Curran Associates, Inc., 2022. 20703-20716 [167] Zheng W Z, Chen W L, Huang Y H, Zhang B R, Duan Y Q, Lu J W. OccWorld: Learning a 3D occupancy world model for autonomous driving. In: Proceedings of the European Conference on Computer Vision (ECCV). Milan, Italy: Springer, 2024. 55-72 [168] Zheng W Z, Song R Q, Guo X D, Zhang C M, Chen L. GenAD: Generative end-to-end autonomous driving. In: Proceedings of the European Conference on Computer Vision (ECCV). Milan, Italy: Springer, 2024. 87-104 [169] Fu H Y, Zhang D K, Zhao Z C, Cui J F, Xie H W, Wang B, et al. MindDrive: A vision-language-action model for autonomous driving via online reinforcement learning. arXiv preprint arXiv: 2512.13636, 2025 [170] 杨林瑶, 陈思远, 王晓, 张俊, 王成红. 数字孪生与平行系统: 发展现状、对比及展望. 自动化学报, 2019, 45(11): 2001−2031 doi: 10.16383/j.aas.2019.y000002Yang Lin-Yao, Chen Si-Yuan, Wang Xiao, Zhang Jun, Wang Cheng-Hong. Digital twins and parallel systems: state of the art, comparisons and prospect. Acta Automatica Sinica, 2019, 45(11): 2001−2031 doi: 10.16383/j.aas.2019.y000002 [171] Yang Z, Chen Y, Wang J K, Manivasagam S, Ma W C, Yang A J, et al. UniSim: A neural closed-loop sensor simulator. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, BC, Canada: IEEE, 2023. 1389-1399 [172] Zhou X Y, Lin Z W, Shan X J, Wang Y T, Sun D Q, Yang M H. DrivingGaussian: Composite Gaussian splatting for surrounding dynamic autonomous driving scenes. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, WA, USA: IEEE, 2024. 21634-21643 [173] Gao H, Chen S Y, Jiang B, Liao B C, Shi Y, Guo X Y, et al. RAD: Training an end-to-end driving policy via large-scale 3DGS-based reinforcement learning. arXiv preprint arXiv: 2502.13144, 2025 [174] Zhang Z J, Liniger A, Dai D X, Yu F, Van Gool L. TrafficBots: Towards world models for autonomous driving simulation and motion prediction. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). London, UK: IEEE, 2023: 1522-1529 [175] Geiger A, Lenz P, Urtasun R. Are we ready for autonomous driving? The KITTI vision benchmark suite. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Providence, RI, USA: IEEE, 2012. 3354-3361 [176] Cordts M, Omran M, Ramos S, Rehfeld T, Enzweiler M, Benenson R, et al. The Cityscapes dataset for semantic urban scene understanding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas, NV, USA: IEEE, 2016. 3213-3223 [177] Huang X Y, Cheng X J, Geng Q C, Cao B B, Zhou D F, Wang P, et al. The Apolloscape dataset for autonomous driving. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPR Workshops). Salt Lake City, UT, USA: IEEE, 2018. 954-960 [178] Yu F, Chen H F, Wang X, Xian W Q, Chen Y Y, Liu F C, et al. BDD100K: A diverse driving dataset for heterogeneous multitask learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, WA, USA: IEEE, 2020. 2636-2645 [179] Richter S R, Hayder Z, Koltun V. Playing for benchmarks. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). Venice, Italy: IEEE, 2017. 2213-2222 [180] Ros G, Sellart L, Materzynska J, Vazquez D, Lopez A M. The SYNTHIA dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas, NV, USA: IEEE, 2016. 3234-3243 [181] Sun P, Kretzschmar H, Dotiwalla X, Chouard A, Patnaik V, Tsui P, et al. Scalability in perception for autonomous driving: Waymo open dataset. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA: IEEE, 2020. 2446-2454 [182] Sheeny M, De Pellegrin E, Mukherjee S, Ahrabian A, Wang S, Wallace A. Radiate: A radar dataset for automotive perception in bad weather. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Xi'an, China: IEEE, 2021. 1-7 [183] Deruyttere T, Vandenhende S, Grujicic D, Van Gool L, Moens M F. Talk2Car: Taking control of your self-driving car. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Hong Kong, China: Association for Computational Linguistics, 2019. 2088-2098 [184] Liu M Y, Yurtsever E, Fossaert J, Zhou X C, Zimmer W, Cui Y N, et al. A survey on autonomous driving datasets: Statistics, annotation quality, and a future outlook. IEEE Transactions on Intelligent Vehicles, 2024, 9(11): 7138−7164 doi: 10.1109/TIV.2024.3394735 [185] Krajewski R, Bock J, Kloeker L, Eckstein L. The highD dataset: A drone dataset of naturalistic vehicle trajectories on German highways for validation of highly automated driving systems. In: Proceedings of the 21st International Conference on Intelligent Transportation Systems (ITSC). Maui, HI, USA: IEEE, 2018. 2118-2125 [186] Moers T, Vater L, Krajewski R, Bock J, Zlocki A, Eckstein L. The exiD dataset: A real-world trajectory dataset of highly interactive highway scenarios in Germany. In: Proceedings of the IEEE Intelligent Vehicles Symposium (IV). Aachen, Germany: IEEE, 2022. 958-964 [187] Rasouli A, Kotseruba I, Kunic T, Tsotsos J K. PIE: A large-scale dataset and models for pedestrian intention estimation and trajectory prediction. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Seoul, Korea (South): IEEE, 2019. 6262-6271 [188] Chang M F, Lambert J, Sangkloy P, Singh J, Bak S, Hartnett A, et al. Argoverse: 3D tracking and forecasting with rich maps. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach, CA, USA: IEEE, 2019. 8748-8757 [189] Qian T W, Chen J J, Zhuo L H, Jiao Y, Jiang Y G. NuScenes-QA: A multi-modal visual question answering benchmark for autonomous driving scenario. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). Vancouver, BC, Canada: AAAI Press, 2024. 38(5): 4542-4550 [190] Sima C, Renz K, Chitta K, Chen L, Zhang H, Xie C, et al. DriveLM: Driving with graph visual question answering. In: Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, Springer, 2024, 15103: 256-274 [191] Benekohal R F, Treiterer J. CARSIM: Car-following model for simulation of traffic in normal and stop-and-go conditions. Transportation Research Record, 1988, 1194: 99−111 [192] Hong C J, Aparow V R. System configuration of human-in-the-loop simulation for Level 3 autonomous vehicle using IPG CarMaker. In: Proceedings of the IEEE International Conference on Internet of Things and Intelligence Systems (IoTaIS). Bandung, Indonesia: IEEE, 2021. 215-221 [193] Shah S, Dey D, Lovett C, Kapoor A. AirSim: High-fidelity visual and physical simulation for autonomous vehicles. In: Proceedings of the 11th International Conference on Field and Service Robotics (FSR). Zurich, Switzerland: Springer, 2018. 5: 621-635 [194] Wu D M, Han W C, Liu Y F, Wang T C, Xu C Z, Zhang X Y, et al. Language prompt for autonomous driving. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). Philadelphia, PA, USA: AAAI Press, 2025. 39(8): 8359-8367 [195] Nie M, Peng R Y, Wang C W, Cai X Y, Han J H, Xu H, et al. Reason2Drive: Towards interpretable and chain-based reasoning for autonomous driving. In: Proceedings of the European Conference on Computer Vision (ECCV). Milan, Italy: Springer, 2024. 292-308 [196] Marcu A M, Chen L, Hünermann J, Karnsund A, Hanotte B, Chidananda P, et al. LingoQA: Visual question answering for autonomous driving. In: Proceedings of the European Conference on Computer Vision (ECCV). Milan, Italy: Springer, 2024. 252-269 [197] Liu P, Ning Q T, Lu X Y, Liu H P, Ma W L, She D, et al. OmniReason: A temporal-guided vision-language-action framework for autonomous driving. arXiv preprint arXiv: 2509.00789, 2025 [198] Ettinger S, Cheng S Y, Caine B, Liu C X, Zhao H, Pradhan S, et al. Large scale interactive motion forecasting for autonomous driving: The Waymo Open Motion Dataset. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, QC, Canada: IEEE, 2021. 9710-9719 [199] Huang Y X, Lin Y B, Yue L S S, Yao Z H, Wang J. From gaze to movement: Predicting visual attention for autonomous driving human-machine interaction based on programmatic imitation learning. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Honolulu, HI, USA: IEEE, 2025. 26146-26155 [200] Zhou Y C, Tang J Y, Xiao X Y, Lin Y Y, Liu L K, Guo Z P, et al. Where, What, Why: Towards Explainable Driver Attention Prediction. arXiv preprint arXiv: 2506.23088, 2025 [201] Xia Y, Zhang D Q, Kim J, Nakayama K, Zipser K, Whitney D. Predicting driver attention in critical situations. In: Proceedings of the Asian Conference on Computer Vision (ACCV). Perth, Australia: Springer, 2018. 658-674 [202] Fang J W, Yan D X, Qiao J H, Xue J R, Wang H, Li S. DADA-2000: Can driving accident be predicted by driver attention? Analyzed by a benchmark. In: Proceedings of the IEEE Intelligent Transportation Systems Conference (ITSC). Auckland, New Zealand: IEEE, 2019. 4303-4309 [203] Sachdeva E, Agarwal N, Chundi S, Roelofs S, Li J, Kochenderfer M, et al. Rank2Tell: A multimodal driving dataset for joint importance ranking and reasoning. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). Waikoloa, HI, USA: IEEE, 2024. 7513-7522 [204] Xie S Y, Kong L D, Dong Y H, Sima C H, Zhang W W, Chen Q A, et al. Are VLMs ready for autonomous driving? An empirical study from the reliability, data, and metric perspectives. arXiv preprint arXiv: 2501.04003, 2025 [205] Li F H, Jin H, Gao B, Fan L Y, Jiang L H, Zeng L. NuGrounding: A multi-view 3D visual grounding framework in autonomous driving. arXiv preprint arXiv: 2503.22436, 2025 [206] Davidson C, Ramanan D, Peri N. RefAV: Towards planning-centric scenario mining. arXiv preprint arXiv: 2505.20981, 2025 [207] Wilson B, Qi W, Agarwal T, Lambert J, Singh J, Khandelwal S, et al. Argoverse 2: Next generation datasets for self-driving perception and forecasting. arXiv preprint arXiv: 2301.00493, 2023 [208] Hawke J, Shen R, Gurau C, Sharma S, Reda D, Nikolov N, et al. Urban driving with conditional imitation learning. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). Paris, France: IEEE, 2020. 251-257 [209] Wang Y, Luo W J, Bai J J, Cao Y L, Che T, Chen K, et al. Alpamayo-R1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail. arXiv preprint arXiv: 2511.00088, 2025 [210] Zhang L J, Yuan Y J, Wu C J, Chang X Y, Cai X, Zeng S, et al. MindDriver: Introducing progressive multimodal reasoning for autonomous driving. arXiv preprint arXiv: 2602.21952, 2026 [211] Wang Y F, Wang T X, Yue T. Uncertainty propagation from sensor data to deep learning models in autonomous driving. Information and Software Technology, 2025, 183: 107735 doi: 10.1016/j.infsof.2025.107735 [212] Feng D, Harakeh A, Waslander S L, Dietmayer K. A review and comparative study on probabilistic object detection in autonomous driving. IEEE Transactions on Intelligent Transportation Systems, 2021, 23(8): 9961−9980 doi: 10.1109/tits.2021.3096854 [213] Lu D, Du H, Wu Z, Yang S. Risk assessment in autonomous driving: A comprehensive survey of risk sources, methodologies, and system architectures. Autonomous Intelligent Systems, 2025(1): Article No. 24 doi: 10.1007/s43684-025-00112-1 [214] Sensoy M, Kaplan L, Kandemir M. Evidential deep learning to quantify classification uncertainty. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems (NeurIPS). Montréal, QC, Canada: Curran Associates, Inc., 2018. 3179-3189 [215] Frantar E, Ashkboos S, Hoefler T, Alistarh D. GPTQ: Accurate post-training quantization for generative pre-trained transformers. In: Proceedings of the International Conference on Learning Representations (ICLR). Kigali, Rwanda, 2023 [216] Sun L Y, Jiang J W, Ding Y F, Li F F, Song Y, Zhang H F, et al. Hardware co-design scaling laws via roofline modelling for on-device LLMs. arXiv preprint arXiv: 2602.10377, 2026 [217] 胡云峰, 曲婷, 刘俊, 施竹清, 朱冰, 曹东璞, 等. 智能汽车人机协同控制的研究现状与展望. 自动化学报, 2019, 45(7): 1261−1280 doi: 10.16383/j.aas.c180136Hu Yun-Feng, Qu Ting, Liu Jun, Shi Zhu-Qing, Zhu Bing, Cao Dong-Pu, et al. Human-machine cooperative control of intelli-gent vehicle: recent developments and future Perspectives. Acta Automatica Sinica, 2019, 45(7): 1261−1280 doi: 10.16383/j.aas.c180136 -
计量
- 文章访问数: 4
- HTML全文浏览量: 2
- 被引次数: 0
下载: