• 中文核心
  • EI
  • 中国科技核心
  • Scopus
  • CSCD
  • 英国科学文摘

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

视听多模态学习综述

宣寒宇 陈强 吴之亮 韩向敏 董文祥 马楠

宣寒宇, 陈强, 吴之亮, 韩向敏, 董文祥, 马楠. 视听多模态学习综述. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250341
引用本文: 宣寒宇, 陈强, 吴之亮, 韩向敏, 董文祥, 马楠. 视听多模态学习综述. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250341
Xuan Han-Yu, Chen Qiang, Wu Zhi-Liang, Han Xiang-Min, Dong Wen-Xiang, Ma Nan. A survey on audio-visual multi-modal learning. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250341
Citation: Xuan Han-Yu, Chen Qiang, Wu Zhi-Liang, Han Xiang-Min, Dong Wen-Xiang, Ma Nan. A survey on audio-visual multi-modal learning. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250341

视听多模态学习综述

doi: 10.16383/j.aas.c250341 cstr: 32138.14.j.aas.c250341
基金项目: 国家自然科学基金(62302006, 62502447, 62371013), 安徽省自然科学基金(2308085QF221), 安徽省高校科研计划项目(2024AH040012)资助
详细信息
    作者简介:

    宣寒宇:安徽大学大数据与统计学院副教授. 主要研究方向为计算机视觉和多模态学习. E-mail: 22176@ahu.edu.cn

    陈强:安徽大学计算机科学与技术学院博士研究生. 主要研究方向为计算机视觉. E-mail: e125111017@stu.ahu.edu.cn

    吴之亮:浙江大学计算机科学与技术学院博士后. 主要研究方向为计算机视觉和多模态技术. E-mail: wu_zhiliang@zju.edu.cn

    韩向敏:清华大学软件学院博士后. 主要研究方向为医学超图计算, 尤其是高阶关联启动的脑网络及病理图像分析. E-mail: simon.xmhan@gmail.com

    董文祥:合肥综合性国家科学中心数据空间研究院研究员. 主要研究方向为网络空间安全, 网络社会认知, 数据智能计算. E-mail: javin0304@foxmail.com

    马楠:北京工业大学信息学院教授. 主要研究方向为交互认知, 具身智能和机器视觉. 本文通信作者. E-mail: manan123@bjut.edu.cn

A Survey on Audio-visual Multi-modal Learning

Funds: Supported by National Natural Science Foundation of China (62302006, 62502447, 62371013), Anhui Natural Science Foundation (2308085QF221), and Anhui Provincial University Scientific Research Project (2024AH040012)
More Information
    Author Bio:

    XUAN Han-Yu Associate professor at the School of Big Data and Statistics, Anhui University. His research interests include computer vision and multi-modal learning

    CHEN Qiang Ph.D. candidate at the School of Computer Science and Technology, Anhui University. His main research interest is computer vision

    WU Zhi-Liang Postdoctor at the School of Computer Science and Technology, Zhejiang University. His research interests include computer vision and multimodal technology

    HAN Xiang-Min Postdoctor at the School of Software, Tsinghua University. His main research interest is medical hypergraph computation, with a particular emphasis on brain networks and pathology image analysis driven by high-order correlations

    DONG Wen-Xiang Researcher at the Institute of Dataspace, Hefei Comprehensive National Science Center. His research interests include cyberspace security, network social cognition, and data intelligent computing

    MA Nan Professor at the School of Information Science and Technology, Beijing University of Technology. Her research interests include interactive cognition, embodied intelligence, and machine vision. Corresponding author of this paper

  • 摘要: 在人类信息获取过程中, 视听感知扮演着重要角色, 大脑通过整合视听信息, 形成统一、连贯且稳定的知觉体验. 视听多模态学习旨在模拟人类的视听多感官整合能力, 近年来受到研究者的广泛关注. 然而, 该领域在应用场景、任务目标和技术方法上呈现出显著的多样性, 目前尚且缺乏对视听多模态学习领域系统性回顾和分析的综合性中文综述. 基于人类的多感官整合机制在视听认知中的重要性以及不同视听多模态学习任务间的内在关联性, 提出一个统一框架, 将现有研究归纳为三类: 视听增强通过引入音频或视觉信息实现对初始单模态任务的增强效应; 跨模态交互旨在探索视听信息间的相互转换; 视听协作致力于探索视听信息的综合理解方法及其协同效应. 在此基础上, 对该领域中的最新研究进展进行系统性综述和总结. 此外, 深入剖析当前视听多模态学习研究所面临的五大核心共性问题和挑战——视听表征、对齐、转换、融合和共同学习; 探讨大模型背景下视听多模态学习的发展现状.
  • 图  1  视听多模态学习示意图

    Fig.  1  Diagram of audio-visual multi-modal learning

    图  2  视听多模态学习文献年度统计分析图

    Fig.  2  Annual statistical analysis chart of audio-visual multi-modal learning

    图  3  视听多模态学习任务及其分类示意图

    Fig.  3  Diagram of audio-visual multi-modal learning tasks and their classification

    图  4  本文结构及其章节安排示意图

    Fig.  4  Diagram of the paper structure and chapter organization

    图  5  视听增强分类及其任务划分示意图

    Fig.  5  Diagram of audio-visual enhancement categorization and its task division

    图  6  跨模态交互分类及其任务划分示意图

    Fig.  6  Diagram of cross-modal interaction categorization and its task division

    图  7  视听协作分类及其任务划分示意图

    Fig.  7  Diagram of audio-visual collaboration categorization and its task division

    图  8  核心问题/挑战示意图

    Fig.  8  Schematic diagram of the core problems/challenges

    表  1  视/听单模态学习与视听多模态学习

    Table  1  Visual/audio uni-modal learning and audio-visual multi-modal learning

    数据处理能力 知识迁移能力 噪声鲁棒能力
    计算机视觉/音频处理(CV/AP)单模态学习仅能处理图像、视频、
    音频单一模态数据
    知识只能从一种模态中
    学习并应用在该模态中
    容易受到模态自身数据
    噪声的影响
    视听多模态学习多模态学习能同时处理图像、视频、
    音频多模态数据
    知识可以从一种模态获取,
    也能应用于不同模态
    各模态数据噪声互不影响,
    且信息可相互补充
    下载: 导出CSV

    表  2  视听学习任务涉及的核心问题/挑战

    Table  2  Core issues/challenges involved in audio-visual learning tasks

    分类 子类 视听表征 视听转换 视听对齐 视听融合 视听共同学习
    视听增强学习音频增强的视觉任务$ \surd $$ \surd $$ \surd $
    视觉增强的音频任务$ \surd $$ \surd $$ \surd $
    视听跨模态学习视听生成任务$ \surd $$ \surd $$ \surd $
    视听检索任务$ \surd $$ \surd $$ \surd $
    视听协作学习视听实例感知任务$ \surd $$ \surd $$ \surd $
    视听场景理解任务$ \surd $$ \surd $$ \surd $
    视听推理与交互任务$ \surd $$ \surd $$ \surd $
    非传统的视听学习任务$ \surd $$ \surd $$ \surd $$ \surd $
    下载: 导出CSV
  • [1] Treichler D G. Are you missing the boat in training aids? Film and Audio-Visual Communication, 1967, 1(48): 14−30
    [2] 文小辉, 刘强, 孙弘进, 张庆林, 尹秦清, 郝明洁, 等. 多感官线索整合的理论模型. 心理科学进展, 2009, 17(4): 659−666

    Wen Xiao-Hui, Liu Qiang, Sun Hong-Jin, Zhang Qing-Lin, Yin Qin-Qing, Hao Ming-Jie, et al. Theoretical models of multisensory cues integration. Advances in Psychological Science, 2009, 17(4): 659−666
    [3] Calvert G A, Brammer M J, Bullmore E T, Campbell R, Iversen S D, David A S. Response amplification in sensory-specific cortices during crossmodal binding. NeuroReport, 1999, 10(12): 2619−2623 doi: 10.1097/00001756-199908200-00033
    [4] Zhu H, Luo M D, Wang R, Zheng A H, He R. Deep audio-visual learning: A survey. International Journal of Automation and Computing, 2021, 18(3): 351−376 doi: 10.1007/s11633-021-1293-0
    [5] Wei Y K, Hu D, Tian Y P, Li X L. Learning in audio-visual context: A review, analysis, and new perspective. arXiv preprint arXiv: 2208.09579, 2022.
    [6] Bedny M. Evidence from blindness for a cognitively pluripotent cortex. Trends in Cognitive Sciences, 2017, 21(9): 637−648 doi: 10.1016/j.tics.2017.06.003
    [7] Calvert G A, Bullmore E T, Brammer M J, Campbell R, Williams S C R, Mcguire P K, et al. Activation of auditory cortex during silent lipreading. Science, 1997, 276(5312): 593−596 doi: 10.1126/science.276.5312.593
    [8] Michael G A, Jacquot L, Millot J L, Brand G. Ambient odors modulate visual attentional capture. Neuroscience Letters, 2003, 352(3): 221−225 doi: 10.1016/j.neulet.2003.08.068
    [9] Shannon C E. A mathematical theory of communication. The Bell System Technical Journal, 1948, 27(3): 379−423 doi: 10.1002/j.1538-7305.1948.tb01338.x
    [10] Sedaghati N, Ardebili S, Ghaffari A. Application of human activity/action recognition: A review. Multimedia Tools and Applications, 2025, 84(28): 33475−33504 doi: 10.1007/s11042-024-20576-2
    [11] Liu Y, Tan Y, Lan H Y. Self-supervised contrastive learning for audio-visual action recognition. In: Proceedings of the IEEE International Conference on Image Processing (ICIP). Kuala Lumpur, Malaysia: IEEE, 2023. 1000−1004
    [12] Shaikh M B, Chai D, Islam S M S, Akhtar N. MAiVAR-T: Multimodal audio-image and video action recognizer using Transformers. In: Proceedings of the 11th European Workshop on Visual Information Processing. Gjovik, Norway: IEEE, 2023. 1−6
    [13] Shahabinejad M, Kezele I, Nabavi S S, Liu W T, Patel S, Yu Y H, et al. Video action recognition with adaptive zooming using motion residuals. In: Proceedings of the International Conference on Computer Vision Workshops (ICCVW). Paris, France: IEEE, 2023. 1206−1215
    [14] Shaikh M B, Chai D, Islam S M S, Akhtar N. Multimodal fusion for audio-image and video action recognition. Neural Computing and Applications, 2024, 36(10): 5499−5513 doi: 10.1007/s00521-023-09186-5
    [15] Gao R H, Oh T H, Grauman K, Torresani L. Listen to look: Action recognition by previewing audio. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE, 2020. 10454−10464
    [16] Song L C, Bi J, Huang C, Xu C L. Audio-visual action prediction with soft-boundary in egocentric videos. In: Proceedings of the International Conference on Computer Vision Workshops. New York, USA: IEEE, 2024. 1−5

    Song L C, Bi J, Huang C, Xu C L. Audio-visual action prediction with soft-boundary in egocentric videos. In: Proceedings of the International Conference on Computer Vision Workshops. New York, USA: IEEE, 2024. 1−5
    [17] Chalk J, Huh J, Kazakos E, Zisserman A, Damen D. TIM: A time interval machine for audio-visual action recognition. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE, 2024. 18153−18163
    [18] Wang K, Hatzinakos D. MOMA: Mixture-of-modality-adaptations for transferring knowledge from image models towards efficient audio-visual action recognition. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, Republic of Korea: IEEE, 2024. 8055−8059
    [19] Han H C, Zheng Q H, Luo M N, Miao K Y, Tian F, Chen Y. Noise-tolerant learning for audio-visual action recognition. IEEE Transactions on Multimedia, 2024, 26: 7761−7774 doi: 10.1109/TMM.2024.3371220
    [20] Wang W G, Shen J B, Guo F, Cheng M M, Borji A. Revisiting video saliency: A large-scale benchmark and a new model. In: Proceedings of the Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA: IEEE, 2018. 4894−4903
    [21] 罗霄骁, 康冠兰, 周晓林. McGurk效应的影响因素与神经基础. 心理科学进展, 2018, 26(11): 1935−1951 doi: 10.3724/SP.J.1042.2018.01935

    Luo Xiao-Xiao, Kang Guan-Lan, Zhou Xiao-Lin. The influential factors and neural mechanisms of McGurk effect. Advances in Psychological Science, 2018, 26(11): 1935−1951 doi: 10.3724/SP.J.1042.2018.01935
    [22] Tavakoli H R, Borji A, Rahtu E, Kannala J. DAVE: A deep audio-visual embedding for dynamic saliency prediction. arXiv preprint arXiv: 1905.10693, 2019.
    [23] Min X K, Zhai G T, Gu K, Yang X K. Fixation prediction through multimodal analysis. ACM Transactions on Multimedia Computing, Communications, and Applications, 2017, 13(1): Article No. 6 doi: 10.1109/vcip.2015.7457921
    [24] Jiang L, Xu M, Liu T, Qiao M L, Wang Z L. DeepVS: A deep learning based video saliency prediction approach. In: Proceedings of the 15th European Conference on Computer Vision. Munich, Germany: Springer, 2018. 625−642
    [25] Jain S, Yarlagadda P, Jyoti S, Karthik S, Subramanian R, Gandhi V. ViNet: Pushing the limits of visual modality for audio-visual saliency prediction. In: Proceedings of the International Conference on Intelligent Robots and Systems (IROS). Prague, Czech Republic: IEEE, 2021. 3520−3527
    [26] Xiong J W, Wang G L, Zhang P, Huang W, Zha Y F, Zhai G T. CASP-Net: Rethinking video saliency prediction from an audio-visual consistency perceptual perspective. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 6441−6450
    [27] Chang Q Y, Zhu S P. Human vision attention mechanism-inspired temporal-spatial feature pyramid for video saliency detection. Cognitive Computation, 2023, 15(3): 856−868 doi: 10.1007/s12559-023-10114-x
    [28] Xie J W, Liu Z, Li G Y, Song Y J. Audio-visual saliency prediction with multisensory perception and integration. Image and Vision Computing, 2024, 143: Article No. 104955 doi: 10.1016/j.imavis.2024.104955
    [29] Chen Z, Zhang K, Cai H, Ding X Y, Jiang C X, Chen Z Z. Audio-visual saliency prediction for movie viewing in immersive environments: Dataset and benchmarks. Journal of Visual Communication and Image Representation, 2024, 100: Article No. 104095 doi: 10.1016/j.jvcir.2024.104095
    [30] Zhu D D, Zhu K, Ding W P, Zhang N N, Min X K, Zhai G T, et al. MTCAM: A novel weakly-supervised audio-visual saliency prediction model with multi-modal Transformer. IEEE Transactions on Emerging Topics in Computational Intelligence, 2024, 8(2): 1756−1771 doi: 10.1109/TETCI.2024.3358184
    [31] Qiao M L, Liu Y F, Xu M, Deng X, Li B, Hu W M, et al. Joint learning of audio-visual saliency prediction and sound source localization on multi-face videos. International Journal of Computer Vision, 2024, 132(6): 2003−2025 doi: 10.1007/s11263-023-01950-3
    [32] Xiong J W, Zhang P, You T, Li C Y, Huang W, Zha Y F. DiffSal: Joint audio and video learning for diffusion saliency prediction. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE, 2024. 27263−27273
    [33] Khan M A, Menouar H, Hamila R. Revisiting crowd counting: State-of-the-art, trends, and future perspectives. Image and Vis-ion Computing, 2023, 129: Article No. 104597 doi: 10.1016/j.imavis.2022.104597
    [34] Hu D, Mou L C, Wang Q Z, Gao J Y, Hua Y S, Dou D J, et al. Ambient sound helps: Audiovisual crowd counting in extreme conditions. In: Proceedings of the Conference on Computer Vision and Pattern Recognition Workshops. New York, USA: IEEE, 2020. 1−4

    Hu D, Mou L C, Wang Q Z, Gao J Y, Hua Y S, Dou D J, et al. Ambient sound helps: Audiovisual crowd counting in extreme conditions. In: Proceedings of the Conference on Computer Vision and Pattern Recognition Workshops. New York, USA: IEEE, 2020. 1−4
    [35] Zou Y, Min W D, Zhao H Y, Han Q. A novel framework for crowd counting using video and audio. Computers and Electrical Engineering, 2023, 109: Article No. 108754 doi: 10.1016/j.compeleceng.2023.108754
    [36] Hu R H, Mo Q L, Xie Y F, Xu Y Q, Chen J Q, Yang Y L, et al. AVMSN: An audio-visual two stream crowd counting framework under low-quality conditions. IEEE Access, 2021, 9: 80500−80510 doi: 10.1109/ACCESS.2021.3074797
    [37] Sajid U, Chen X Y, Sajid H, Kim T, Wang G H. Audio-visual Transformer based crowd counting. In: Proceedings of the International Conference on Computer Vision Workshops (ICCVW). Montreal, Canada: IEEE, 2021. 2249−2259
    [38] Hu D, Li X H, Mou L C, Jin P, Chen D, Jing L P, et al. Cross-task transfer for geotagged audiovisual aerial scene recognition. In: Proceedings of the 16th European Conference on Computer Vision. Glasgow, UK: Springer, 2020. 68−84
    [39] Heidler K, Mou L C, Hu D, Jin P, Li G Y, Gan C, et al. Self-supervised audiovisual representation learning for remote sensing data. International Journal of Applied Earth Observation and Geoinformation, 2023, 116: Article No. 103130 doi: 10.1016/j.jag.2022.103130
    [40] Sun X, Gao J Y, Yuan Y. Alignment and fusion using distinct sensor data for multimodal aerial scene classification. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: Article No. 5626811 doi: 10.1109/tgrs.2024.3406697
    [41] Han F Z, Yu T Y, Zhang L M, Si L Y, Zhang Y Q. SlotFusion: Object-centric audiovisual feature fusion with slot attention for remote sensing scene recognition. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Hyderabad, India: IEEE, 2025. 1−5
    [42] Khanal S, Xing E, Sastry S, Dhakal A, Xiong Z X, Ahmad A, et al. PSM: Learning probabilistic embeddings for multi-scale zero-shot soundscape mapping. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024. 1361−1369
    [43] Corley I, Robinson C, Dodhia R, Ferres J M L, Najafirad P. Revisiting pre-trained remote sensing model benchmarks: Resizing and normalization matters. In: Proceedings of the Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Seattle, USA: IEEE, 2024. 3162−3172
    [44] Liu X L, Yu Y, Li X L, Zhao Y. MCL: Multimodal contrastive learning for deepfake detection. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(4): 2803−2813 doi: 10.1109/TCSVT.2023.3312738
    [45] Rana S, Nobi M N, Murali B, Sung A H. Deepfake detection: A systematic literature review. IEEE Access, 2022, 10: 25494−25513 doi: 10.1109/ACCESS.2022.3154404
    [46] Hashmi A, Shahzad S A, Lin C W, Tsao Y, Wang H M. AVTENet: Audio-visual Transformer-based ensemble network exploiting multiple experts for video deepfake detection. arXiv preprint arXiv: 2310.13103, 2023.
    [47] Zhang Y B, Lin W G, Xu J F. Joint audio-visual attention with contrastive learning for more general deepfake detection. ACM Transactions on Multimedia Computing, Communications, and Applications, 2024, 20(5): Article No. 137 doi: 10.1145/3625100
    [48] Wang R, Ye D P, Tang L, Zhang Y M, Deng J C. AVT.2-DWF: Improving deepfake detection with audio-visual fusion and dynamic weighting strategies. IEEE Signal Processing Letters, 2024, 31: 1960−1964 doi: 10.1109/LSP.2024.3433596
    [49] Wang Y F, Wu X C, Zhang J, Jing M H, Lu K D, Yu J, et al. Building robust video-level deepfake detection via audio-visual local-global interactions. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024. 11370−11376
    [50] Liu W F, She T Y, Liu J W, Li B H, Yao D Y, Liang Z Y, et al. Lips are lying: Spotting the temporal inconsistency between audio and visual in lip-syncing deepfakes. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates Inc., 2024. 91131−91155
    [51] Astrid M, Ghorbel E, Aouada D. Audio-visual deepfake detection with local temporal inconsistencies. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Hyderabad, India: IEEE, 2025. 1−5
    [52] Koutlis C, Papadopoulos S. DiMoDif: Discourse modality-information differentiation for audio-visual deepfake detection and localization. arXiv preprint arXiv: 2411.10193, 2024.
    [53] Shahzad S A, Hashmi A, Peng Y T, Tsao Y, Wang H M. AV-Lip-Sync+: Leveraging AV-HuBERT to exploit multimodal inconsistency for deepfake detection of frontal face videos. IEEE Transactions on Human-Machine Systems, 2025, 55(6): 973−982 doi: 10.1109/THMS.2025.3618409
    [54] Chugh K, Gupta P, Dhall A, Subramanian R. Not made for each other-audio-visual dissonance-based deepfake detection and localization. In: Proceedings of the 28th ACM International Conference on Multimedia. Seattle, USA: ACM, 2020. 439−447
    [55] Zou H Q, Shen M, Hu Y C, Chen C, Chng E S, Rajan D. Cross-modality and within-modality regularization for audio-visual deepfake detection. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, Republic of Korea: IEEE, 2024. 4900−4904
    [56] Bekheet A A, Ghoneim A, Khoriba G. A comprehensive comparative analysis of deepfake detection techniques in visual, audio, and audio-visual domains. In: Proceedings of the Intelligent Methods, Systems, and Applications (IMSA). Giza, Egypt: IEEE, 2024. 122−129
    [57] Nie F, Ni J Q, Zhang J, Zhang B, Zhang W Z. FRADE: Forgery-aware audio-distilled multimodal learning for deepfake detection. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024. 6297−6306
    [58] Zhao H Q, Zhou W B, Chen D D, Zhang W M, Guo Y, Cheng Z, et al. Audio-visual contrastive pre-train for face forgery detection. ACM Transactions on Multimedia Computing, Communications and Applications, 2025, 21(2): Article No. 45 doi: 10.1145/3651311
    [59] Liang Y C, Yu M, Li G, Jiang J G, Li B Q, Yu F, et al. SpeechForensics: Audio-visual speech representation learning for face forgery detection. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates Inc., 2024. Article No. 2735
    [60] Oorloff T, Koppisetti S, Bonettini N, Solanki D, Colman B, Yacoob Y, et al. AVFF: Audio-visual feature fusion for video deepfake detection. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE, 2024. 27092−27102
    [61] Li X L, Liu Z H, Chen C, Li L T, Guo L, Wang D. Zero-shot fake video detection by audio-visual consistency. In: Proceedings of the 25th Annual Conference of the International Speech Communication Association. Kos, Greece: ISCA, 2024.
    [62] Muppalla S, Jia S, Lyu S W. Integrating audio-visual features for multimodal deepfake detection. In: Proceedings of the MIT Undergraduate Research Technology Conference (URTC). Cam-bridge, USA: IEEE, 2023. 1−5
    [63] Yang W Y, Zhou X Y, Chen Z K, Guo B F, Ba Z J, Xia Z H, et al. AVoiD-DF: Audio-visual joint learning for detecting deepfake. IEEE Transactions on Information Forensics and Security, 2023, 18: 2015−2029 doi: 10.1109/TIFS.2023.3262148
    [64] Yu C, Chen P, Tian J H, Liu J, Dai J, Wang X, et al. Modality-agnostic audio-visual deepfake detection. arXiv preprint arXiv: 2307.14491, 2023.
    [65] 李金新, 黄志勇, 李文斌, 周登文. 基于多层次特征融合的图像超分辨率重建. 自动化学报, 2023, 49(1): 161−171 doi: 10.16383/j.aas.c200585

    Li Jin-Xin, Huang Zhi-Yong, Li Wen-Bin, Zhou Deng-Wen. Image super-resolution based on multi-hierarchical features fusion network. Acta Automatica Sinica, 2023, 49(1): 161−171 doi: 10.16383/j.aas.c200585
    [66] Sanguineti V, Thakur S, Morerio P, del Bue A, Murino A. Audio-visual inpainting: Reconstructing missing visual information with sound. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece: IEEE, 2023. 1−5
    [67] Lu Y F, Wang Z P, Liu M J, Wang H J, Wang L. Learning spatial-temporal implicit neural representations for event-guided video super-resolution. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 1557−1567
    [68] Chen Y X, Zhao P C, Qi M B, Zhao Y, Jia W, Wang R G. Audio matters in video super-resolution by implicit semantic guidance. IEEE Transactions on Multimedia, 2022, 24: 4128−4142 doi: 10.1109/TMM.2022.3152941
    [69] Xiao J, Jiang X Y, Zheng N X, Yang H, Yang Y F, Yang Y Q, et al. Online video super-resolution with convolutional kernel bypass grafts. IEEE Transactions on Multimedia, 2023, 25: 8972−8987 doi: 10.1109/TMM.2023.3243615
    [70] Chakraborty C, Talukdar P H. Issues and limitations of HMM in speech processing: A survey. International Journal of Computer Applications, 2016, 141(7): 13−17 doi: 10.5120/ijca2016909693
    [71] Fang H J, Frintrop S, Gerkmann T. Uncertainty-driven hybrid fusion for audio-visual phoneme recognition. In: Proceedings of the 15th ITG Conference on Speech Communication. Aachen, Germany: VDE, 2023. 255−259
    [72] Richter J, Liebold J, Gerkmann T. Continuous phoneme recognition based on audio-visual modality fusion. In: Proceedings of the International Joint Conference on Neural Networks (IJCNN). Padua, Italy: IEEE, 2022. 1−8
    [73] Biswas A, Sahu P K, Bhowmick A, Chandra M. VidTIMIT audio visual phoneme recognition using AAM visual features and human auditory motivated acoustic wavelet features. In: Proceedings of the 2nd International Conference on Recent Trends in Information Systems (ReTIS). Kolkata, India: IEEE, 2015. 428−433
    [74] Hu Y C, Li R Z, Chen C, Qin C W, Zhu Q S, Chng E S. Hearing lips in noise: Universal viseme-phoneme mapping and transfer for robust audio-visual speech recognition. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Toronto, Canada: Association for Computational Linguistics, 2023. 15213−15232
    [75] Kim M, Yeo J, Park S J, Rha H, Ro Y M. Efficient training for multilingual visual speech recognition: Pre-training with discretized visual speech representation. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024. 1311−1320
    [76] Pai L T, Wang Y C, Yan B C, Wang H W, Lu J L, Lin C H, et al. An effective contextualized automatic speech recognition approach leveraging self-supervised phoneme features. In: Proceedings of the Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). Macao, China: IEEE, 2024. 1−6
    [77] Vincent E, Virtanen T, Gannot S. Audio Source Separation and Speech Enhancement. Hoboken: John Wiley & Sons, 2018. 110−234
    [78] Li G N, Deng J J, Geng M Z, Jin Z R, Wang T Z, Hu S J, et al. Audio-visual end-to-end multi-channel speech separation, dereverberation and recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023, 31: 2707−2723 doi: 10.1109/TASLP.2023.3294705
    [79] Tan R B, Ray A, Burns A, Plummer B A, Salamon J, Nieto O, et al. Language-guided audio-visual source separation via trimodal consistency. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 10575−10584
    [80] Gao R H, Grauman K. VisualVoice: Audio-visual speech separation with cross-modal consistency. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2021. 15490−15500
    [81] Yoshinaga T, Tanaka K, Morishima S. Audio-visual speech enhancement with selective off-screen speech extraction. In: Proceedings of the 31st European Signal Processing Conference (EUSIPCO). Helsinki, Finland: IEEE, 2023. 595−599
    [82] Pan Z X, Wichern G, Masuyama Y, Germain F G, Khurana S, Hori C, et al. Scenario-aware audio-visual TF-Gridnet for target speech extraction. In: Proceedings of the Automatic Speech Recognition and Understanding Workshop (ASRU). Taipei, China: IEEE, 2023. 1−8
    [83] Chen J B, Zhang R R, Lian D Z, Yang J Q, Zeng Z Y, Shi J B. iQuery: Instruments as queries for audio-visual sound separation. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 14675−14686
    [84] Chatterjee M, le Roux J, Ahuja N, Cherian A. Visual scene graphs for audio source separation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Montreal, Canada: IEEE, 2021. 1184−1193
    [85] Ye Y X, Yang W M, Tian Y P. LAVSS: Location-guided audio-visual spatial audio separation. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA: IEEE, 2024. 5496−5507
    [86] Pan T R, Liu J, Wang B H, Tang J, Wu G S. RAVSS: Robust audio-visual speech separation in multi-speaker scenarios with missing visual cues. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024. 4748−4756
    [87] Li K, Xie F H, Chen H, Yuan K X, Hu X L. An audio-visual speech separation model inspired by cortico-thalamo-cortical circuits. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(10): 6637−6651 doi: 10.1109/TPAMI.2024.3384034
    [88] Liu Y G, Deng Y J, Wei Y. A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin. IEEE Journal of Selected Topics in Signal Processing, 2024, 18(3): 459−472 doi: 10.1109/JSTSP.2024.3427424
    [89] Pian W G, Nan Y Y, Deng S J, Mo S T, Guo Y H, Tian Y P. Continual audio-visual sound separation. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates Inc., 2024. Article No. 2421
    [90] Kalkhorani V A, Kumar A, Tan K, Xu B Y, Wang D L. Audiovisual speaker separation with full-and sub-band modeling in the time-frequency domain. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, Republic of Korea: IEEE, 2024. 12001−12005
    [91] Fan C H, Xiang W, Tao J H, Yi J Y, Lv Z. Cross-modal knowledge distillation with multi-stage adaptive feature fusion for speech separation. IEEE Transactions on Audio, Speech and Language Processing, 2025, 33: 935−948 doi: 10.1109/TASLPRO.2025.3533359
    [92] Boll S. Suppression of acoustic noise in speech using spectral subtraction. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing. Piscataway, USA: IEEE, 2020. 7314−7318

    Boll S. Suppression of acoustic noise in speech using spectral subtraction. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing. Piscataway, USA: IEEE, 2020. 7314−7318
    [93] Isik Y, le Roux J, Chen Z, Watanabe S, Hershey J R. Single-channel multi-speaker separation using deep clustering. In: Proceedings of the 17th Annual Conference of the International Speech Communication Association. San Francisco, USA: ISCA, 2016. 545−549
    [94] Zhu Z R, Yang H M, Tang M, Yang Z Y, Eskimez S E, Wang H M. Real-time audio-visual end-to-end speech enhancement. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece: IEEE, 2023. 1−5
    [95] Balasubramanian S, Rajavel R, Kar A. Ideal ratio mask estimation based on cochleagram for audio-visual monaural speech enhancement. Applied Acoustics, 2023, 211: Article No. 109524 doi: 10.1016/j.apacoust.2023.109524
    [96] Li Y K, Zhang X M. Lip landmark-based audio-visual speech enhancement with multimodal feature fusion network. Neurocomputing, 2023, 549: Article No. 126432 doi: 10.1016/j.neucom.2023.126432
    [97] Hussain T, Dashtipour K, Tsao Y, Hussain A. Audio-visual speech enhancement in noisy environments via emotion-based contextual cues. arXiv preprint arXiv: 2402.16394, 2024.
    [98] Zheng R C, Ai Y, Ling Z H. Incorporating ultrasound tongue images for audio-visual speech enhancement. IEEE/ACM Tran-sactions on Audio, Speech, and Language Processing, 2024, 32: 1430−1444 doi: 10.1109/TASLP.2024.3361376
    [99] Gogate M, Dashtipour K, Hussain A. Robust real-time audio-visual speech enhancement based on DNN and GAN. IEEE Transactions on Artificial Intelligence, 2025, 6(11): 2860−2869 doi: 10.1109/TAI.2024.3366141
    [100] Chen H L, Mira R, Petridis S, Pantic M. RT-LA-VocE: Real-time low-SNR audio-visual speech enhancement. arXiv preprint arXiv: 2407.07825, 2024.
    [101] Jung C, Lee S, Kim J H, Chung J S. FlowAVSE: Efficient audio-visual speech enhancement with conditional flow matching. arXiv preprint arXiv: 2406.09286, 2024.
    [102] Passos L A, Papa J P, del Ser J, Hussain A, Adeel A. Multimodal audio-visual information fusion using canonical-correlated graph neural network for energy-efficient speech enhancement. Information Fusion, 2023, 90: 1−11 doi: 10.2139/ssrn.4184514
    [103] Morrone G, Michelsanti D, Tan Z H, Jensen J. Audio-visual speech inpainting with deep learning. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Toronto, Canada: IEEE, 2021. 6653−6657
    [104] Chen H, Wang Q, Du J, Yin B C, Pan J, Lee C H. Optimizing audio-visual speech enhancement using multi-level distortion measures for audio-visual speech recognition. IEEE/ACM Tran-sactions on Audio, Speech, and Language Processing, 2024, 32: 2508−2521 doi: 10.1109/TASLP.2024.3393732
    [105] Chen S, Kirton-Wingate J, Doctor F, Arshad U, Dashtipour K, Gogate M, et al. Context-aware audio-visual speech enhancement based on neuro-fuzzy modeling and user preference learning. IEEE Transactions on Fuzzy Systems, 2024, 32(10): 5400−5412 doi: 10.1109/TFUZZ.2024.3435050
    [106] Ahlawat H, Aggarwal N, Gupta D. Automatic speech recognition: A survey of deep learning techniques and approaches. International Journal of Cognitive Computing in Engineering, 2025, 6: 201−237 doi: 10.1016/j.ijcce.2024.12.007
    [107] Hong J, Kim M, Choi J, Ro Y M. Watch or listen: Robust audio-visual speech recognition with visual corruption modeling and reliability scoring. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 18783−18794
    [108] Wang X M, Mi J C, Li B Q, Zhao Y X, Meng J X. CATNet: Cross-modal fusion for audio-visual speech recognition. Pattern Recognition Letters, 2024, 178: 216−222 doi: 10.1016/j.patrec.2024.01.002
    [109] Wang H, Guo P C, Zhou P, Xie L. MLCA-AVSR: Multi-layer cross attention fusion based audio-visual speech recognition. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, Republic of Korea: IEEE, 2024. 8150−8154
    [110] Wang J D, Qian X Y, Li H Z. Predict-and-update network: Audio-visual speech recognition inspired by human speech perception. IEEE Transactions on Audio, Speech and Language Processing, 2025, 33: 11−22 doi: 10.1109/TASLP.2024.3507575
    [111] Ma P C, Haliassos A, Fernandez-Lopez A, Chen H L, Petridis S, Pantic M. Auto-AVSR: Audio-visual speech recognition with automatic labels. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece: IEEE, 2023. 1−5
    [112] Yeo J H, Kim M, Choi J, Kim D H, Ro Y M. AKVSR: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model. IEEE Transactions on Multimedia, 2024, 26: 6462−6474 doi: 10.1109/TMM.2024.3352388
    [113] Lian J C, Baevski A, Hsu W N, Auli M. Av-Data2Vec: Self-supervised learning of audio-visual speech representations with contextualized target representations. In: Proceedings of the Automatic Speech Recognition and Understanding Workshop (ASRU). Taipei, China: IEEE, 2023. 1−8
    [114] Dai Y S, Chen H, Du J, Wang R Y, Chen S H, Wang H T, et al. A study of dropout-induced modality bias on robustness to missing video frames for audio-visual speech recognition. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE, 2024. 27435−27445
    [115] Rouditchenko A, Gong Y, Thomas S, Karlinsky L, Kuehne H, Feris R, et al. Whisper-Flamingo: Integrating visual features into whisper for audio-visual speech recognition and translation. In: Proceedings of the 25th Annual Conference of the International Speech Communication Association. Kos, Greece: ISCA, 2024. 2420−2424
    [116] Wang J D, Pan Z X, Zhang M L, Tan R T, Li H Z. Restoring speaking lips from occlusion for audio-visual speech recognition. In: Proceedings of the 38th AAAI Conference on Artificial Intelligence. Vancouver, Canada: AAAI Press, 2024. 19144−19152
    [117] Li J H, Li C D, Wu Y F, Qian Y M. Unified cross-modal attention: Robust audio-visual speech recognition and beyond. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024, 32: 1941−1953 doi: 10.1109/TASLP.2024.3375641
    [118] Fu D J, Cheng X Z, Yang X D, Wang H T, Zhao Z, Jin T. Boosting speech recognition robustness to modality-distortion with contrast-augmented prompts. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024. 3838−3847
    [119] Burchi M, Puvvada K C, Balam J, Ginsburg B, Timofte R. Multilingual audio-visual speech recognition with hybrid CTC/RNN-T fast conformer. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, Republic of Korea: IEEE, 2024. 10211−10215
    [120] Kabir M M, Mridha M F, Shin J, Jahan I, Ohi A Q. A survey of speaker recognition: Fundamental theories, recognition meth-ods and opportunities. IEEE Access, 2021, 9: 79236−79263 doi: 10.1109/ACCESS.2021.3084299
    [121] Chelali F Z. Bimodal fusion of visual and speech data for audiovisual speaker recognition in noisy environment. International Journal of Information Technology, 2023, 15(6): 3135−3145 doi: 10.1007/s41870-023-01291-x
    [122] Tang X O, Li Z F. Audio-guided video-based face recognition. IEEE Transactions on Circuits and Systems for Video Technology, 2009, 19(7): 955−964 doi: 10.1109/TCSVT.2009.2022694
    [123] Gong D H, Li N, Li Z F, Qiao Y. Multi-feature subspace analysis for audio-vidoe based multi-modal person recognition. In: Proceedings of the 4th IEEE International Conference on Information Science and Technology. Shenzhen, China: IEEE, 2014. 776−779
    [124] Tao R J, Lee K A, Shi Z, Li H Z. Speaker recognition with two-step multi-modal deep cleansing. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece: IEEE, 2023. 1−5
    [125] Gebru I D, Ba S, Li X F, Horaud R. Audio-visual speaker diarization based on spatiotemporal Bayesian fusion. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018, 40(5): 1086−1099 doi: 10.1109/TPAMI.2017.2648793
    [126] Lin Y K, Cheng M, Zhang F L, Gao Y Y, Zhang S L, Li M. VoxBlink2: A 100k+ speaker recognition corpus and the open-set speaker-identification benchmark. arXiv preprint arXiv: 2407.11510, 2024.
    [127] Clarke J, Gotoh Y, Goetze S. Speaker embedding informed audiovisual active speaker detection for egocentric recordings. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Hyderabad, India: IEEE, 2025. 1−5
    [128] Tao R J, Qian X Y, Jiang Y D, Li J J, Wang J D, Li H Z. Audio-visual target speaker extraction with selective auditory attention. IEEE Transactions on Audio, Speech and Language Processing, 2025, 33: 797−811 doi: 10.1109/TASLPRO.2025.3527766
    [129] 李韩超, 沈成泽, 刘新国. 带虚拟边约束的面部表情基生成方法. 计算机学报, 2023, 46(11): 2453−2462

    Li Han-Chao, Shen Cheng-Ze, Liu Xin-Guo. Facial blendshape generation method with virtual edge constraints. Chinese Journal of Computers, 2023, 46(11): 2453−2462
    [130] Agarwal M, Mukhopadhyay R, Namboodiri V P, Jawahar C V. Audio-visual face reenactment. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA: IEEE, 2023. 5167−5176
    [131] Wang S Z, Li L C, Ding Y, Yu X. One-shot talking face generation from single-speaker audio-visual correlation learning. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence. Virtual Event: AAAI Press, 2022. 2531−2539

    Wang S Z, Li L C, Ding Y, Yu X. One-shot talking face generation from single-speaker audio-visual correlation learning. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence. Virtual Event: AAAI Press, 2022. 2531−2539
    [132] Jang Y, Rho K, Woo J, Lee H, Park J, Lim Y, et al. That's what I said: Fully-controllable talking face generation. In: Proceedings of the 31st ACM International Conference on Multimedia. Ottawa, Canada: ACM, 2023. 3827−3836
    [133] Park S J, Kim M, Hong J, Choi J, Ro Y M. SyncTalkFace: Talking face generation with precise lip-syncing via audio-lip memory. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence. Virtual Event: AAAI Press, 2022. 2062−2070

    Park S J, Kim M, Hong J, Choi J, Ro Y M. SyncTalkFace: Talking face generation with precise lip-syncing via audio-lip memory. In: Proceedings of the 36th AAAI Conference on Artificial Intelligence. Virtual Event: AAAI Press, 2022. 2062−2070
    [134] Wang G L, Zhang P, Xie L, Huang W, Zha Y F. Attention-based lip audio-visual synthesis for talking face generation in the wild. arXiv preprint arXiv: 2203.03984, 2022.
    [135] Zhang J Y, Liu Y Z, Li X, Li W, Tang Y. Talking face generation driven by time-frequency domain features of speech audio. Displays, 2023, 80: Article No. 102558 doi: 10.1016/j.displa.2023.102558
    [136] Liu S G, Wang H X. Talking face generation via facial anatomy. ACM Transactions on Multimedia Computing, Communications and Applications, 2023, 19(3): Article No. 125 doi: 10.1145/3571746
    [137] Yaman D, Eyiokur F I, Bärmann L, Aktı S, Ekenel H K, Waibel A. Audio-visual speech representation expert for enhanced talking face video generation and evaluation. In: Proceedings of the Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Seattle, USA: IEEE, 2024. 6003−6013.
    [138] Jang Y, Kim J H, Ahn J, Kwak D, Yang H S, Ju Y C, et al. Faces that speak: Jointly synthesising talking face and speech from text. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE, 2024. 8818−8828
    [139] Xu S C, Chen G J, Guo Y X, Yang J L, Li C, Zang Z Y, et al. VASA-1: Lifelike audio-driven talking faces generated in real time. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates Inc., 2024. Article No. 21
    [140] Zhang Z J, Zhang J, Mai W J. VPT: Video portraits Transformer for realistic talking face generation. Neural Networks, 2025, 184: Article No. 107122 doi: 10.1016/j.neunet.2025.107122
    [141] Nyatsanga S, Kucherenko T, Ahuja C, Henter G E, Neff M. A comprehensive review of data-driven co-speech gesture generation. Computer Graphics Forum, 2023, 42(2): 569−596
    [142] Cassell J, Vilhjálmsson H H, Bickmore T. BEAT: The behavior expression animation toolkit. In: Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques. Los Angeles, USA: ACM, 2001. 477−486
    [143] Neff M, Kipp M, Albrecht I, Seidel H P. Gesture modeling and animation based on a probabilistic re-creation of speaker style. ACM Transactions on Graphics, 2008, 27(1): Article No. 5 doi: 10.1145/1330511.1330516
    [144] Ferstl Y, McDonnell R. Investigating the use of recurrent motion modelling for speech gesture generation. In: Proceedings of the 18th International Conference on Intelligent Virtual Age-nts. Sydney, Australia: ACM, 2018. 93−98
    [145] Liang Y Z, Feng Q Y, Zhu L C, Hu L, Pan P, Yang Y. SEEG: Semantic energized co-speech gesture generation. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, USA: IEEE, 2022. 10463−10472
    [146] Hasegawa D, Kaneko N, Shirakawa S, Sakuta H, Sumi K. Evaluation of speech-to-gesture generation using bi-directional LSTM network. In: Proceedings of the 18th International Conference on Intelligent Virtual Agents. Sydney, Australia: ACM, 2018. 79−86
    [147] Yoon Y, Cha B, Lee J H, Jang M, Lee J, Kim J, et al. Speech gesture generation from the trimodal context of text, audio, and speaker identity. ACM Transactions on Graphics, 2020, 39(6): Article No. 222 doi: 10.1145/3414685.3417838
    [148] Qi X Q, Liu C, Li L C, Hou J, Xin H R, Yu X. EmotionGesture: Audio-driven diverse emotional co-speech 3D gesture generation. IEEE Transactions on Multimedia, 2024, 26: 10420−10430 doi: 10.1109/TMM.2024.3407692
    [149] Liu P X, Zhang P F, Kim H, Garrido P, Shapiro A, Olszewski K. Contextual gesture: Co-speech gesture video generation through context-aware gesture representation. In: Proceedings of the 33rd ACM International Conference on Multimedia. Dublin, Ireland: ACM, 2025. 9803−9812
    [150] Gao Z L, Li Y H, Wu S J, Cao Y Q, Duan H Y, Zhai G T. GES-QA: A multidimensional quality assessment dataset for audio-to-3D gesture generation. In: Proceedings of the International Conference on Visual Communications and Image Processing (VCIP). Klagenfurt, Austria: IEEE, 2025. 1−5
    [151] Liu H Y, Zhu Z H, Becherini G, Peng Y C, Su M Y, Zhou Y, et al. EMAGE: Towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE, 2024. 1144−1154
    [152] Chen J H, Yang H, Shi R H, Ding C F, Mo X Q, Xiong S Y, et al. Audio-driven gesture generation via deviation feature in the latent space. In: Proceedings of the IEEE International Conference on Multimedia and Expo. Nantes, France: IEEE, 2025. 1−6
    [153] Lee M, Lee K, Park J. Music similarity-based approach to generating dance motion sequence. Multimedia Tools and Applications, 2013, 62(3): 895−912 doi: 10.1007/s11042-012-1288-5
    [154] Huang Y H, Zhang J J, Liu S Y, Bao Q, Zeng D, Chen Z N, et al. Genre-conditioned long-term 3D dance generation driven by music. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Singapore: IEEE, 2022. 4858−4862
    [155] Wang T, Li L J, Lin K, Zhai Y H, Lin C C, Yang Z Y, et al. Disco: Disentangled control for realistic human dance generation. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE, 2024. 9326−9336
    [156] Tseng J, Castellon R, Liu C K. EDGE: Editable dance generation from music. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 448−458
    [157] Yang Z P, Wen Y H, Chen S Y, Liu X, Gao Y, Liu Y J, et al. Keyframe control of music-driven 3D dance generation. IEEE Transactions on Visualization and Computer Graphics, 2024, 30(7): 3474−3486 doi: 10.1109/TVCG.2023.3235538
    [158] Li S Y, Yu W J, Gu T P, Lin C Z, Wang Q, Qian C, et al. Bailando: 3D dance generation by actor-critic GPT with choreographic memory. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(12): 14192−14207
    [159] Yin W J, Yin H, Baraka K, Kragic D, Björkman M. Dance style transfer with cross-modal Transformer. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA: IEEE, 2023. 5047−5056
    [160] Habibie I, Xu W P, Mehta D, Liu L J, Seidel H P, Pons-Moll G, et al. Learning speech-driven 3D conversational gestures from video. In: Proceedings of the 21st ACM International Conference on Intelligent Virtual Agents. Virtual Event: ACM, 2021. 101−108

    Habibie I, Xu W P, Mehta D, Liu L J, Seidel H P, Pons-Moll G, et al. Learning speech-driven 3D conversational gestures from video. In: Proceedings of the 21st ACM International Conference on Intelligent Virtual Agents. Virtual Event: ACM, 2021. 101−108
    [161] Yi H W, Liang H L, Liu Y F, Cao Q, Wen Y D, Bolkart T, et al. Generating holistic 3D human motion from speech. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 469−480
    [162] Zhu H, Li Y, Zhu F X, Zheng A H, He R. Let's play music: Audio-driven performance video generation. In: Proceedings of the International Conference on Pattern Recognition (ICPR). Milan, Italy: IEEE, 2021. 3574−3581
    [163] Xu S Y, Dou Z Y, Shi M Y, Pan L, Ho L, Wang J B, et al. MOSPA: Human motion generation driven by spatial audio. In: Proceedings of the 39th Annual Conference on Neural Information Processing Systems. San Diego, USA: OpenReview.net, 2025.
    [164] Zhang Z Y, Wang Y R, Mao W, Li D N, Zhao R, Wu B, et al. Motion anything: Any to motion generation. arXiv preprint arXiv: 2503.06955, 2025.
    [165] Zhang M Y, Jin D S, Gu C Y, Hong F Z, Cai Z G, Huang J F, et al. Large motion model for unified multi-modal motion generation. In: Proceedings of the 18th European Conference on Computer Vision. Milan, Italy: Springer, 2024. 397−421
    [166] Shim J Y, Kim J, Kim J K. S2I-Bird: Sound-to-image generation of bird species using generative adversarial networks. In: Proceedings of the International Conference on Pattern Recognition (ICPR). Milan, Italy: IEEE, 2021. 2226−2232
    [167] Song C J, Zhang Y C, Peng W, Mohaghegh P, Wandt B, Rhodin H. AudioViewer: Learning to visualize sounds. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA: IEEE, 2023. 2205−2215
    [168] Sung-Bin K, Senocak A, Ha H, Owens A, Oh T H. Sound to visual scene generation by audio-to-visual latent alignment. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 6430−6440
    [169] Lee S H, Roh W, Byeon W, Yoon S H, Kim C, Kim J, et al. Sound-guided semantic image manipulation. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, USA: IEEE, 2022. 3367−3376
    [170] Zhuang Y G, Kang Y H, Fei T, Bian M, Du Y Y. From hearing to seeing: Linking auditory and visual place perceptions with soundscape-to-image generative artificial intelligence. Computers, Environment and Urban Systems, 2024, 110: Article No. 102122 doi: 10.1016/j.compenvurbsys.2024.102122
    [171] Ephrat A, Peleg S. Vid2speech: Speech reconstruction from silent video. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). New Orleans, USA: IEEE, 2017. 5095−5099
    [172] Kumar Y, Aggarwal M, Nawal P, Satoh S, Shah R R, Zimmermann R. Harnessing AI for speech reconstruction using multi-view silent video feed. In: Proceedings of the 26th ACM International Conference on Multimedia. Seoul, Republic of Korea: ACM, 2018. 1976−1983
    [173] Dong Z P, Xu Y, Abel A, Wang D. Lip2Speech: Lightweight multi-speaker speech reconstruction with Gabor features. Applied Sciences, 2024, 14(2): Article No. 798 doi: 10.3390/app14020798
    [174] Kefalas T, Panagakis Y, Pantic M. Audio-visual video-to-speech synthesis with synthesized input audio. arXiv preprint arXiv: 2307.16584, 2023.
    [175] Hong J, Kim M, Ro Y M. VisageSynTalk: Unseen speaker video-to-speech synthesis via speech-visage feature selection. In: Proceedings of the 17th European Conference on Computer Vision. Tel Aviv, Israel: Springer, 2022. 452−468
    [176] Kameoka H, Tanaka K, Puche A V, Ohishi Y, Kaneko T. Crossmodal voice conversion. arXiv preprint arXiv: 1904.04540, 2019.
    [177] Weng S E, Shuai H H, Cheng W H. Zero-shot face-based voice conversion: Bottleneck-free speech disentanglement in the real-world scenario. In: Proceedings of the 37th AAAI Conference on Artificial Intelligence. Washington, USA: AAAI Press, 2023. 13718−13726
    [178] Wang D S, Yang S, Su D, Liu X Y, Yu D, Meng H. VCVTS: Multi-speaker video-to-speech synthesis via cross-modal knowledge transfer from voice conversion. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Singapore: IEEE, 2022. 7252−7256
    [179] Lu H H, Weng S E, Yen Y F, Shuai H H, Cheng W H. Face-based voice conversion: Learning the voice behind a face. In: Proceedings of the 29th ACM International Conference on Multimedia. Virtual Event: ACM, 2021. 496−505

    Lu H H, Weng S E, Yen Y F, Shuai H H, Cheng W H. Face-based voice conversion: Learning the voice behind a face. In: Proceedings of the 29th ACM International Conference on Multimedia. Virtual Event: ACM, 2021. 496−505
    [180] Mira R, Vougioukas K, Ma P C, Petridis S, Schuller B W, Pantic M. End-to-end video-to-speech synthesis using generative adversarial networks. IEEE Transactions on Cybernetics, 2023, 53(6): 3454−3466 doi: 10.1109/TCYB.2022.3162495
    [181] Chen X Y, Wang Y J, Wu X X, Wang D S, Wu Z Y, Liu X Y, et al. Exploiting audio-visual features with pretrained AV-HuBERT for multi-modal dysarthric speech reconstruction. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, Republic of Korea: IEEE, 2024. 12341−12345
    [182] Pham L K, Tran T V T, Pham M T, Nguyen V. RESOUND: Speech reconstruction from silent videos via acoustic-semantic decomposed modeling. arXiv preprint arXiv: 2505.22024, 2025.
    [183] Rong Y, Liu L. Seeing your speech style: A novel zero-shot identity-disentanglement face-based voice conversion. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, Pennsylvania: AAAI Press, 2025. 25092−25100
    [184] Zhu Y, Olszewski K, Wu Y, Achlioptas P, Chai M L, Yan Y, et al. Quantized GAN for complex music generation from dance videos. In: Proceedings of the 17th European Conference on Computer Vision. Tel Aviv, Israel: Springer, 2022. 182−199
    [185] Su K, Li J Y, Huang Q Q, Kuzmin D, Lee J, Donahue C, et al. V2Meow: Meowing to the visual beat via video-to-music generation. In: Proceedings of the 38th AAAI Conference on Artificial Intelligence. Vancouver, Canada: AAAI Press, 2024. 4952−4960
    [186] Liu X H, Tu T, Ma Y S, Chua T S. Extending visual dynamics for video-to-music generation. arXiv preprint arXiv: 2504.07594, 2025.
    [187] Lin Y B, Tian Y, Yang L J, Bertasius G, Wang H. VMAs: Video-to-music generation via semantic alignment in web music videos. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Tucson, USA: IEEE, 2025. 1155−1165
    [188] Wang B S, Zhuo L, Wang Z K, Bao C X, Wu C J, Nie X C, et al. Multimodal music generation with explicit bridges and retrieval augmentation. arXiv preprint arXiv: 2412.09428. 2024.

    Wang B S, Zhuo L, Wang Z K, Bao C X, Wu C J, Nie X C, et al. Multimodal music generation with explicit bridges and retrieval augmentation. arXiv preprint arXiv: 2412.09428. 2024.
    [189] Tian Z Y, Liu Z Y, Yuan R B, Pan J H, Liu Q F, Tan X, et al. VidMuse: A simple video-to-music generation framework with long-short-term modeling. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2025. 18782−18793
    [190] Owens A, Isola P, McDermott J, Torralba A, Adelson E H, Freeman W T. Visually indicated sounds. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas, USA: IEEE, 2016. 2405−2413
    [191] Zhou Y P, Wang Z W, Fang C, Bui T, Berg T L. Visual to sound: Generating natural sound for videos in the wild. In: Proceedings of the Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA: IEEE, 2018. 3550−3558
    [192] Du Y X, Chen Z Y, Salamon J, Russell B, Owens A. Conditional generation of audio from video via Foley analogies. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 2426−2436
    [193] Iashin V, Rahtu E. Taming visually guided sound generation. In: Proceedings of the British Machine Vision Conference. London, UK: BMVA, 2021. 1−15

    Iashin V, Rahtu E. Taming visually guided sound generation. In: Proceedings of the British Machine Vision Conference. London, UK: BMVA, 2021. 1−15
    [194] Sheffer R, Adi Y. I hear your true colors: Image guided audio generation. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece: IEEE, 2023. 1−5
    [195] Chen P H, Zhang Y, Tan M K, Xiao H D, Huang D, Gan C. Generating visually aligned sound from videos. IEEE Transactions on Image Processing, 2020, 29: 8292−8302 doi: 10.1109/TIP.2020.3009820
    [196] Liu H D, Luo K C, Wang J L, Wang W, Chen Q, Zhao Z, et al. ThinkSound: Chain-of-thought reasoning in multimodal LLMs for audio generation and editing. In: Proceedings of the 39th Annual Conference on Neural Information Processing Systems (NeurIPS). San Diego, USA: OpenReview.net, 2025.
    [197] Jeong Y, Kim Y, Chun S, Lee J. Read, watch and scream! Sound generation from text and video. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, Pennsylvania, USA: AAAI Press, 2025. 17590−17598
    [198] Xie Z F, Yu S Y, He Q L, Li M T. SonicVisionLM: Playing sound with vision language models. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Seattle, USA: IEEE, 2024. 26856−26865
    [199] Chen Z Y, Geng D, Owens A. Images that sound: Composing images and sounds on a single canvas. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates Inc., 2024. Article No. 2700
    [200] Hawley M L, Litovsky R Y, Culling J F. The benefit of binaural hearing in a cocktail party: Effect of location and type of interferer. The Journal of the Acoustical Society of America, 2004, 115(2): 833−843 doi: 10.1121/1.1639908
    [201] 王睿琦, 程皓楠, 叶龙. 分层特征编解码驱动的视觉引导立体声生成方法. 软件学报, 2024, 35(5): 2165−2175 doi: 10.13328/j.cnki.jos.007027

    Wang Rui-Qi, Cheng Hao-Nan, Ye Long. Visually guided binaural audio generation method based on hierarchical feature encoding and decoding. Journal of Software, 2024, 35(5): 2165−2175 doi: 10.13328/j.cnki.jos.007027
    [202] Cheng C I, Wakefield G H. Introduction to head-related transfer functions (HRTFs): Representations of HRTFs in time, frequency, and space. Journal of the Audio Engineering Society, 2001, 49(4): 5026−5035
    [203] Lin Y Q, Lee D D. Bayesian regularization and nonnegative deconvolution for room impulse response estimation. IEEE Transactions on Signal Processing, 2006, 54(3): 839−847 doi: 10.1109/TSP.2005.863030
    [204] Morgado P, Vasconcelos N, Langlois T, Wang O. Self-supervised generation of spatial audio for 360° video. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems. Montréal, Canada: Curran Associates Inc., 2018. 360−370
    [205] Parida K K, Srivastava S, Sharma G. Beyond mono to binaural: Generating binaural audio from mono audio with depth and cross modal attention. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA: IEEE, 2022. 2151−2160
    [206] Li Z J, Zhao B, Yuan Y. Cross-modal generative model for visual-guided binaural stereo generation. Knowledge-Based Systems, 2024, 296: Article No. 111814 doi: 10.1016/j.knosys.2024.111814
    [207] Chen M F, Shlizerman E. AV-Cloud: Spatial audio rendering through audio-visual cloud splatting. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates Inc., 2024. Article No. 4477
    [208] Xie S Y, Zhu H X, Chen X Y, He T Y, Li X, Chen Z B. Sonic4D: Spatial audio generation for immersive 4D scene exploration. arXiv preprint arXiv: 2506.15759, 2025.
    [209] Marinoni C, Gramaccioni R F, Shimada K, Shibuya T, Mitsufuji Y, Comminiello D. StereoSync: Spatially-aware stereo audio generation from video. In: Proceedings of the International Joint Conference on Neural Networks. Rome, Italy: IEEE, 2025. 1−8
    [210] Li X D, Zhuo F, Luo D, Chen J, Kang S Y, Wu Z Y, et al. Generating stereophonic music with single-stage language models. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, Republic of Korea: IEEE, 2024. 1471−1475
    [211] Kim J, Yun H, Kim G. ViSAGe: Video-to-spatial audio generation. In: Proceedings of the 13th International Conference on Learning Representations (ICLR). Singapore: OpenReview.net, 2025.
    [212] Nagrani A, Albanie S, Zisserman A. Seeing voices and hearing faces: Cross-modal biometric matching. In: Proceedings of the Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA: IEEE, 2018. 8427−8436
    [213] Fang Z, Liu Z, Hung C C, Sekhavat Y A, Liu T T, Wang X. Learning coordinated emotion representation between voice and face. Applied Intelligence, 2023, 53(11): 14470−14492 doi: 10.1007/s10489-022-04216-6
    [214] Zeng D H, Yu Y, Oyama K. Audio-visual embedding for cross-modal music video retrieval through supervised deep CCA. In: Proceedings of the International Symposium on Multimedia (ISM). Taichung, China: IEEE, 2018. 143−150
    [215] Guo M, Zhou C H, Liu J H. Jointly learning of visual and auditory: A new approach for RS image and audio cross-modal retrieval. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2019, 12(11): 4644−4654 doi: 10.1109/JSTARS.2019.2949220
    [216] Chen G Y, Zhang D Y, Liu T, Du X Y. Self-lifting: A novel framework for unsupervised voice-face association learning. In: Proceedings of the International Conference on Multimedia Retrieval. Newark, USA: ACM, 2022. 527−535
    [217] Oncescu A M, Henriques J F, Zisserman A, Albanie S, Koepke A S. A sound approach: Using large language models to generate audio descriptions for egocentric text-audio retrieval. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Seoul, Republic of Korea: IEEE, 2024. 7300−7304
    [218] Lin J A, Liu D Z, Chen X K, Qu X Y, Yang X, Zhu J X, et al. Audio does matter: Importance-aware multi-granularity fusion for video moment retrieval. In: Proceedings of the 33rd ACM International Conference on Multimedia. Dublin Ireland: ACM, 2025. 6027−6036
    [219] Li X L, Hu D, Lu X Q. Image2song: Song retrieval via bridging image content and lyric words. In: Proceedings of the International Conference on Computer Vision (ICCV). Venice, Italy: IEEE, 2017. 5650−5650
    [220] Hong S, Im W, Yang H S. CBVMR: Content-based video-music retrieval using soft intra-modal structure constraint. In: Proceedings of the ACM International Conference on Multimedia Retrieval. Yokohama, Japan: ACM, 2018. 353−361
    [221] Nakatsuka T, Hamasaki M, Goto M. Content-based music-image retrieval using self-and cross-modal feature embedding memory. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA: IEEE, 2023. 2173−2183
    [222] Cheng X X, Zhu Z H, Li H X, Li Y W, Zou Y X. SSVMR: Saliency-based self-training for video-music retrieval. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece: IEEE, 2023. 1−5
    [223] Era Y, Togo R, Maeda K, Ogawa T, Haseyama M. Video-music retrieval with fine-grained cross-modal alignment. In: Proceedings of the International Conference on Image Processing (ICIP). Kuala Lumpur, Malaysia: IEEE, 2023. 2005−2009
    [224] McKee D, Salamon J, Sivic J, Russell B. Language-guided music recommendation for video via prompt analogies. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 14784−14793
    [225] Chen Y X, Du C, Zi Y F, Xiong S W, Lu X Q. Scale-aware adaptive refinement and cross-interaction for remote sensing audio-visual cross-modal retrieval. IEEE Transactions on Geosci-ence and Remote Sensing, 2024, 62: Article No. 4706914
    [226] Huang J H, Chen Y X, Xiong S W, Lu X Q. Cross-modal remote sensing image-audio retrieval with adaptive learning for aligning correlation. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: Article No. 4705213
    [227] Wu P, Su W S, He X T, Wang P, Zhang Y N. VarCMP: Adapting cross-modal pre-training models for video anomaly retrieval. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, Pennsylvania, USA: AAAI Press, 2025. 8423−8431
    [228] Tian Y P, Shi J, Li B C, Duan Z Y, Xu C L. Audio-visual event localization in unconstrained videos. In: Proceedings of the 15th European Conference on Computer Vision (ECCV). Munich, Germany: Springer, 2018. 252−268
    [229] Xuan H Y, Zhang Z Y, Chen S, Yang J, Yan Y. Cross-modal attention network for temporal inconsistent audio-visual event localization. In: Proceedings of the 34th AAAI Conference on Artificial Intelligence. New York, USA: AAAI Press, 2020. 279−286
    [230] Mahmud T, Marculescu D. AVE-CLIP: AudioCLIP-based multi-window temporal Transformer for audio visual event localization. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA: IEEE, 2023. 5147−5156
    [231] Bao P J, Yang W H, Ng B P, Er M H, Kot A C. Cross-modal label contrastive learning for unsupervised audio-visual event localization. In: Proceedings of the 37th AAAI Conference on Artificial Intelligence. Washington, USA: AAAI Press, 2023. 215−222
    [232] Zhou Z H, Zhou J X, Qian W, Tang S G, Chang X J, Guo D. Dense audio-visual event localization under cross-modal consistency and multi-temporal granularity collaboration. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, Pennsylvania, USA: AAAI Press, 2025. 10905−10913
    [233] Sun C, Chen M, Zhu C B, Zhang S, Lu P, Chen J C. Listen with seeing: Cross-modal contrastive learning for audio-visual event localization. IEEE Transactions on Multimedia, 2025, 27: 2650−2665 doi: 10.1109/TMM.2025.3535359
    [234] Liu L, Li S Y, Zhu Y Q. Audio-visual semantic graph network for audio-visual event localization. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2025. 23957−23966
    [235] Zhang P F, Wang J X, Wan M, Chang S J, Ding L H, Shi P. Multi-relation learning network for audio-visual event localization. Knowledge-Based Systems, 2025, 310: Article No. 112925 doi: 10.1016/j.knosys.2024.112925
    [236] Lin Y B, Tseng H Y, Lee H Y, Lin Y Y, Yang M H. Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. Virtual Event: Curran Associates Inc., 2021. Article No. 875

    Lin Y B, Tseng H Y, Lee H Y, Lin Y Y, Yang M H. Exploring cross-video and cross-modality signals for weakly-supervised audio-visual video parsing. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. Virtual Event: Curran Associates Inc., 2021. Article No. 875
    [237] Fan Y Y, Wu Y, Du B, Lin Y T. Revisit weakly-supervised audio-visual video parsing from the language perspective. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2023. Article No. 1767
    [238] Fu J, Gao J Y, Bao B K, Xu C S. Multimodal imbalance-aware gradient modulation for weakly-supervised audio-visual video parsing. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(6): 4843−4856 doi: 10.1109/TCSVT.2023.3337134
    [239] Geng T T, Wang T, Duan J M, Cong R M, Zheng F. Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 22942−22951
    [240] Yu J S, Cheng Y, Zhao R W, Feng R, Zhang Y J. MM-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video parsing. In: Proceedings of the 30th ACM International Conference on Multimedia. Lisboa, Portugal: ACM, 2022. 6241−6249
    [241] Lai Y H, Ebbers J, Wang Y C F, Germain F O, Jones M J, Chatterjee M. UWAV: Uncertainty-weighted weakly-supervised audio-visual video parsing. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2025. 13561−13570
    [242] Gao Y B, Sun X C, Lv G H, Yu D, Niu S J. Reinforced label denoising for weakly-supervised audio-visual video parsing. In: Proceedings of the 13th International Conference on Computational Visual Media. Hong Kong, China: Springer, 2025. 107−124
    [243] Gao J Y, Chen M Y, Xu C S. Learning probabilistic presence-absence evidence for weakly-supervised audio-visual event perception. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(6): 4787−4802 doi: 10.1109/TPAMI.2025.3546312
    [244] Zhao P C, Zhou J X, Zhao Y, Guo D, Chen Y X. Multimodal class-aware semantic enhancement network for audio-visual video parsing. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, Pennsylvania, USA: AAAI Press, 2025. 10448−10456
    [245] Senocak A, Oh T H, Kim J, Yang M H, Kweon I S. Learning to localize sound source in visual scenes. In: Proceedings of the Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA: IEEE, 2018. 4358−4366
    [246] Xuan H Y, Wu Z L, Yang J, Yan Y, Alameda-Pineda X. A proposal-based paradigm for self-supervised sound source localization in videos. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). New Orleans, USA: IEEE, 2022. 1019−1028
    [247] Liu J X, Ju C, Xie W D, Zhang Y. Exploiting transformation invariance and equivariance for self-supervised sound localisation. In: Proceedings of the 30th ACM International Conference on Multimedia. Lisboa, Portugal: ACM, 2022. 3742−3753
    [248] Qian R, Hu D, Dinkel H, Wu M Y, Xu N, Lin W Y. Multiple sound sources localization from coarse to fine. In: Proceedings of the 16th European Conference on Computer Vision (ECCV). Glasgow, UK: Springer, 2020. 292−308
    [249] Mo S T, Tian Y P. Audio-visual grouping network for sound localization from mixtures. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 10565−10574
    [250] Ryu H, Kim S, Chung J S, Senocak A. Seeing speech and sound: Distinguishing and locating audio sources in visual scenes. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2025. 13540−13549
    [251] Senocak A, Ryu H, Kim J, Oh T H, Pfister H, Chung J S. Sound source localization is all about cross-modal alignment. In: Proceedings of the International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 7743−7753
    [252] Senocak A, Ryu H, Kim J, Oh T H, Pfister H, Chung J S. Toward interactive sound source localization: Better align sight and sound! IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(9): 7643−7659 doi: 10.1109/TPAMI.2025.3573994
    [253] Kim I, Song Y, Park J, Kim W H, Kwak S. Improving sound source localization with joint slot attention on image and audio. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2025. 3121−3130
    [254] Liu T Y, Zhang P, Xiong J W, Li C Y, Huo Y, Huang W, et al. Less means more: Single stream audio-visual sound source localization via shared-parameter network. IEEE Transactions on Audio, Speech and Language Processing, 2025, 33: 4350−4360 doi: 10.1109/TASLPRO.2025.3619850
    [255] Zhou J X, Shen X Y, Wang J Y, Zhang J Y, Sun W X, Zhang J, et al. Audio-visual segmentation with semantics. International Journal of Computer Vision, 2025, 133(4): 1644−1664 doi: 10.1007/s11263-024-02261-x
    [256] Mo S T, Raj B. Weakly-supervised audio-visual segmentation. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2023. Article No. 753
    [257] Bhosale S, Yang H S, Kanojia D, Zhu X T. Leveraging foundation models for unsupervised audio-visual segmentation. arXiv preprint arXiv: 2309.06728, 2023.
    [258] Liu J X, Wang Y, Ju C, Ma C F, Zhang Y, Xie W D. Annotation-free audio-visual segmentation. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA: IEEE, 2024. 5592−5602
    [259] Chen T X, Tan Z T, Gong T, Chu Q, Wu Y, Liu B, et al. Bootstrapping audio-visual video segmentation by strengthening audio cues. IEEE Transactions on Circuits and Systems for Video Technology, 2025, 35(3): 2398−2409 doi: 10.1109/TCSVT.2024.3486344
    [260] Mao Y X, Zhang J, Xiang M C, Zhong Y R, Dai Y C. Multimodal variational auto-encoder based audio-visual segmentation. In: Proceedings of the International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 954−965
    [261] Liu C, Li P K, Yang L Y, Wang D D, Li L C, Yu X. Robust audio-visual segmentation via audio-guided visual convergent alignment. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2025. 28922−28931
    [262] Bhosale S, Yang H S, Kanojia D, Deng J K, Zhu X T. Unsupervised audio-visual segmentation with modality alignment. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, Pennsylvania, USA: AAAI Press, 2025. 15567−15575
    [263] Zhu Y, Li K, Yang Z X. Exploiting EfficientSAM and temporal coherence for audio-visual segmentation. IEEE Transactions on Multimedia, 2025, 27: 2999−3008 doi: 10.1109/TMM.2025.3557637
    [264] Xuan H Y, Liu T X, Dong W X, Li Z H, Chen S. X-STA: Cross-modal spatial-temporal alignment network for unified audio-visual segmentation. IEEE Signal Processing Letters, 2025, 32: 2883−2887 doi: 10.1109/LSP.2025.3586552
    [265] Lei Y Y, Cao H W. Audio-visual emotion recognition with preference learning based on intended and multi-modal perceived labels. IEEE Transactions on Affective Computing, 2023, 14(4): 2954−2969 doi: 10.1109/TAFFC.2023.3234777
    [266] Goncalves L, Leem S G, Lin W C, Sisman B, Busso C. Versatile audio-visual learning for emotion recognition. IEEE Transactions on Affective Computing, 2025, 16(1): 306−318 doi: 10.1109/TAFFC.2024.3433386
    [267] Hsu J H, Wu C H. Applying segment-level attention on bi-modal Transformer encoder for audio-visual emotion recognition. IEEE Transactions on Affective Computing, 2023, 14(4): 3231−3243 doi: 10.1109/TAFFC.2023.3258900
    [268] Mocanu B, Tapu R, Zaharia T. Multimodal emotion recognition using cross modal audio-video fusion with attention and deep metric learning. Image and Vision Computing, 2023, 133: Article No. 104676 doi: 10.1016/j.imavis.2023.104676
    [269] Praveen R G, Cardinal P, Granger E. Audio-visual fusion for emotion recognition in the valence-arousal space using joint cross-attention. IEEE Transactions on Biometrics, Behavior, and Identity Science, 2023, 5(3): 360−373 doi: 10.1109/TBIOM.2022.3233083
    [270] Praveen R G, Alam J, Charton E. United we stand, divided we fall: Handling weak complementarity for audio-visual emotion recognition in valence-arousal space. In: Proceedings of the Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Nashville, USA: IEEE, 2025. 5741−5751
    [271] Pan B, Hirota K, Jia Z Y, Zhao L H, Jin X M, Dai Y P. Multimodal emotion recognition based on feature selection and extreme learning machine in video clips. Journal of Ambient Intelligence and Humanized Computing, 2023, 14(3): 1903−1917 doi: 10.1007/s12652-021-03407-2
    [272] Shi T, Ge X R, Jose J M, Pugeault N, Henderson P. Detail-enhanced intra-and inter-modal interaction for audio-visual emotion recognition. In: Proceedings of the 27th International Conference on Pattern Recognition. Kolkata, India: Springer, 2024. 451−465
    [273] Ding S Y, Tang T B, Lu C K. Lightweight spatio-temporal convolutional neural network for audio-visual emotion recognition. IEEE Transactions on Affective Computing, 2025, 16(4): 2721−2734 doi: 10.1109/TAFFC.2025.3566773
    [274] Sharafi M, Yazdchi M, Rasti J. Audio-visual emotion recognition using K-means clustering and spatio-temporal CNN. In: Proceedings of the 6th International Conference on Pattern Recognition and Image Analysis (IPRIA). Qom, Islamic Republic of Iran: IEEE, 2023. 1−6
    [275] Wang A J, Fang Z J, Jiang X Y, Gao Y B, Cao G F, Ma S W. Depth estimation of multi-modal scene based on multi-scale modulation. In: Proceedings of the International Conference on Image Processing (ICIP). Kuala Lumpur, Malaysia: IEEE, 2023. 2795−2799
    [276] Gao R H, Chen C G, Al-Halah Z, Schissler C, Grauman K. VISUALECHOES: Spatial image representation learning through echolocation. In: Proceedings of the 16th European Conference on Computer Vision (ECCV). Glasgow, UK: Sprin-ger, 2020. 658−676
    [277] Liu X H, Hornauer S, Moutarde F, Lu J L. AVS-Net: Audio-visual scale net for self-supervised monocular metric depth estimation. arXiv preprint arXiv: 2412.01637, 2024.
    [278] Karaoguz C, Weisswange T H, Rodemann T, Wrede B, Rothkopf C A. Reward-based learning of optimal cue integration in audio and visual depth estimation. In: Proceedings of the International Conference on Advanced Robotics (ICAR). Tallinn, Estonia: IEEE, 2011. 389−395
    [279] Zhang C H, Tian K, Ni B L, Meng G F, Fan B, Zhang Z X, et al. Stereo depth estimation with echoes. In: Proceedings of the 17th European Conference on Computer Vision (ECCV). Tel Aviv, Israel: Springer, 2022. 496−513
    [280] Sun W, Qiu L L. Visual timing for sound source depth estimation in the wild. In: Proceedings of the International Conference on Intelligent Robots and Systems (IROS). Abu Dhabi, United Arab Emirates: IEEE, 2024. 12348−12355
    [281] Parida K K, Srivastava S, Sharma G. Beyond image to depth: Improving depth prediction using echoes. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2021. 8264−8273
    [282] Liang S S, Huang C, Tian Y P, Kumar A, Xu C L. AV-NeRF: Learning neural fields for real-world audio-visual scene synthesis. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2023. Article No. 1629
    [283] Purushwalkam S, Garí S V A, Ithapu V K, Schissler C, Robinson P, Gupta A, et al. Audio-visual floorplan reconstruction. In: Proceedings of the International Conference on Computer Vision (ICCV). Montreal, Canada: IEEE, 2021. 1163−1172
    [284] Wilson J, Rewkowski N, Lin M C, Fuchs H. Echo-reconstruction: Audio-augmented 3D scene reconstruction. arXiv preprint arXiv: 2110.02405, 2021.
    [285] Kim H, Remaggi L, Jackson P J B, Fazi F M, Hilton A. 3D room geometry reconstruction using audio-visual sensors. In: Proceedings of the International Conference on 3D Vision (3DV). Qingdao, China: IEEE, 2017. 621−629
    [286] Alawadh M. 3D Audio-visual Indoor Scene Reconstruction and Completion for Virtual Reality From a Single Image [Ph.D. dissertation], University of Southampton, UK, 2025.
    [287] Konno T, Nishida K, Itoyama K, Nakadai K. Audio-visual 3D reconstruction framework for dynamic scenes. In: Proceedings of the International Symposium on System Integration. Honolulu, USA: IEEE, 2020. 802−807
    [288] Younes A, Honerkamp D, Welschehold T, Valada A. Catch me if you hear me: Audio-visual navigation in complex unmapped environments with moving sounds. IEEE Robotics and Automation Letters, 2023, 8(2): 928−935 doi: 10.1109/LRA.2023.3234766
    [289] Shi Z B, Zhang L, Li L F, Shen Y. Towards audio-visual navigation in noisy environments: A large-scale benchmark dataset and an architecture considering multiple sound-sources. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, Pennsylvania, USA: AAAI Press, 2025. 14673−14680
    [290] Wang H C, Wang Y X, Zhong F W, Wu M D, Zhang J W, Wang Y Z, et al. Learning semantic-agnostic and spatial-aware representation for generalizable visual-audio navigation. IEEE Robotics and Automation Letters, 2023, 8(6): 3900−3907 doi: 10.1109/LRA.2023.3272518
    [291] Huang C G, Mees O, Zeng A, Burgard W. Audio visual language maps for robot navigation. In: Proceedings of the 18th International Symposium on Experimental Robotics (ISER). Chiang Mai, Thailand: Springer, 2023. 105−117
    [292] Chen C G, Al-Halah Z, Grauman K. Semantic audio-visual navigation. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2021. 15511−15520
    [293] Chen C G, Jain U, Schissler C, Gari S V A, Al-Halah Z, Ithapu V K, et al. SoundSpaces: Audio-visual navigation in 3D environments. In: Proceedings of the 16th European Conference on Computer Vision (ECCV). Glasgow, UK: Springer, 2020. 17−36
    [294] Kondoh H, Kanezaki A. Multi-goal audio-visual navigation using sound direction map. In: Proceedings of the International Conference on Intelligent Robots and Systems (IROS). Detroit, USA: IEEE, 2023. 5219−5226
    [295] Yu Y F, Huang W B, Sun F C, Chen C G, Wang Y K, Liu X H. Sound adversarial audio-visual navigation. In: Proceedings of the 10th International Conference on Learning Representations. Virtual Event: OpenReview.net, 2022.

    Yu Y F, Huang W B, Sun F C, Chen C G, Wang Y K, Liu X H. Sound adversarial audio-visual navigation. In: Proceedings of the 10th International Conference on Learning Representations. Virtual Event: OpenReview.net, 2022.
    [296] Liu X L, Paul S, Chatterjee M, Cherian A. CAVEN: An embodied conversational agent for efficient audio-visual navigation in noisy environments. In: Proceedings of the 38th AAAI Conference on Artificial Intelligence. Vancouver, Canada: AAAI Press, 2024. 3765−3773
    [297] 包希港, 周春来, 肖克晶, 覃飙. 视觉问答研究综述. 软件学报, 2021, 32(8): 2522−2544 doi: 10.13328/j.cnki.jos.006215

    Bao Xi-Gang, Zhou Chun-Lai, Xiao Ke-Jing, Qin Biao. Survey on visual question answering. Journal of Software, 2021, 32(8): 2522−2544 doi: 10.13328/j.cnki.jos.006215
    [298] Yang P C, Wang X, Duan X G, Chen H, Hou R Z, Jin C, et al. AVQA: A dataset for audio-visual question answering on videos. In: Proceedings of the 30th ACM International Conference on Multimedia. Lisboa, Portugal: ACM, 2022. 3480−3491
    [299] Zhao Y H, Xi W, Bai G R, Liu X H, Zhao J Z. Heterogeneous interactive graph network for audio-visual question answering. Knowledge-Based Systems, 2024, 300: Article No. 112165 doi: 10.1016/j.knosys.2024.112165
    [300] Li G Y, Hou W X, Hu D. Progressive spatio-temporal perception for audio-visual question answering. In: Proceedings of the 31st ACM International Conference on Multimedia. Ottawa, Canada: ACM, 2023. 7808−7816
    [301] Li Z B, Zhou J X, Zhang J, Tang S G, Li K, Guo D. Patch-level sounding object tracking for audio-visual question answering. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, Pennsylvania, USA: AAAI Press, 2025. 5075−5083
    [302] Li Z B, Guo D, Zhou J X, Zhang J, Wang M. Object-aware adaptive-positivity learning for audio-visual question answering. In: Proceedings of the 38th AAAI Conference on Artificial Intelligence. Vancouver, Canada: AAAI Press, 2024. 3306−3314
    [303] Li L J, Jin T, Lin W, Jiang H, Pan W W, Wang J, et al. Multi-granularity relational attention network for audio-visual question answering. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(8): 7080−7094 doi: 10.1109/TCSVT.2023.3264524
    [304] Jiang Y Y, Yin J Q. Target-aware spatio-temporal reasoning via answering questions in dynamic audio-visual scenarios. In: Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023. Singapore: ACL, 2023. 9399−9409
    [305] Ma J, Hu M, Wang P H, Sun W C, Song L Y, Pei H B, et al. Look, listen, and answer: Overcoming biases for audio-visual question answering. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. Vancouver, Canada: Curran Associates Inc., 2024. Article No. 302
    [306] Lao M R, Pu N, Liu Y, He K, Bakker E M, Lew M S. COCA: Collaborative causal regularization for audio-visual question answering. In: Proceedings of the 37th AAAI Conference on Artificial Intelligence. Washington, USA: AAAI Press, 2023. 12995−13003
    [307] Li G Y, Du H H, Hu D. Boosting audio visual question answering via key semantic-aware cues. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024. 5997−6005
    [308] Pei B Q, Huang Y F, Chen G, Xu J L, Wang Y L, Wang L M, et al. Guiding audio-visual question answering with collective question reasoning. International Journal of Computer Vision, 2025, 133(10): 6912−6929 doi: 10.1007/s11263-025-02510-7
    [309] Alamri H, Cartillier V, Das A, Wang J, Cherian A, Essa I, et al. Audio visual scene-aware dialog. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach, USA: IEEE, 2019. 7550−7559
    [310] Heo Y, Kang S, Seo J. Natural-language-driven multimodal representation learning for audio-visual scene-aware dialog system. Sensors, 2023, 23(18): Article No. 7875 doi: 10.3390/s23187875
    [311] Chen Z, Liu H C, Wang Y. DialogMCF: Multimodal context flow for audio visual scene-aware dialog. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024, 32: 753−764 doi: 10.1109/TASLP.2023.3284511
    [312] Li Z K, Li Z J, Zhang J C, Feng Y, Zhou J. Bridging text and video: A universal multimodal Transformer for audio-visual scene-aware dialog. IEEE/ACM Transactions on Audio, Spe-ech, and Language Processing, 2021, 29: 2476−2483 doi: 10.1109/TASLP.2021.3065823
    [313] Ye M C, You Q Z, Ma F L. QUALIFIER: Question-guided self-attentive multimodal fusion network for audio visual scene-aware dialog. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Waikoloa, USA: IEEE, 2022. 2503−2511
    [314] Park S J, Kim Y, Rha H, Godiva B, Ro Y M. AV-EmoDialog: Chat with audio-visual users leveraging emotional cues. arXiv preprint arXiv: 2412.17292, 2024.
    [315] Shah A, Geng S J, Gao P, Cherian A, Hori T, Marks T K, et al. Audio-visual scene-aware dialog and reasoning using audio-visual Transformers with joint student-teacher learning. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Singapore: IEEE, 2022. 7732−7736
    [316] Ye Q L, Yu Z T, Liu X. Answering diverse questions via text attached with key audio-visual clues. arXiv preprint arXiv: 2403.06679, 2024.
    [317] 侯静怡, 齐雅昀, 吴心筱, 贾云得. 跨语言知识蒸馏的视频中文字幕生成. 计算机学报, 2021, 44(9): 1907−1921 doi: 10.11897/SP.J.1016.2021.01907

    Hou Jing-Yi, Qi Ya-Yun, Wu Xin-Xiao, Jia Yun-De. Cross-lingual knowledge distillation for Chinese video captioning. Chinese Journal of Computers, 2021, 44(9): 1907−1921 doi: 10.11897/SP.J.1016.2021.01907
    [318] Shen X Y, Li D, Zhou J X, Qin Z, He B W, Han X D, et al. Fine-grained audible video description. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 10585−10596
    [319] Xie Z Y, Yang Y, Yu Y K, Liu Y. Exploring audio-visual concepts for dense video captioning. In: Proceedings of the International Conference on Digital Society and Intelligent Systems (DSInS). Sydney, Australia: IEEE, 2024: 334−338
    [320] Çayli Ö, Liu X B, Kiliç V, Wang W W. Knowledge distillation for efficient audio-visual video captioning. In: Proceedings of the European Signal Processing Conference (EUSIPCO). Helsinki, Finland: IEEE, 2023. 745−749
    [321] Xie Y L, Niu J J, Zhang Y, Ren F. Global-shared text representation based multi-stage fusion Transformer network for multi-modal dense video captioning. IEEE Transactions on Multimedia, 2024, 26: 3164−3179 doi: 10.1109/TMM.2023.3307972
    [322] Yang A, Nagrani A, Seo P H, Miech A, Pont-Tuset J, Laptev I, et al. Vid2Seq: Large-scale pretraining of a visual language model for dense video captioning. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 10714−10726
    [323] Han S X, Liu J, Zhang J Y M, Gong P Z, Zhang X L, He H H. Lightweight dense video captioning with cross-modal attention and knowledge enhanced unbiased scene graph. Complex & Intelligent Systems, 2023, 9(5): 4995−5012 doi: 10.1007/s40747-023-00998-5
    [324] Kim J, Shin J, Kim J. AVCap: Leveraging audio-visual features as text tokens for captioning. In: Proceedings of the 25th Annual Conference of the International Speech Communication Association. Kos, Greece: ISCA, 2024.
    [325] AlSuwat M, Al-Shareef S, AlGhamdi M. Audio-visual self-supervised representation learning: A survey. Neurocomputing, 2025, 634: Article No. 129750 doi: 10.1016/j.neucom.2025.129750
    [326] Arandjelović R, Zisserman A. Objects that sound. In: Proceedings of the 15th European Conference on Computer Vision (ECCV). Munich, Germany: Springer, 2018. 451−466
    [327] Owens A, Efros A A. Audio-visual scene analysis with self-supervised multisensory features. In: Proceedings of the 15th European Conference on Computer Vision (ECCV). Munich, Germany: Springer, 2018. 639−658
    [328] Korbar B, Tran D, Torresani L. Cooperative learning of audio and video models from self-supervised synchronization. In: Proceedings of the 32nd International Conference on Neural Information Processing Systems. Montréal, Canada: Curran Associates Inc., 2018. 7774−7785
    [329] Huang C, Tian Y P, Kumar A, Xu C L. Egocentric audio-visual object localization. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 22910−22921
    [330] Sun W X, Zhang J Y, Wang J Y, Liu Z Y, Zhong Y R, Feng T P, et al. Learning audio-visual source localization via false negative aware contrastive learning. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE, 2023. 6420−6429
    [331] Morgado P, Vasconcelos N, Misra I. Audio-visual instance discrimination with cross-modal agreement. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2021. 12470−12481
    [332] Ma S, Zeng Z Y, McDuff D, Song Y L. Active contrastive learning of audio-visual video representations. In: Proceedings of the International Conference on Learning Representations (ICLR). Vienna, Austria: OpenReview, 2021. 1−19
    [333] Xuan H Y, Xu Y H, Chen S, Wu Z L, Yang J, Yan Y, et al. Active contrastive set mining for robust audio-visual instance discrimination. In: Proceedings of the 31st International Joint Conference on Artificial Intelligence. Vienna, Austria: IJCAI, 2022. 3643−3649
    [334] Xuan H Y, Wu Z L, Yang J, Jiang B, Luo L, Alameda-Pineda X, et al. Robust audio-visual contrastive learning for proposal-based self-supervised sound source localization in videos. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(7): 4896−4907 doi: 10.1109/TPAMI.2024.3363508
    [335] Li Z J, Zhao B, Yuan Y. Bio-inspired audiovisual multi-representation integration via self-supervised learning. In: Proceedings of the 31st ACM International Conference on Multimedia. Ottawa, Canada: ACM, 2023. 3755−3764
    [336] Sarkar P, Etemad A. Self-supervised audio-visual representation learning with relaxed cross-modal synchronicity. In: Proceedings of the 37th AAAI Conference on Artificial Intelligence. Washington, USA: AAAI Press, 2023. 9723−9732
    [337] Jenni S, Black A, Collomosse J. Audio-visual contrastive learning with temporal self-supervision. In: Proceedings of the 37th AAAI Conference on Artificial Intelligence. Washington, USA: AAAI Press, 2023. 7996−8004
    [338] Zhang J X, Wan G S, Gao J Q, Ling Z H. Audio-visual representation learning via knowledge distillation from speech foundation models. Pattern Recognition, 2025, 162: Article No. 111432 doi: 10.1016/j.patcog.2025.111432
    [339] Zhu B Q, Wang C J, Xu K L, Feng D W, Zhou Z M, Zhu X Q. Learning incremental audio-visual representation for continual multimodal understanding. Knowledge-Based Systems, 2024, 304: Article No. 112513 doi: 10.1016/j.knosys.2024.112513
    [340] Zuo Y K, Yao H T, Zhuang L S, Xu C S. Hierarchical augmentation and distillation for class incremental audio-visual video recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(11): 7348−7362 doi: 10.1109/TPAMI.2024.3387946
    [341] Cui M, Yue X H, Qian X Y, Zhao J Z, Liu H H, Liu X B, et al. Audio-visual class-incremental learning for fish feeding intensity assessment in aquaculture. arXiv preprint arXiv: 2504.15171, 2025.
    [342] Mo S T, Pian W G, Tian Y P. Class-incremental grouping network for continual audio-visual learning. In: Proceedings of the International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 7754−7764
    [343] Pian W G, Mo S T, Guo Y H, Tian Y P. Audio-visual class-incremental learning. In: Proceedings of the International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 7765−7777
    [344] Yue X H, Zhang X Y, Chen Y M, Zhang C W, Lao M R, Zhuang H P, et al. MMAL: Multi-modal analytic learning for exemplar-free audio-visual class incremental tasks. In: Proceedings of the 32nd ACM International Conference on Multimedia. Melbourne, Australia: ACM, 2024. 2428−2437
    [345] Cui Y W, Liu L, Yu Z T, Huang G J, Hong X P. Few-shot audio-visual class-incremental learning with temporal prompting and regularization. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, Pennsylvania, USA: AAAI Press, 2025. 16118−16126
    [346] 张鲁宁, 左信, 刘建伟. 零样本学习研究进展. 自动化学报, 2020, 46(1): 1−23

    Zhang Lu-Ning, Zuo Xin, Liu Jian-Wei. Research and development on zero-shot learning. Acta Automatica Sinica, 2020, 46(1): 1−23
    [347] Parida K K, Matiyali N, Guha T, Sharma G. Coordinated joint multimodal embeddings for generalized audio-visual zero-shot classification and retrieval of videos. In: Proceedings of the Winter Conference on Applications of Computer Vision (WACV). Snowmass, USA: IEEE, 2020. 3240−3249
    [348] Li Y P, Luo Y, Du B. Audio-visual generalized zero-shot learning based on variational information bottleneck. In: Proceedings of the International Conference on Multimedia and Expo (ICME). Brisbane, Australia: IEEE, 2023. 450−455
    [349] Zheng Q C, Hong J, Farazi M. A generative approach to audio-visual generalized zero-shot learning: Combining contrastive and discriminative techniques. In: Proceedings of the International Joint Conference on Neural Networks (IJCNN). Gold Coast, Australia: IEEE, 2023. 1−8
    [350] Mo S T, Morgado P. Audio-visual generalized zero-shot learning the easy way. In: Proceedings of the 18th European Conference on Computer Vision (ECCV). Milan, Italy: Springer, 2024. 377−395
    [351] Dong Y J, Chen S M, Duan B W, Ding W P, Wang Y S, You X G. Object-aware image augmentation for audio-visual zero-shot learning. IEEE Transactions on Emerging Topics in Computational Intelligence, 2025, 9(6): 4106−4118 doi: 10.1109/TETCI.2024.3485624
    [352] Li W R, Wang P H, Xiong R Q, Fan X P. Spiking tucker fusion Transformer for audio-visual zero-shot learning. IEEE Transactions on Image Processing, 2024, 33: 4840−4852 doi: 10.1109/TIP.2024.3430080
    [353] Yang Z, Li W R, Hou J X, Cheng G H. Multi-modal spiking tensor regression network for audio-visual zero-shot learning. Neurocomputing, 2025, 629: Article No. 129636 doi: 10.1016/j.neucom.2025.129636
    [354] Li W R, Ma Z Y, Deng L J, Man H Y, Fan X P. Modality-fusion spiking Transformer network for audio-visual zero-shot learning. In: Proceedings of the International Conference on Multimedia and Expo (ICME). Brisbane, Australia: IEEE, 2023. 426−431
    [355] Zhang K W, Zhao K C, Tian Y N. Temporal-semantic aligning and reasoning Transformer for audio-visual zero-shot learning. Mathematics, 2024, 12(14): Article No. 2200 doi: 10.3390/math12142200
    [356] Li W R, Wang P H, Wang X T, Zuo W M, Fan X P, Tian Y H. Multi-timescale motion-decoupled spiking Transformer for audio-visual zero-shot learning. IEEE Transactions on Circuits and Systems for Video Technology, 2025, 35(11): 10772−10786 doi: 10.1109/TCSVT.2025.3574499
    [357] Larsen-Freeman D. Transfer of learning transformed. Language Learning, 2013, 63(S1): 107−129 doi: 10.1111/j.1467-9922.2012.00740.x
    [358] Hajavi A, Etemad A. Audio representation learning by distilling video as privileged information. IEEE Transactions on Artificial Intelligence, 2024, 5(1): 446−456 doi: 10.1109/TAI.2023.3243596
    [359] Yun H, Na J, Kim G. Dense 2D-3D indoor prediction with sound via aligned cross-modal distillation. In: Proceedings of the International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 7829−7838
    [360] Kim J U, Kim S T. Towards robust audio-based vehicle detection via importance-aware audio-visual learning. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Rhodes Island, Greece: IEEE, 2023. 1−5
    [361] Chen J Y, Wang W G, Liu S, Li H S, Yang Y. Omnidirectional information gathering for knowledge transfer-based audio-visual navigation. In: Proceedings of the International Conference on Computer Vision (ICCV). Paris, France: IEEE, 2023. 10959−10969
    [362] Chen M C, Zhang B M, Han Z B, Jiang W Y, Wang Y M, Feng S, et al. Test-time selective adaptation for uni-modal distribution shift in multi-modal data. In: Proceedings of the 42nd International Conference on Machine Learning. Vancouver, Canada: OpenReview.net, 2025. 1−10
    [363] Duan H Y, Xia Y, Zhou M Z, Tang L, Zhu J M, Zhao Z. Cross-modal prompts: Adapting large pre-trained models for audio-visual downstream tasks. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2023. Article No. 2445
    [364] Mo S T, Morgado P. A unified audio-visual learning framework for localization, separation, and recognition. In: Proceedings of the 40th International Conference on Machine Learning. Honolulu, USA: Curran Associates Inc., 2023. Article No.1041
    [365] 蔡朝阳, 周黎婧. 认知心理学视角下学习迁移与能力生成研究. 教育进展, 2025, 15(1): 1226−1237

    Cai Chao-Yang, Zhou Li-Jing. Study on learning transfer and ability generation from the perspective of cognitive psychology. Advances in Education, 2025, 15(1): 1226−1237
    [366] 陈光, 郭军. 大语言模型时代的人工智能: 技术内涵、行业应用与挑战. 北京邮电大学学报, 2024, 47(4): 20−28

    Chen Guang, Guo Jun. Artificial intelligence in the era of large language models: Technical significance, industry applications, and challenges. Journal of Beijing University of Posts and Telecommunications, 2024, 47(4): 20−28
    [367] Peng P Y, Huang P Y, Li S W, Mohamed A, Harwath D. VoiceCraft: Zero-shot speech editing and text-to-speech in the wild. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Bangkok, Thailand: ACL, 2024. 12442−12462
    [368] Andreyev A. Quantization for OpenAI's whisper models: A comparative analysis. arXiv preprint arXiv: 2503.09905, 2025.
    [369] Su Y, Bai J S, Xu Q S, Xu K L, Dou Y. Audio-language models for audio-centric tasks: A systematic survey. arXiv preprint arXiv: 2501.15177, 2025.
    [370] Jiang F, Lin Z Y, Bu F, Du Y H, Wang B Y, Li H Z. S2S-arena, evaluating speech2speech protocols on instruction following with paralinguistic information. arXiv preprint arXiv: 2503.05085, 2025.
    [371] He H R, Zhang Y, Lin L, Xu Z W, Pan L. Pre-trained video generative models as world simulators. arXiv preprint arXiv: 2502.07825, 2025.
    [372] Wang Y M, Liu X Y, Pang W, Ma L, Yuan S, Debevec P, et al. Survey of video diffusion models: Foundations, implementations, and applications. arXiv preprint arXiv: 2504.16081, 2025.
    [373] Zhang Y B, Wei Y X, Lin X H, Hui Z, Ren P R, Xie X S, et al. VideoElevator: Elevating video generation quality with versatile text-to-image diffusion models. In: Proceedings of the 39th AAAI Conference on Artificial Intelligence. Philadelphia, Pennsylvania: AAAI Press, 2025. 10266−10274
    [374] Wang Y L, Deng Y J, Zheng Y J, Chattopadhyay P, Wang L P. Vision Transformers for image classification: A comparative survey. Technologies, 2025, 13(1): Article No. 32 doi: 10.3390/technologies13010032
    [375] Yu Z P, Ananiadou S. Understanding multimodal LLMs: The mechanistic interpretability of Llava in visual question answering. arXiv preprint arXiv: 2411.10950, 2024.
    [376] Zhan J, Dai J Q, Ye J S, Zhou Y H, Zhang D, Liu Z G, et al. AnyGPT: Unified multimodal LLM with discrete sequence modeling. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Bangkok, Thailand: ACL, 2024. 9637−9662
    [377] Wu S Q, Fei H, Qu L G, Ji W, Chua T S. NExT-GPT: Any-to-any multimodal LLM. In: Proceedings of the 41st International Conference on Machine Learning. Vienna, Austria: JMLR.org, 2024. Article No. 2187
    [378] Chen F L, Han M L, Zhao H Z, Zhang Q Y, Shi J, Xu S, et al. X-LLM: Bootstrapping advanced large language models by treating multi-modalities as foreign languages. arXiv preprint arXiv: 2305.04160, 2023.
    [379] Fu C Y, Lin H J, Long Z W, Shen Y H, Dai Y H, Zhao M, et al. VITA: Towards open-source interactive Omni multimodal LLM. arXiv preprint arXiv: 2408.05211, 2024.
    [380] Akbari H, Yuan L Z, Qian R, Chuang W H, Chang S F, Cui Y, et al. VATT: Transformers for multimodal self-supervised learning from raw video, audio and text. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. Virtual Event: Curran Associates Inc., 2021. Article No. 1853

    Akbari H, Yuan L Z, Qian R, Chuang W H, Chang S F, Cui Y, et al. VATT: Transformers for multimodal self-supervised learning from raw video, audio and text. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. Virtual Event: Curran Associates Inc., 2021. Article No. 1853
    [381] Lin Y B, Bertasius G. Siamese vision Transformers are scalable audio-visual learners. In: Proceedings of the 18th European Conference on Computer Vision (ECCV). Milan, Italy: Springer, 2024. 303−321
    [382] Huang P Y, Sharma V, Xu H, Ryali C, Fan H Q, Li Y H, et al. MAViL: Masked audio-video learners. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. New Orleans, USA: Curran Associates Inc., 2023. Article No. 894
    [383] Chowdhury S, Nag S, Dasgupta S, Chen J, Elhoseiny M, Gao R H, et al. MEERKAT: Audio-visual large language model for grounding in space and time. In: Proceedings of the 18th European Conference on Computer Vision (ECCV). Milan, Italy: Springer, 2024. 52−70
    [384] Chen K, Gou Y H, Huang R H, Liu Z L, Tan D X, Xu J, et al. EMOVA: Empowering language models to see, hear and speak with vivid emotions. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2025. 5455−5466
    [385] Sun G Z, Yu W Y, Tang C L, Chen X Z, Tan T, Li W, et al. video-SALMONN: Speech-enhanced audio-visual large language models. In: Proceedings of the 41st International Conference on Machine Learning. Vienna, Austria: Curran Associates Inc., 2024. Article No. 1921
    [386] Chu Y Q, Liao L Z, Zhou Z Y, Ngo C W, Hong R H. Towards multimodal emotional support conversation systems. IEEE Transactions on Multimedia, 2025, 27: 8276−8287 doi: 10.1109/TMM.2025.3604951
    [387] Sun G Z, Yang Y D, Zhuang J M, Tang C L, Li Y X, Li W, et al. video-SALMONN-o1: Reasoning-enhanced audio-visual large language model. In: Proceedings of the 42nd International Conference on Machine Learning (ICML). Vancouver, Canada: OpenReview.net, 2025. 1−10
    [388] Cappellazzo U, Kim M, Chen H J, Ma P C, Petridis S, Falavigna D, et al. Large language models are strong audio-visual speech recognition learners. In: Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP). Hyderabad, India: IEEE, 2025. 1−5
    [389] Mao Y R, Ge Y H, Fan Y J, Xu W Y, Mi Y, Hu Z H, et al. A survey on LoRA of large language models. Frontiers of Computer Science, 2025, 19(7): Article No. 197605 doi: 10.1007/s11704-024-40663-9
    [390] Du H H, Li G Y, Zhou C, Zhang C J, Zhao A, Hu D. Crab: A unified audio-visual scene understanding model with explicit cooperation. In: Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). Nashville, USA: IEEE, 2025. 18804−18814
    [391] Jin D, Zhou Y H, Zhou J X, Ma J Q, Guo R H, Guo D. SimToken: A simple baseline for referring audio-visual segmentation. arXiv preprint arXiv: 2509.17537, 2025.
    [392] Cappellazzo U, Kim M, Petridis S. Adaptive audio-visual speech recognition via Matryoshka-based multimodal LLMs. arXiv preprint arXiv: 2503.06362, 2025.
    [393] Tang C J, Li Y X, Yang Y D, Zhuang J M, Sun G Z, Li W, et al. video-SALMONN 2: Captioning-enhanced audio-visual large language models. In: Proceedings of the 14th International Conference on Learning Representations. Rio de Janeiro, Brazil: OpenReview.net, 2025.
    [394] Huang G J, Lin W L, Liu L. Content-aware efficient learner for audio-visual emotion recognition. In: Proceedings of the 16th International Conference on Social Robotics. Shenzhen, China: Springer, 2024. 31−40
  • 加载中
图(8) / 表(2)
计量
  • 文章访问数:  1141
  • HTML全文浏览量:  2409
  • PDF下载量:  45
  • 被引次数: 0
出版历程
  • 收稿日期:  2025-07-24
  • 录用日期:  2025-12-31
  • 网络出版日期:  2026-03-13

目录

    /

    返回文章
    返回