• 中文核心
  • EI
  • 中国科技核心
  • Scopus
  • CSCD
  • 英国科学文摘

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

面向心理健康的生成式自动问答研究综述

张丽园 刘德喜 彭文忠 万齐智 陈启 方玉明

张丽园, 刘德喜, 彭文忠, 万齐智, 陈启, 方玉明. 面向心理健康的生成式自动问答研究综述. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250731
引用本文: 张丽园, 刘德喜, 彭文忠, 万齐智, 陈启, 方玉明. 面向心理健康的生成式自动问答研究综述. 自动化学报, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250731
Zhang Li-Yuan, Liu De-Xi, Peng Wen-Zhong, Wan Qi-Zhi, Chen Qi, Fang Yu-Ming. A survey on generative automatic question answering for mental health. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250731
Citation: Zhang Li-Yuan, Liu De-Xi, Peng Wen-Zhong, Wan Qi-Zhi, Chen Qi, Fang Yu-Ming. A survey on generative automatic question answering for mental health. Acta Automatica Sinica, xxxx, xx(x): x−xx doi: 10.16383/j.aas.c250731

面向心理健康的生成式自动问答研究综述

doi: 10.16383/j.aas.c250731 cstr: 32138.14.j.aas.c250731
基金项目: 国家自然科学基金(62272206, 62272205, 62562033), 江西省自然科学基金(20242BAB25119, 20262BAC240389), 江西省教育厅科技项目(GJJ2502406), 豫章师范学院江西省人才战略发展研究基地重点项目(YZRCZD-25-04)资助
详细信息
    作者简介:

    张丽园:江西财经大学计算机与人工智能学院博士研究生. 主要研究方向为自然语言处理和心理健康自动问答. E-mail: 2202110072@stu.jxufe.edu.cn

    刘德喜:江西财经大学计算机与人工智能学院教授. 主要研究方向为自然语言处理和信息检索. 本文通信作者. E-mail: liudexi@jxufe.edu.cn

    彭文忠:江西财经大学计算机与人工智能学院讲师. 主要研究方向为信息抽取, 自然语言处理和数据挖掘. E-mail: pengwenzhong@jxufe.edu.cn

    万齐智:江西财经大学计算机与人工智能学院讲师. 主要研究方向为人工智能, 深度学习, 信息抽取, 自然语言处理和文本数据挖掘. E-mail: wanqizhi@jxufe.edu.cn

    陈启:江西财经大学计算机与人工智能学院博士研究生. 主要研究方向为方面级情感分析和心理健康支持. E-mail: 2202210079@stu.jxufe.edu.cn

    方玉明:江西财经大学计算机与人工智能学院教授. 主要研究方向为计算机视觉, 多媒体信号处理和视觉质量评估. E-mail: fa0001ng@e.ntu.edu.sg

  • 中图分类号: Y

A Survey on Generative Automatic Question Answering for Mental Health

Funds: Supported by National Natural Science Foundation of China (62272206, 62272205, 62562033), Natural Science Foundation of Jiangxi Province (20242BAB25119, 20262BAC240389), Science and Technology Project of the Education Department of Jiangxi Province (GJJ2502406), and Key Project of the Jiangxi Provincial Research Base for Talent Strategy Development at Yuzhang Normal University (YZRCZD-25-04)
More Information
    Author Bio:

    ZHANG Li-Yuan Ph.D. candidate at the School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics. Her research interests include natural language processing and automatic question answering for mental health

    LIU De-Xi Professor at the School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics. His research interests include natural language processing and information retrieval. Corresponding author of this paper

    PENG Wen-Zhong Lecturer at the School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics. His research interests include information extraction, natural language processing, and data mining

    WAN Qi-Zhi Lecturer at the School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics. His research interests include artificial intelligence, deep learning, information extraction, natural language processing, and text data mining

    CHEN Qi Ph.D. candidate at the School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics. His research interests include aspect-based sentiment analysis and mental health support

    FANG Yu-Ming Professor at the School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics. His research interests include computer vision, multimedia signal processing, and visual quality assessment

  • 摘要: 心理健康自动问答通过自动化或人机协同模式, 为心理困惑或情绪困扰用户提供及时可靠的信息支持与情感陪伴. 该技术弥补了传统心理健康服务在覆盖范围与响应效率上的不足, 近年成为研究热点. 然而, 现有综述主要聚焦于开放域问答或广泛的情感支持任务, 尚未有针对心理健康自动问答进行系统性梳理的综述文献. 为此, 对该领域相关研究进行全面回顾和总结. 首先, 从技术视角系统梳理心理健康自动问答的发展历史, 并分别从功能目标与心理问题类别两个维度对现有研究进行分类与评述. 其次, 聚焦文本模态, 根据技术方法将生成式心理健康自动问答任务划分为“LLMs之前”与“基于LLMs” 两大方向, 展开深入综述, 分析各类方法的原理与优劣. 在此基础上, 进一步拓展至基于LLMs的多模态相关研究, 进行初步分析与梳理. 接着, 介绍中英文心理健康问答数据集及其特点, 并对多种评估方法进行详细总结. 最后, 探讨该领域当前面临的关键挑战以及应对与优化机制, 并对未来研究方向进行展望.
    1)  11https://chat.openai.com2https://gemini.google.com/3https://llama.meta.com/4https://chat.deepseek.com/5https://chatglm.cn6https://tongyi.aliyun.com/7https://www.doubao.com/chat/8https://www.baichuan-ai.com/home
    2)  2https://gemini.google.com/
    3)  3https://llama.meta.com/
    4)  4https://chat.deepseek.com/
    5)  5https://chatglm.cn
    6)  6https://tongyi.aliyun.com/
    7)  7https://www.doubao.com/chat/
    8)  8https://www.baichuan-ai.com/home
    9)  99https://github.com/chatopera/efaqa-corpus-zh10https://www.xinli001.com/qa?source=pc-home
    10)  10https://www.xinli001.com/qa?source=pc-home
    11)  1111http://dxy.com/12http://www.xywy.com/
    12)  12http://www.xywy.com/
    13)  1313https://csn.cancer.org/
    14)  1414https://huggingface.co/datasets/heliosbrahma/mental_health_chatbot_dataset.
  • 图  1  组织结构图

    Fig.  1  Organizational structure diagram

    图  2  心理健康问答示例

    Fig.  2  Example of mental health QA

    图  3  心理健康自动问答发展历史脉络图

    Fig.  3  Evolutionary roadmap of mental health automatic QA

    表  1  心理健康问答方法对比分析

    Table  1  Comparative analysis of mental health QA methods

    分类 定义 优点 缺点 代表性模型/方法
    规则类 基于规则和模板进行操作, 将用户的问题与预定义的回答进行匹配 执行效率高、响应迅速、准确率高 受限于预设规则范围, 规则的构建与维护成本较高、扩展性和适应性差 ELIZA[22]、Parry[23]
    检索类 依据预先构建的知识库或文本语料, 通过信息检索技术实现问答匹配 效率高、准确性和可靠性高 受限于预构建的知识库、灵活性不足、泛化能力弱、无法回答预构知识库之外的问题 心理文档检测框架[26]、PAL[27]、CRISISbot[28]
    生成类 从海量文本语料中学习语言模式与结构, 通过语言模型直接生成回答 语言创造力和语境适应性强、可输出更为多样且贴合对话情境的回答 受限于模型结构、参数规模和训练数据质量, 上下文语义理解和逻辑推理能力尚不完善, 易生成不准确或不合逻辑的“幻觉”内容 ESC[30]、MISC[33]、EmoDS[42]、知识增强生成模型[43]、PARNER[44]、VA[45]、Supporter[46]
    检索+生成类 结合检索式和生成式两种方法, 发挥二者的优势 生成回答的准确性、相关性和专业性更强 依赖于外部知识库质量, 仍可能生成不准确或不合逻辑的“幻觉”内容; 在复杂场景下的可靠性、泛化能力存在局限 检索—判别—重写框架[34]、KEMI[35]、检索增强生成模型[47]、DialogueCRN[48]、Iseeq[49]、CoQUAD[50]
    LLMs类 使用LLMs生成回答 强大的上下文理解能力. 生成回答流畅性高、专业性强、接近人类水平 “幻觉”问题、高度依赖提示工程且存在不可解释性, 深层情感和复杂心理状态的理解与推理能力有限, 存在隐私安全与偏见等伦理风险 ChatCounselor[36]、SuDoSys[38]、PsyRA[40]、PsyDT[51]、CPsyCoun[52]、PsyAdvisor[53]、CBT-LLM[54]、SoulChat[55]、Psy-LLM[56]、MentaLLaMA[57]、PATIENT-Ψ[58]、Mental-LLM[59]
    多模态LLMs类 使用多模态数据作为输入, 基于LLMs生成回答 融合了多模态信息、情感感知全面、准确 数据隐私风险、多模态对齐问题、算力成本高、延迟高 SMES[41]、FIRES[60]、Empatheia[61]、MultiMood[62]、DELTA[63]
    注: “代表性模型/方法”列中, 原文献已命名的模型或框架沿用原名, 未命名则以核心方法代称.
    下载: 导出CSV

    表  2  心理健康自动问答任务分类

    Table  2  Mental health automatic QA task classification

    分类依据 分类 代表性模型/方法
    根据功能目标 诊断 PsyGUARD[64]、BERT EnsembleCNN框架[65]、大模型法[66]、知识增强法[67]、MDD[68]、POST[69]、智能诊断机器人[70]
    干预 Motibot[73]、自我导向系统[75]、Popbots[77]、多维决策支持系统[78]
    监测 Woebot[79]、LISSA[81]、共情—表露方法[82]、BLISS[83]、Fedtherapist[84]
    咨询 PsyDT[51]、CPsyCoun[52]、FIRES[60]、Empatheia[61]、MultiMood[62]、DELTA[63]、SOPHIE[85]、Empathic-Chatbot[86]、Multi-Persona-Chatbot[87]、PsyChatbot[88]、Aroha[89]、Cactus[90]、Psy-Insight[91]、视觉—文本融合框架[92]、IntelliCare[93]
    多功能混合 ESC[30]、SMES[41]、PsyAdvisor[53]、DELTA[63]、Woebot[94]
    根据心理问题类别 针对广泛心理问题 ESC[30]、PsyDT[51]、CPsyCoun[52]、PsyAdvisor[53]、CBT-LLM[54]、SoulChat[55]、FIRES[60]、Empatheia[61]、MultiMood[62]、DELTA[63]、Psy-Insight[91]、IntelliCare[93]
    针对抑郁与焦虑 知识增强法[67]、PsyChatbot[88]、Mindful Diary[98]
    针对自杀风险 PsyGUARD[64]、BERT Ensemble-CNN框架[65]、LLM诊断法[66]、自杀风险评测框架[100]
    针对压力管理 Chat Counselor[36]、Popbots[77]、Atena[101]
    针对特定
    公共卫生事件
    Aroha[89]
    针对其他心理问题 1) 情感性障碍: Sofia[103]
    2) 性虐待: ReClaim[105]和SCD[106]、3) 饮食失调: KIT[107]和LLM辅助主题建模方法[108]、4) 创伤后应激障碍PTSD: PTSDialogue[109]、5) 老年人社会隔离: 虚拟同伴系统[111]
    注: “代表性模型/方法”列中, 原文献已命名的模型或框架沿用原名, 未命名则以核心方法代称.
    下载: 导出CSV

    表  3  基于生成的心理健康问答方法对比

    Table  3  Comparison of generation-based methods for mental health QA

    方法 定义 优点 缺点 细分方法及
    技术路径
    策略
    可控性
    误差传递
    风险
    数据
    依赖
    主要特点 代表性模型/
    方法
    基于
    管道
    由多个串联、功能不同的子模块构成, 各模块分阶段处理不同任务并生成回答 模块化、专业性和可控性高、数据需求灵活 误差传递、工程复杂、局部最优不等于全局最优 情绪—意图
    识别
    从输入文本中提取并识别用户的情感与意图特征 EmoDS[42]、任务自适应标记法[115]
    策略行为
    优化
    根据历史信息或状态, 优化对话策略或行为, 对生成过程加以约束 E-SCBA[116]、CDL[117]
    外部知识融合 引入外部知识, 提升模型自身常识推理能力 S2SPMN[118]、常识—知识整合框架[43]
    基于
    端到端
    采用一体化建模方式, 依托预训练架构, 根据会话历史或相关知识直接生成回答 避免误差传递、端到端可微、更灵活 数据质量要求高、存在幻觉与不准确输出、可解释性差 模型微调 使用领域数据集微调, 优化模型参数, 内化领域知识 心理治疗会话
    机器人[119]
    融合心理学
    知识
    在模型生成过程中融入专业知识, 引导模型学习人类咨询师咨询策略技巧, 优化模型输出 ESC[30]、DialogueCRN[48]
    强化学习 针对特定的生成目标, 设计多重奖励机制对模型输出进行约束与优化 PARNER[44]、VA[45]、Supporter[46]
    注: ○表示低/弱, 表示中等/一般, ●表示高/强. “代表性模型/方法”列中, 原文献已命名的模型或框架沿用原名, 未命名则以核心方法代称.
    下载: 导出CSV

    表  4  基于LLMs的心理健康问答方法对比

    Table  4  Comparison of LLM-Based Methods for Mental Health QA

    方法 定义 优点 缺点 细分方法及
    技术路径
    算力
    需求
    专业性 共情度 稳定性 可解
    释性
    幻觉
    风险
    主要特点 代表性模型/
    方法
    基于提示 无需微调, 通过提示引导LLMs调用内置知识, 适配心理健康领域任务 无需微调、效率高、领域相关性高 高度依赖提示语设计、稳定性差 zero-shot 直接通过提示指令引导模型生成, 无需示例 SuDoSys[38]、PATIENT-Ψ[58]
    few-shot 提供少量领域相关示例辅助模型理解和生成 few-shot-LLM[145]
    CoT及拓展 分步骤引导模型进行推理, 包括标准CoT、共情链、交互链等 Cue-CoT[146]、ECoT[147]、DoT[149]、CoE[150]、CoI[151]
    基于微调 微调优化模型参数, 内化领域知识, 将通用LLMs转化为领域专用LLMs 提示语依赖弱、稳定性和专业性强、领域知识内化与泛化能力强 局限于开源LLMs、硬件要求高、有限的透明度 通用目标导向微调 侧重于提升LLMs在心理支持任务上的整体指标与泛化表现 Mental-LLM[59]、Chat Counselor[36]
    特定目标导向微调 侧重于特定目标或指标的提升, 更聚焦于特定需求 CPsy Coun[52]、MentaL LaMA[57]、Psy-LLM[56]
    提示与微调融合 先用提示框架约束生成或扩展数据, 再利用微调内化LLMs专业能力 专业性强、稳定可控、领域知识内化与泛化能力强 依赖提示设计、局限于开源LLMs、硬件要求高 心理学理论驱动融合 侧重于支持过程符合心理学派的理论规范和实践范式 POST[69]、PsyAdvisor[53]、CBT-LLM[54]
    个性化目标驱动融合 侧重支持结果契合求助者个性化与心理特质 PsyDT[51]、SoulChat[55]
    长程多会话驱动融合 侧重于跨越多个会话周期的长期心理状态追踪与动态记忆管理 MusPsy[152]、Anna Agent[153]
    基于检索增强生成 引入外部专业知识库作为上下文检索信息, 增强LLMs生成内容 专业性强、稳定性强、无需微调 高度依赖检索知识库、过程复杂度提高 情感导向检索 构建共情反应数据库, 检索匹配情感状态的支持策略 Aptness[154]
    认知导向检索 整合认知常识知识, 检索与用户心理状态相关的认知关系 D2RCU[155]
    其他方法 强化学习、对比学习或多智能体协同等训练协作机制, 优化LLMs生成性能 可靠性高、多维协同与动态纠偏能力强 模型架构复杂、部署与运维复杂度高、技术门槛高 强化/对比学习的优化 引入强化学习或对比学习, 针对特定任务进行奖惩反馈 DP_2O[156]、Muffin[157]
    多智能体协作与架构创新 构建智能体协同或大小模型混合架构, 实现细化分工与协作 Safeguard GPT[158]、HEF[159]
    注: ○表示低/弱, 表示中等/一般, ●表示高/强; “代表性模型/方法”列中, 原文献已命名的模型或框架沿用原名; 未命名则以核心方法代称.
    下载: 导出CSV

    表  5  基于LLMs的多模态心理健康问答方法对比

    Table  5  Comparison of LLM-based multimodal methods for mental health QA

    方法 主要思想 优点 缺点 跨模态融合 跨模态推理 专业对齐度 实时响应性 代表性模型/
    方法
    特征级联与结构化融合 使用专门的模型提取各模态的显性特征, 然后通过结构化的管道方法将其融合 独立感知模型, 特征提取目标明确, 工程实现结构清晰 缺乏深层跨模态逻辑的动态推理能力 视觉−文本融合框
    [92]、IntelliCare[93]
    多模态认知链与动态推理 引入多模态思维链机制, 引导模型分步骤进行动态推理 避免信息过载、深层理解隐晦线索 依赖复杂多模态提示设计 Empatheia[61]
    基于强化学习与多智能体协同的方法 结合强化学习进行专业对齐, 并利用多智能体进行跨模态协同与交叉验证 可靠性强, 降低幻觉与策略偏移 系统架构复杂, 多智能体交互存在延迟 FIRES[60]、MultiMood[62]、DELTA[63]
    注: ○ 表示低/弱, 表示中等/一般, ● 表示高/强; “代表性模型/方法”列中, 原文献已命名的模型或框架沿用原名; 未命名则以核心方法代称.
    下载: 导出CSV

    表  6  中文心理健康问答数据集

    Table  6  Chinese mental health QA datasets

    类型 名称(文献) 来源 形式 数量 优缺点 可用性与伦理
    面向广泛类别心理问题 Efaqa (斯坦福/加州/辅仁大学) 真实咨询对话 多轮对话 20 K样本 优点: 带有烦恼类型/心理疾病类型/SOS紧急情况标签
    缺点: 回复较为简短笼统, 缺乏具体建议
    获取: 需购买证书下载(https://github.com/chatopera/efaqa-corpus-zh)
    协议: 春松许可证v1.0
    伦理: 人工主观标注, 官方明确声明免责
    PsyQA[21] “壹心理”平台 单轮对话 22 K问题和55 K条回答 优点: 带有支持策略标签
    缺点: 异步非实时问答数据, 难以模拟真实咨询会话场景
    获取: 需签署协议并通过邮件申请(https://github.com/thu-coai/PsyQA)
    协议: 仅限学术及非商业用途
    伦理: 已进行数据脱敏与匿名化处理
    SmileChat[166] PsyQA 多轮对话 55 K样本, 平均对话轮次24 优点: 弥补PsyQA单轮问答不足
    缺点: 仅基于PsyQA数据集改写生成, 缺乏原生语料的引入; 支持策略标签缺失
    获取: https://github.com/qiuhuachuan/smile(可公开获取)
    协议: CC0 1.0
    伦理: LLMs合成数据, 明确不可替代专业干预
    Xinling[167] 自建平台 多轮对话 2.4 K样本, 平均对话轮次78.49 优点: 标注了咨询师意图与策略、来访者反应与行为模式
    缺点: 数据规模较小
    获取: 原始数据未公开, 可邮件申请(https://github.com/dll-wu/Client-Reactions)
    协议: 数据未开源, 无相关数据协议
    伦理: 基于严格的伦理和隐私保护原则, 未公开原始数据
    Psy-Insight[91] 博客和书籍的真实咨询对话 中/英文431/520样本, 平均对话轮次77/55 优点: 标注了心理治疗方法、情绪状态、支持策略与咨询主题标签, 并包含轮次推理与会话指导信息; 平均轮次高
    缺点: 数据规模较小
    获取: https://github.com/ckqqqq/Psy-Insight(可公开获取)
    协议: MIT许可证
    伦理: 已严格匿名化, 无患者隐私泄露风险
    CPsyCounD[52] “壹心理”平台 多轮对话 3.13 K样本, 平均对话轮次8.7 优点: 隐私保护、专业性、真实性和安全性强
    缺点: 数据规模较小
    获取: https://github.com/CAS-SIAT-XinHai/CPsyCoun(可公开获取)
    协议: CC BY 4.0协议
    伦理: LLMs重构生成, 规避了隐私泄露风险
    ProPsyC[53] Xinling/CPsyCounD/Psy-Insight 多轮对话 2 K样本, 平均对话轮次16.12 优点: 标注了咨询师主动和被动策略、来访者消极和积极反应
    缺点: 数据规模较小
    获取: https://github.com/EthanHu777/PsyAdvisor(可公开获取)
    协议: 仅限学术及非商业用途
    伦理: 已进行数据脱敏与匿名化处理
    SoulChatCorpus[55] 众包收集与LLMs提示 多轮对话 1 200 K样本 优点: 语料规模百万级别、共情度高
    缺点: 未包含任何标签
    获取: https://github.com/scutcyr/SoulChat?tab=readme-ov-file(可公开获取)
    协议: 仅限学术与非商业研究用途
    伦理: LLMs合成, 已剔除敏感与隐私内容
    CBTQA[54] PsyQA 单轮问答 22 K样本 优点: 基于CBT理论构建, 具备较强的干预指导性
    缺点: 仅基于PsyQA数据集进行拓展, 缺乏原生语料的引入; 策略标签缺失
    获取: 需联系作者申请(https://huggingface.co/Hongbin37/CBT-LLM)
    协议: Apache-2.0协议
    伦理: LLMs合成, 数据已脱敏与匿名化处理
    PsyDTCorpus[51] 真实咨询案例 多轮对话 5 K样本, 平均对话轮次18 优点: 融合咨询师风格与来访者人格特质
    缺点: 数据规模较小
    获取: https://github.com/scutcyr/SoulChat2.0(可公开获取)
    协议: 仅限学术与非商业研究用途
    伦理: 已剔除敏感与隐私内容
    面向特定类别心理问题 Chinese-Psychological-QA[168] 丁香园、寻医问药平台 单轮问答 44 K条问题、45.8 K条回答 优点: 涵盖11个二级科室、12类心理疾病, 标注了心理问题类别
    缺点: 单轮问答, 难以模拟真实咨询会话场景
    获取: https://github.com/flyrae/Chinese-Psychological-QA-DataSet(可公开获取)
    协议: 未明确声明开源协议
    伦理: 未明确, 存在伦理风险
    D4[169] 真人模拟 多轮对话 1.34 K样本, 平均对话轮次21.6 优点: 标注了临床医生提供的抑郁风险和自杀风险评分
    缺点: 数据规模较小
    获取: 需签署使用协议申请(https://x-lance.github.io/D4/)
    协议: 仅限学术研究用途
    伦理: 由专业医师监督构建, 要求实名签署协议
    未命名[69] D4 多轮对话 1.34 K样本, 平均对话轮次21.6 优点: 基于ABC理论补充标注了心理状态跟踪
    缺点: 仅基于D4数据集进行二次标注, 缺乏原生语料的引入; 数据规模较小
    获取: 未公开, 可邮件申请
    协议: 仅限学术研究用途
    伦理: 数据来源于D4, 遵循D4伦理准则
    多模态情感支持 M3ED[170] 电视剧 多轮对话 9.08 K样本,
    24.4 K条话语
    优点: 标注了细粒度混合情感标签
    缺点: 基于电视剧片段, 非真实会话场景
    获取: https://github.com/AIM3-UC/RUCM3ED(可公开获取)
    协议: 仅限学术研究用途
    伦理: 数据来源公开, 无伦理隐私风险
    EmotionTalk[171] 专业演员录制 多轮对话 744段对话,
    19.6 K条话语
    优点: 标注了7类离散情感、5维情感极性及4类说话风格描述
    缺点: 人工录制的模拟数据, 非真实会话场景
    获取: https://github.com/NKU-LT/EmotionTalk(可公开获取)
    协议: 仅限学术研究用途
    伦理: 官方要求需严格遵守伦理与隐私保护原则
    CPED[172] 电视剧 多轮对话 12 K样本,
    133 K条话语
    优点: 标注了10种对话场景、19种对话行为、13类情绪和人格特质
    缺点: 基于电视剧片段, 非真实会话场景
    获取: https://github.com/scutcyr/CPED(可公开获取)
    协议: 仅限学术研究用途
    伦理: 已剔除敏感与隐私内容
    下载: 导出CSV

    表  7  英文心理健康问答数据集

    Table  7  English mental health QA datasets

    类型 名称(文献) 来源 形式 数量 优缺点 可用性与伦理
    面向广泛类别
    心理问题
    未命名[173] 真实短信危机干预咨询对话 多轮对话 80 K样本 优点: 标注了三类干预结果和主要问题类型
    缺点: 仅来源短信咨询的用户群体
    获取: 需提交申请并审核后才可获取(http://snap.stanford.edu/counseling)
    协议: 严格的数据使用协议, 禁止商业用途, 仅限学术研究
    伦理: 包含敏感的真实危机干预短信记录, 官方对数据访问隐私保护和伦理审查极其严格
    EPITOME[174] TalkLife和Reddit 多轮问答 8 000 K样本 优点: 多维度共情标注、跨平台
    缺点: 异步非实时问答数据, 难以模拟真实咨询会话场景
    获取: https://github.com/behavioral-data/Empathy-Mental-Health (可公开获取)
    协议: 官方自定义的学术与署名许可
    伦理: 已剔除敏感与隐私内容
    Psych8k[36] 真实咨询
    对话录音
    单轮问答 8.2 K样本 优点: 专业、强调互动性、蕴含反思、积极倾听等核心咨询技巧
    缺点: 数据规模较小
    获取: 需提交申请才可下载(https://huggingface.co/datasets/EmoCareAI/Psych8k)
    协议: 要求署名, 禁止商业用途, 仅限学术研究
    伦理: 真实数据与LLMs合成, 已剔除敏感与隐私内容
    ConvCounsel[175] YouTube心理频道和模拟咨询对话 多轮对话 5 K样本 优点: 强调积极倾听、标注学生情绪、咨询师策略和角色等多维度标签
    缺点: 数据规模较小
    获取: 需联系作者申请
    协议: CC BY 4.0
    伦理: 公开平台数据, 已剔除敏感与隐私内容
    IEMPATHIZE[176] 在线癌症幸存者平台(CSN) 多轮问答 5 K样本 优点: 标注细粒度共情方向
    缺点: 数据规模较小, 异步非实时问答数据
    获取: https://github.com/Mahhos/Empathy (可公开获取)
    协议: 未明确声明开源协议
    伦理: 来源真实数据, 要求严格遵守隐私保护
    CACTUS[90] AI模拟
    咨询对话
    多轮对话 1 000 K样本, 平均对话轮次16.6 优点: 大规模、遵循CBT理论、结构化
    缺点: 来源于AI模拟的咨询对话, 非真实咨询数据
    获取: https://github.com/coding-groot/cactus (可公开获取)
    协议: GPL-2.0
    伦理: LLMs合成, 已剔除敏感与隐私内容
    面向特定类别
    心理问题
    MotiVAte[177] 在线抑郁症论坛 多轮对话 4 K样本, 平均对话轮次4 优点: 聚焦抑郁症的专业对话
    缺点: 数据规模较小
    获取: 未明确获取途径
    协议: 未明确声明开源协议
    伦理: 已剔除敏感与隐私内容
    MHQA[178] 心理健康的
    文章摘要
    单轮问答 580 K样本 优点: 聚焦焦虑、抑郁、创伤和强迫的专业数据
    缺点: 基于书面摘要构建, 非真实咨询数据
    获取: https://github.com/joshiprashanthd/mhqa (可公开获取)
    协议: 仅限非商业性学术研究使用
    伦理: LLMs合成与人工校对, 剔除敏感与隐私内容
    面向情感
    支持任务
    ED[179] 众包方式 多轮对话 25 K样本, 平均对话轮次6 优点: 标注32种情绪类别、话题围绕特定情境开展
    缺点: 只标注情绪, 标注类别较少
    获取: https://github.com/facebookresearch/EmpatheticDialogues (可公开获取)
    协议: 仅限非商业性学术研究使用
    伦理: 公开平台数据, 已剔除敏感与隐私内容
    ESConv[30] 真人模拟 多轮对话 1 K样本, 平均对话轮次29.8 优点: 标注求助者问题类型、情绪类型和强度、支持者的支持策略, 以及支持质量反馈
    缺点: 数据规模较小
    获取: https://github.com/thu-coai/Emotional-Support-Conversation (可公开获取)
    协议: 仅限非商业性学术研究使用
    伦理: 已规避敏感与隐私内容
    Augesc[180] ESConv和ED 多轮对话 65 K个样本 优点: 规模大、主题广
    缺点: 基于公开数据构建, 缺乏原生语料的引入
    获取: https://github.com/thu-coai/AugESC (可公开获取)
    协议: 仅限非商业性学术研究使用
    伦理: 已规避敏感与隐私内容
    ExTES[181] ESConv/
    ETMHS/
    Reddit和网络
    多轮对话 11.2 K样本, 平均对话轮次18.2 优点: 标注16种支持策略、32种情感场景
    缺点: 基于公开数据构建, 缺乏原生语料的引入
    获取: https://github.com/pandazzh2020/ExTES (可公开获取)
    协议: 未明确声明开源协议
    伦理: 来源公开数据, 已剔除敏感与隐私内容
    其他 TREC QA[182]、BioASQ 2019[183]和PubMedQA[184]等生物医学问答数据集也包含心理健康相关数据,
    MedMCQA[185]设有专门的精神病学类别
    MESC[41] 电视剧 多轮对话 1 K样本, 28 762条话语 优点: 标注了咨询场景、支持策略和情绪状态
    缺点: 基于电视剧片段, 非真实咨询场景
    获取: https://github.com/chuyq/MESC (可公开获取)
    协议: 仅限学术研究用途
    伦理: 数据来源公开, 无伦理隐私风险
    多模态心理健康/情感支持(未包含生理信号) IEMOCAP[186] 专业演员
    即兴对话
    多轮对话 5个长会话片段, 约10 039条话语 优点: 标注9类离散情绪和3维连续情感属性
    缺点: 数据规模较小、基于专业演员的既定场景对话, 非真实咨询场景
    获取: 提交申请并签署用户许可协议(https://sail.usc.edu/iemocap/)
    协议: 仅限学术研究用途
    伦理: 已规避敏感与隐私内容
    MELD[187] 电视剧 多轮对话 1.4 K段对话、13 000条话语 优点: 标注7类离散情绪和3类情感倾向标签
    缺点: 基于喜剧片段, 非真实咨询场景
    获取: https://github.com/declare-lab/MELD (可公开获取)
    协议: 仅限学术研究用途
    伦理: 数据来源公开, 已规避敏感与隐私内容
    AvaMERG[61] ED为基础, 真实志愿者录制 多轮对话 33 K样本, 152 K条话语, 平均对话轮次4.6 优点: 继承了ED标签, 并包含多维度用户画像、数据分布平衡
    缺点: 基于公开数据人工模拟, 非真实咨询场景
    获取: https://avamerg.github.io/ (可公开获取)
    协议: 仅限学术研究用途
    伦理: 基于已有数据, LLMs+人工录制, 已规避敏感与隐私内容
    EMOTyDA[188] MELD/
    IEMOCAP
    多轮对话 1.3 K段对话, 19 K条话语 优点: 继承了MELD和IEMOCAP标签, 并标注了12类对话行为
    缺点: 基于公开数据, 非真实咨询场景
    获取: https://github.com/sahatulika15/EMOTyDA (可公开获取)
    协议: 仅限学术研究用途
    伦理: 已规避敏感与隐私内容
    多模态心理健康/情感支持(包含生理信号) Badalona Corpus[189] 5组双人真实环境下的对话 对话+生理信号 单次会话0.5 h
    (总计约7.5 h)
    优点: 脑电+生理+音视频数据, 长轮次
    缺点: 数据规模较小
    获取: 需向研究机构提交正式申请后获取
    协议: 仅限学术研究用途
    伦理: 经过了严格的伦理委员会审批与受试者的同意, 官方要求使用者必须遵守伦理与隐私安全
    K-Emo Con[190] 真实对话 对话+生理信号 单次会话10 min
    (总计约160 min)
    优点: 脑电+生理+音视频数据, 标注多维度情感评分
    缺点: 数据规模较小
    获取: 邮件提交申请且需审核
    协议: 仅限学术研究用途
    伦理: 经过了严格的伦理委员会审批与受试者的同意, 官方要求使用者必须遵守伦理与隐私安全
    下载: 导出CSV

    表  8  自动评估指标对比

    Table  8  Comparison of automatic evaluation metrics

    方法 原理 优势 局限 与人工评估相关性
    PPL 预测下一个词的概率 简单直观 无法衡量语义连贯性 中等偏高
    BLEU n-gram精确度 简单快速、可解释 表面形式匹配, 忽略语义、位置不敏感 中等
    ROUGE n-gram召回率 强调召回率、解释性强 忽略语义等同性 中等
    Distinct-n n-gram不重复率 简单快速、可解释 表面形式匹配, 忽略语义、位置不敏感 中等
    BERT Score 上下文嵌入 语义级匹配、上下文感知、跨语言支持 计算密集, 复杂度高
    下载: 导出CSV

    A1  检索式、生成式及LLMs相关文献纳入清单

    A1  List of included literature on retrieval-based, generative, and LLMs

    作者 年份 标题 类型 方法类别
    Liu等[27] 2013 PAL: A chatterbot system for answering domain-specific questions 会议 检索式
    Demasi等[28] 2019 Towards augmenting crisis counselor training by improving message retrieval 会议 检索式
    Collins等[24] 2022 Covid connect: Chat-driven anonymous story-sharing for peer support 会议 检索式
    Sun等[197] 2022 Comparing experts and novices for ai data work: Insights on allocating human intelligence to design a conversational agent 会议 检索式
    Maharjan等[112] 2021 Can we talk? Design implications for the questionnaire-driven self-report of health and wellbeing via conversational agent 会议 检索式
    Boyd等[113] 2022 Usability testing and trust analysis of a mental health and wellbeing chatbot 会议 检索式
    Zhang等[12] 2020 Recent advances and challenges in task-oriented dialog systems 期刊 生成式
    (管道)
    Zhou等[121] 2018 Emotional chatting machine: Emotional conversation generation with internal and external memory 会议 生成式
    (管道)
    Liu等[115] 2023 Task-adaptive tokenization: Enhancing long-form text generation efficacy in mental health and beyond 会议 生成式
    (管道)
    Song等[42] 2019 Generating responses with a specific emotion in dialog 会议 生成式
    (管道)
    Li等[116] 2018 A syntactically constrained bidirectional-asynchronous approach for emotional conversation generation 会议 生成式
    (管道)
    Shen等[117] 2020 CDL: Curriculum dual learning for emotion-controllable response generation 会议 生成式
    (管道)
    Pei等[118] 2018 S2SPMN: A simple and effective framework for response generation with relevant information 会议 生成式
    (管道)
    Shen等[43] 2022 Knowledge enhanced reflection generation for counseling dialogues 会议 生成式
    (管道)
    Ji等[129] 2022 Mentalbert: Publicly available pretrained language models for mental healthcare 会议 生成式
    (端到端)
    Ji等[130] 2023 Domain-specific continued pretraining of language models for capturing long context in mental health 预印本 生成式
    (端到端)
    Das等[119] 2022 Conversational bots for psychotherapy: A study of generative transformer models using domain-specific dialogues 会议 生成式
    (端到端)
    Zhai等[133] 2024 Chinese MentalBERT: Domain-adaptive pre-training on social media for Chinese mental health text analysis 会议 生成式
    (端到端)
    Liu等[30] 2021 Towards emotional support dialog systems 会议 生成式
    (端到端)
    Sharma等[44] 2021 Towards facilitating empathic conversations in online mental health support: A reinforcement learning approach 会议 生成式
    (端到端)
    Hu等[48] 2021 Dialoguecrn: Contextual reasoning networks for emotion recognition in conversations 会议 生成式
    (端到端)
    Saha等[45] 2022 Towards motivational and empathetic response generation in online mental health support 会议 生成式
    (端到端)
    Zhou等[46] 2023 Facilitating multi-turn emotional support conversation with positive emotion elicitation: A reinforcement learning approach 会议 生成式
    (端到端)
    Sun等[21] 2021 Psyqa: A Chinese dataset for generating long counseling text for mental health support 会议 生成式
    (端到端)
    Haider等[47] 2025 AI-Driven mental health chatbot: Empowering Well-Being with conversational AI and retrieval-augmented generation 会议 生成式
    (混合方法)
    Raza等[50] 2022 CoQUAD: A COVID-19 question answering dataset system, facilitating research, benchmarking, and practice 期刊 生成式
    (混合方法)
    Deng等[35] 2023 Knowledge-enhanced mixed-initiative dialogue system for emotional support conversations 会议 生成式
    (混合方法)
    Chen等[38] 2024 Structured dialogue system for mental health: An LLM chatbot leveraging the pm+ guidelines 会议 LLMs
    (提示方法)
    Wang等[58] 2024 Patient-Ψ: Using large language models to simulate patients for training mental health professionals 会议 LLMs
    (提示方法)
    Liu等[145] 2024 Large language models are few-shot health learners 预印本 LLMs
    (提示方法)
    Wang等[146] 2023 Cue-CoT: Chain-of-thought prompting for responding to in-depth dialogue questions with LLMs 预印本 LLMs
    (提示方法)
    Li等[147] 2024 Enhancing emotional generation capability of large language models via emotional chain-of-thought 预印本 LLMs
    (提示方法)
    Chen等[149] 2023 Empowering psychotherapy with large language models: Cognitive distortion detection through diagnosis of thought prompting 预印本 LLMs
    (提示方法)
    Han等[151] 2024 Chain-of-interaction: Enhancing large language models for psychiatric behavior understanding by dyadic contexts 会议 LLMs
    (提示方法)
    Xu等[59] 2024 Mental-LLM: Leveraging large language models for mental health prediction via online text data 会议 LLMs
    (微调)
    Liu等[36] 2024 ChatCounselor: A large language models for mental health support 预印本 LLMs
    (微调)
    Hu等[160] 2023 LLM-adapters: An adapter family for parameter-efficient fine-tuning of large language models 会议 LLMs
    (微调)
    Qiu等[161] 2025 PsyDial: A large-scale long-term conversational dataset for mental health support 会议 LLMs
    (微调)
    Zhang等[52] 2024 CPsyCoun: A report-based multi-turn dialogue reconstruction and evaluation framework for chinese psychological counseling 会议 LLMs
    (微调)
    Chen等[91] 2025 Psy-Insight: Explainable multi-turn bilingual dataset for mental health counseling 预印本 LLMs
    (微调)
    Yang等[57] 2024 MentaLLaMA: Interpretable mental health analysis on social media with large language models 会议 LLMs
    (微调)
    Lai等[56] 2023 Psy-LLM: Scaling up global mental health psychological services with ai-based large language models 预印本 LLMs
    (微调)
    Gu等[69] 2025 Enhancing depression-diagnosis-oriented chat with psychological state tracking 会议 LLMs
    (提示+微调)
    Hu等[53] 2025 PsyAdvisor: A plug-and-play strategy advice planner with proactive questioning in psychological conversations 会议 LLMs
    (提示+微调)
    Na等[54] 2024 CBT-LLM: A Chinese large language model for cognitive behavioral therapy-based mental health question answering 会议 LLMs
    (提示+微调)
    Xie等[51] 2025 PsyDT: Using LLMs to construct the digital twin of psychological counselor with personalized counseling style for psychological counseling 会议 LLMs
    (提示+微调)
    Chen等[55] 2023 SoulChat: Improving LLMs' empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations 会议 LLMs
    (提示+微调)
    Hu等[154] 2024 Aptness: Incorporating appraisal theory and emotion support strategies for empathetic response generation 会议 LLMs
    (检索增强)
    Xu等[155] 2024 Dynamic demonstration retrieval and cognitive understanding for emotional support conversation 会议 LLMs
    (检索增强)
    Li等[156] 2024 Dialogue for prompting: A policy-gradient-based discrete prompt generation for few-shot learning 会议 LLMs
    (强化学习)
    Wang等[157] 2024 Muffin: Mitigating unhelpfulness in emotional support conversations with multifaceted AI feedback 会议 LLMs
    (对比学习)
    Lin等[158] 2024 Towards healthy AI: Large language models need therapists too 会议 LLMs
    (多智能体)
    Yang等[159] 2024 Enhancing empathetic response generation by augmenting LLMs with small-scale empathetic models 预印本 LLMs
    (大小模型)
    Zhu等[92] 2025 Integrating visual modalities with large language models for mental health support 会议 LLMs
    (多模态)
    Gounder等[93] 2024 A hybrid-multimodal mental health chatbot for psychological counselling 会议 LLMs
    (多模态)
    Wang等[60] 2025 Flexible thinking for multimodal emotional support conversation via reinforcement learning 会议 LLMs
    (多模态)
    Zhang等[61] 2025 Towards multimodal empathetic response generation: A rich text-speech-vision avatar-based benchmark 会议 LLMs
    (多模态)
    Le等[62] 2025 Reinforcing trustworthiness in multimodal emotional support systems 预印本 LLMs
    (多模态)
    Yang等[63] 2026 DELTA: Deliberative multi-agent reasoning with reinforcement learning for multimodal psychological counseling 预印本 LLMs
    (多模态)
    下载: 导出CSV

    B1  排除的主要文献及原因

    B1  List of excluded references with reasons

    作者 年份 标题 类型 排除原因
    Thekkekara等 2024 An attention-based CNN-BiLSTM model for depression detection on social media text 期刊 心理健康分类任务,
    非问答任务
    Dong等 2024 How abilities in large language models are affected by supervised fine-tuning data composition 会议 研究主题不相关
    Ehsan等 2023 Evaluating open-domain question answering in the era of large language models 会议 开放域问答技术方法,
    非心理健康
    Feng等 2023 Joint constrained learning with boundary-adjusting for emotion-cause pair extraction 会议 仅涉及情感原因分析,
    未涉及问答生成
    Julio等 2024 ChatGPT and social psychiatry: A commentary on the article “Old dog, new tricks? Exploring the potential functionalities of ChatGPT in supporting educational methods in social psychiatry” 期刊 评论述评文章,
    无模型/数据/实验
    Singhal等 2023 Large language models encode clinical knowledge 期刊 生物医学领域问答
    下载: 导出CSV
  • [1] World Health Organization. World Mental Health Day 2022: Make Mental Health and Well-Being for All a Global Priority[Online], available: https://www.who.int/campaigns/world-mental-health-day/2022, October 13, 2022.
    [2] 傅小兰, 张侃, 陈雪峰, 陈祉妍. 中国国民心理健康发展报告(2021—2022). 北京: 社会科学文献出版社, 2023.

    Fu Xiao-Lan, Zhang Kan, Chen Xue-Feng, Chen Zhi-Yan. Report on National Mental Health Development in China (2021-2022). Beijing: Social Sciences Academic Press (China), 2023.
    [3] 姚元杰, 龚毅光, 刘佳, 徐闯, 朱栋梁. 基于深度学习的智能问答系统综述. 计算机系统应用, 2023, 32(4): 1−15 doi: 10.15888/j.cnki.csa.009038

    Yao Yuan-Jie, Gong Yi-Guang, Liu Jia, Xu Chuang, Zhu Dong-Liang. Survey on intelligent question answering system based on deep learning. Computer Systems & Applications, 2023, 32(4): 1−15 doi: 10.15888/j.cnki.csa.009038
    [4] 车万翔, 张伟男. 人机对话系统综述. 人工智能, 2018(1): 76−82 doi: 10.16453/j.cnki.issn2096-5036.2018.01.007

    Che Wan-Xiang, Zhang Wei-Nan. Overview of human computer dialogue system. Artificial Intelligence, 2018(1): 76−82 doi: 10.16453/j.cnki.issn2096-5036.2018.01.007
    [5] Diefenbach D, Lopez V, Singh K D, Maret P. Core techniques of question answering systems over knowledge bases: A survey. Knowledge and Information Systems, 2018, 55(3): 529−569 doi: 10.1007/s10115-017-1100-y
    [6] Casu M, Triscari S, Battiato S, Guarnera L, Caponnetto P. AI chatbots for mental health: A scoping review of effectiveness, feasibility, and applications. Applied Sciences, 2024, 14(13): 5889 doi: 10.3390/app14135889
    [7] 杨州, 陈志豪, 蔡铁城, 王宇峰, 廖湘文. 基于深度学习的情感对话响应综述. 计算机学报, 2023, 46(12): 2489−2519 doi: 10.11897/SP.J.1016.2023.02489

    Yang Zhou, Chen Zhi-Hao, Cai Tie-Cheng, Wang Yu-Feng, Liao Xiang-Wen. A survey of deep learning based emotional dialogue response. Chinese Journal of Computers, 2023, 46(12): 2489−2519 doi: 10.11897/SP.J.1016.2023.02489
    [8] Zhao W X, Zhou K, Li J, Tang T, Wang X, Hou Y, et al. A survey of large language models. arXiv preprint arXiv: 2303.18223, 2023.
    [9] Guo Z, Lai A, Thygesen J H, Farrington J, Keen T, Li K. Large language models for mental health applications: Systematic review. JMIR Mental Health, 2024, 11(1): e57400 doi: 10.2196/preprints.57400
    [10] Ma Y, Nguyen K L, Xing F Z, Cambria E. A survey on empathetic dialogue systems. Information Fusion, 2020, 64: 50−70 doi: 10.1016/j.inffus.2020.06.011
    [11] Sorin V, Brin D, Barash Y, Konen E, Charney A, Nadkarni G, et al. Large language models and empathy: Systematic review. Journal of Medical Internet Research, 2024, 26: e52597 doi: 10.2196/52597
    [12] Zhang Z, Takanobu R, Zhu Q, Huang M, Zhu X. Recent advances and challenges in task-oriented dialog systems. Science China Technological Sciences, 2020, 63(10): 2011−2027 doi: 10.1007/s11431-020-1692-3
    [13] 赵芸, 刘德喜, 万常选, 刘喜平, 廖国琼. 检索式自动问答研究综述. 计算机学报, 2021, 44(6): 1214−1232 doi: 10.11897/SP.J.1016.2021.01214

    Zhao Yun, Liu De-Xi, Wan Chang-Xuan, Liu Xi-Ping, Liao Guo-Qiong. Retrieval-based automatic question answer: A literature survey. Chinese Journal of Computers, 2021, 44(6): 1214−1232 doi: 10.11897/SP.J.1016.2021.01214
    [14] Shi X, Liu Z, Du L, Wang Y, Wang H, Guo Y, et al. Medical dialogue: A survey of categories, methods, evaluation and challenges. In: Proceedings of Findings of the Association for Computational Linguistics. Bangkok, Thailand: Association for Computational Linguistics, 2024. 2840−2861.
    [15] Valizadeh M, Parde N. The AI doctor is in: A survey of task-oriented dialogue systems for healthcare applications. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin, Ireland: Association for Computational Linguistics, 2022. 6638−6660.
    [16] Xu B, Zhuang Z. Survey on psychotherapy chatbots. Concurrency and Computation: Practice and Experience, 2022, 34(7): e6170 doi: 10.1002/cpe.6170
    [17] Abd-Alrazaq A A, Alajlani M, Alalwan A A, Bewick B M, Gardner P, Househ M. An overview of the features of chatbots in mental health: A scoping review. International Journal of Medical Informatics, 2019, 132: 103978 doi: 10.1016/j.ijmedinf.2019.103978
    [18] Abd-Alrazaq A A, Rababeh A, Alajlani M, Bewick B M, Househ M. Effectiveness and safety of using chatbots to improve mental health: Systematic review and meta-analysis. Journal of Medical Internet Research, 2020, 22(7): e16021 doi: 10.2196/16021
    [19] Li J, Li Y, Hu Y, Ma D C F, Mei X, Chan E A, et al. Chatbot-delivered interventions for improving mental health among young people: A systematic review and meta-analysis. Worldviews On Evidence-Based Nursing, 2025, 22(4): e70059 doi: 10.1111/wvn.70059
    [20] Bucher A, Egger S, Vashkite I, Wu W, Schwabe G. “it's not only attention we need”: Systematic review of large language models in mental health care. JMIR Mental Health, 2025, 12: e78410 doi: 10.2196/78410
    [21] Sun H, Lin Z, Zheng C, Liu S, Huang M. PsyQA: A Chinese dataset for generating long counseling text for mental health support. In: Proceedings of Findings of the 59th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2021. 1489−1503.
    [22] Weizenbaum J. ELIZA—a computer program for the study of natural language communication between man and machine. Communications of the ACM, 1966, 9(1): 36−45 doi: 10.1145/365153.365168
    [23] Colby K M. PARRYing. Behavioral and Brain Sciences, 1981, 4(4): 550−560 doi: 10.1017/S0140525X00000224
    [24] Collins C, Arbour S, Beals N, Yama S, Laffier J, Zhao Z. COVID connect: Chat-driven anonymous story-sharing for peer support. In: Proceedings of the 2022 ACM Designing Interactive Systems Conference. New York, USA: ACM, 2022. 301−318.
    [25] Morris R R, Kouddous K, Kshirsagar R, Schueller S M. Towards an artificially empathic conversational agent for mental health applications: System design and user perceptions. Journal of Medical Internet Research, 2018, 20(6): e10148 doi: 10.2196/10148
    [26] Yu L, Wu C, Jang F. Psychiatric document retrieval using a discourse-aware model. Artificial Intelligence, 2009, 173(7-8): 817−829 doi: 10.1016/j.artint.2008.12.004
    [27] Liu Y, Liu M, Wang X, Wang L, Li J. PAL: A chatterbot system for answering domain-specific questions. In: Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics: System Demonstrations. Sofia, Bulgaria: Association for Computational Linguistics, 2013. 67−72.
    [28] Demasi O, Hearst M A, Recht B. Towards augmenting crisis counselor training by improving message retrieval. In: Proceedings of the Sixth Workshop on Computational Linguistics and Clinical Psychology. Minneapolis, Minnesota: Association for Computational Linguistics, 2019. 1−11.
    [29] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, et al. Attention is all you need. Advances in Neural Information Processing Systems, 2017, 30: 5998−6008
    [30] Liu S, Zheng C, Demasi O, Sabour S, Li Y, Yu Z, et al. Towards emotional support dialog systems. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2021. 3469−3483.
    [31] Zhang Y, Sun S, Galley M, Chen Y C, Brockett C, Gao X, et al. DialoGPT: Large-scale generative pre-training for conversational response generation. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations. Online: Association for Computational Linguistics, 2020. 270−278.
    [32] Roller S, Dinan E, Goyal N, Ju D, Williamson M, Liu Y, et al. Recipes for building an open-domain chatbot. In: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2021. 300−325.
    [33] Tu Q, Li Y, Cui J, Wang B, Wen J R, Yan R. MISC: A mixed strategy-aware model integrating COMET for emotional support conversation. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. Dublin, Ireland: Association for Computational Linguistics, 2022. 308−319.
    [34] Lu X, Tian Y, Zhao Y, Qin B. Retrieve, discriminate and rewrite: A simple and effective framework for obtaining affective response in retrieval-based chatbots. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Online: Association for Computational Linguistics, 2021. 1956−1969.
    [35] Deng Y, Zhang W, Yuan Y, Lam W. Knowledge-enhanced mixed-initiative dialogue system for emotional support conversations. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. Toronto, Canada: Association for Computational Linguistics, 2023. 4079−4095.
    [36] Liu J M, Li D, Cao H, Ren T, Liao Z, Wu J. ChatCounselor: A large language models for mental health support. arXiv preprint arXiv: 2309.15461, 2023.
    [37] Touvron H, Lavril T, Izacard G, Martinet X, Lachaux M A, Lacroix T, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv: 2302.13971, 2023.
    [38] Chen Y, Zhang X, Wang J, Xie X, Yan N, Chen H, et al. Structured dialogue system for mental health: An LLM chatbot leveraging the PM+ guidelines. In: Proceedings of the 16th International Conference on Social Robotics. Singapore: Springer Nature Singapore, 2024. 262−271.
    [39] World Health Organization. Problem Management Plus (PM+): Individual psychological help for adults impaired by distress in communities exposed to adversity (generic field-trial version 1.0)[R]. Geneva: World Health Organization, 2016.
    [40] Hassaan M, Shah M S A, Uzaif M, Afridi Y S, Khattak R U. PsyRA–a Retrieval-Augmented dialogue system for mental health support. International Journal of Innovations in Science & Technology, 2025, 7(7): 279−288
    [41] Chu Y, Liao L, Zhou Z, Ngo C W, Hong R. Towards multimodal emotional support conversation systems. IEEE Transactions on Multimedia, 2025, 27: 8276−8287 doi: 10.1109/TMM.2025.3604951
    [42] Song Z, Zheng X, Liu L, Xu M, Huang X. Generating responses with a specific emotion in dialog. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics, 2019. 3685−3695.
    [43] Shen S, Perez-Rosas V, Welch C, Poria S, Mihalcea R. Knowledge enhanced reflection generation for counseling dialogues. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin, Ireland: Association for Computational Linguistics, 2022. 3096−3107.
    [44] Sharma A, Lin I W, Miner A S, Atkins D C, Althoff T. Towards facilitating empathic conversations in online mental health support: A reinforcement learning approach. In: Proceedings of the Web Conference 2021. Ljubljana, Slovenia: ACM, 2021. 194−205.
    [45] Saha T, Gakhreja V, Das A S, Chakraborty S, Saha S. Towards motivational and empathetic response generation in online mental health support. In: Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. Madrid, Spain: ACM, 2022. 2650−2656.
    [46] Zhou J, Chen Z, Wang B, Huang M. Facilitating multi-turn emotional support conversation with positive emotion elicitation: A reinforcement learning approach. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Toronto, Canada: Association for Computational Linguistics, 2023. 1714−1729.
    [47] Haider T, Shafi N, Haq N U. AI-driven mental health chatbot: Empowering well-being with conversational AI and retrieval-augmented generation. In: Proceedings of the 2025 IEEE Conference on Artificial Intelligence. Santa Clara, USA: IEEE, 2025. 1596−1602.
    [48] Hu D, Wei L, Huai X. DialogueCRN: Contextual reasoning networks for emotion recognition in conversations. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Online: Association for Computational Linguistics, 2021. 7042−7052.
    [49] Gaur M, Gunaratna K, Srinivasan V, Jin H. ISeeQ: Information seeking question generation using dynamic meta-information retrieval and knowledge graphs. In: Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto, USA: AAAI Press, 2022. 10672−10680.
    [50] Raza S, Schwartz B, Rosella L C. CoQUAD: A COVID-19 question answering dataset system, facilitating research, benchmarking, and practice. BMC Bioinformatics, 2022, 23(1): 210 doi: 10.1186/s12859-022-04751-6
    [51] Xie H, Chen Y, Xing X, Lin J, Xu X. PsyDT: Using LLMs to construct the digital twin of psychological counselor with personalized counseling style for psychological counseling. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Vienna, Austria: Association for Computational Linguistics, 2025. 1081−1115.
    [52] Zhang C, Li R, Tan M, Yang M, Zhu J, Yang D, et al. CPsyCoun: A report-based multi-turn dialogue reconstruction and evaluation framework for Chinese psychological counseling. In: Proceedings of Findings of the Association for Computational Linguistics. Bangkok, Thailand: Association for Computational Linguistics, 2024. 13947−13966.
    [53] Hu Y, Liu D, Liu B, Chen Y, Cao J, Liu Y. PsyAdvisor: A Plug-and-Play strategy advice planner with proactive questioning in psychological conversations. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Vienna, Austria: Association for Computational Linguistics, 2025. 12205−12229.
    [54] Na H. CBT-LLM: A Chinese large language model for cognitive behavioral therapy-based mental health question answering. In: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation. Torino, Italia: European Language Resources Association and International Committee on Computational Linguistics, 2024. 2930−2940.
    [55] Chen Y, Xing X, Lin J, Zheng H, Wang Z, Liu Q, et al. SoulChat: Improving LLMs' empathy, listening, and comfort abilities through fine-tuning with multi-turn empathy conversations. In: Proceedings of Findings of the Association for Computational Linguistics. Singapore: Association for Computational Linguistics, 2023. 1170−1183.
    [56] Lai T, Shi Y, Du Z, Wu J, Fu K, Dou Y, et al. Psy-LLM: Scaling up global mental health psychological services with AI-based large language models. arXiv preprint arXiv: 2307.11991, 2023.
    [57] Yang K, Zhang T, Kuang Z, Xie Q, Huang J, Ananiadou S. MentaLLaMA: Interpretable mental health analysis on social media with large language models. In: Proceedings of the ACM Web Conference 2024. Singapore: ACM, 2024. 4489−4500.
    [58] Wang R, Milani S, Chiu J C, Zhi J, Eack S M, Labrum T, et al. PATIENT-Ψ: Using large language models to simulate patients for training mental health professionals. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Miami, USA: Association for Computational Linguistics, 2024. 12772−12797.
    [59] Xu X, Yao B, Dong Y, Gabriel S, Yu H, Hendler J, et al. Mental-LLM: Leveraging large language models for mental health prediction via online text data. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 2024, 8(1): 1−32
    [60] Wang F, Shen X, Yu J, Xia R. Flexible thinking for multimodal emotional support conversation via reinforcement learning. In: Proceedings of Findings of the Association for Computational Linguistics. Suzhou, China: Association for Computational Linguistics, 2025. 1341−1356.
    [61] Zhang H, Meng Z, Luo M, Han H, Liao L, Cambria E, et al. Towards multimodal empathetic response generation: A rich text-speech-vision avatar-based benchmark. In: Proceedings of the ACM on Web Conference. Sydney, Australia: ACM, 2025. 2872−2881.
    [62] Le H M, Nguyen D T, Vo N T T, Nguyen T D Q, Le N B, Nguyen D M H, et al. Reinforcing trustworthiness in multimodal emotional support systems. arXiv preprint arXiv: 2511.10011, 2025.
    [63] Yang J, Chen J, Wang F, Nie Y, Liu Y, Duan Z, et al. DELTA: Deliberative Multi-Agent reasoning with reinforcement learning for multimodal psychological counseling. arXiv preprint arXiv: 2602.04112, 2026.
    [64] Qiu H, Ma L, Lan Z. PsyGUARD: An automated system for suicide detection and risk assessment in psychological counseling. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Miami, USA: Association for Computational Linguistics, 2024. 4581−4607.
    [65] Gorai J, Shaw D K. A BERT-encoded ensembled CNN model for suicide risk identification in social media posts. Neural Computing and Applications, 2024, 36(18): 10955−10970 doi: 10.1007/s00521-024-09642-w
    [66] Zhou W, Prater L C, Goldstein E V, Mooney S J. Identifying rare circumstances preceding female firearm suicides: Validating a large language model approach. JMIR Mental Health, 2023, 10(1): e49359 doi: 10.2196/preprints.49359
    [67] Wang X, Liu K, Wang C. Knowledge-enhanced pre-training large language model for depression diagnosis and treatment. In: Proceedings of the 9th International Conference on Cloud Computing and Intelligent Systems. Dali, China: IEEE, 2023. 532−536.
    [68] Chen S, Zhang Z, Wu M, Zhu K. Detection of multiple mental disorders from social media with two-stream psychiatric experts. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore: Association for Computational Linguistics, 2023. 9071−9084.
    [69] Gu Y, Zhou Y, Chen Q, Zhou N, Zhou J, Zhou A, et al. Enhancing depression-diagnosis-oriented chat with psychological state tracking. In: Proceedings of the 14th CCF International Conference on Natural Language Processing and Chinese Computing. Singapore: Springer Nature Singapore, 2025. 107−119.
    [70] Patel F, Thakore R, Nandwani I, Bharti S K. Combating depression in students using an intelligent chatbot: A cognitive behavioral therapy. In: Proceedings of the IEEE 16th India Council International Conference. Rajkot, India: IEEE, 2019. 1−4.
    [71] Molloy A, Anderson P L. Engagement with mobile health interventions for depression: A systematic review. Internet Interventions, 2021, 26: 100454 doi: 10.1016/j.invent.2021.100454
    [72] De Angel V, Lewis S, White K, Oetzmann C, Leightley D, Oprea E, et al. Digital health tools for the passive monitoring of depression: A systematic review of methods. NPJ Digital Medicine, 2022, 5(1): 3 doi: 10.1038/s41746-021-00548-8
    [73] Bassi G, Giuliano C, Perinelli A, Forti S, Gabrielli S, Salcuni S. A virtual coach (Motibot) for supporting healthy coping strategies among adults with diabetes: Proof-of-concept study. JMIR Human Factors, 2022, 9(1): e32211 doi: 10.2196/32211
    [74] Terblanche N, Molyn J, De Haan E, Nilsson V O. Coaching at scale: Investigating the efficacy of artificial intelligence coaching. International Journal of Evidence Based Coaching & Mentoring, 2022, 20(2): 20−36
    [75] Kamita T, Matsumoto A, Sun B, Inoue T. Promotion of continuous use of a self-guided mental healthcare system by a chatbot. In: Companion Publication of the 2020 Conference on Computer Supported Cooperative Work and Social Computing. New York, USA: ACM, 2020. 293−298.
    [76] Philippe T J, Sikder N, Jackson A, Koblanski M E, Liow E, Pilarinos A, et al. Digital health interventions for delivery of mental health care: Systematic and comprehensive meta-review. JMIR Mental Health, 2022, 9(5): e35159 doi: 10.2196/35159
    [77] Mauriello M L, Tantivasadakarn N, Mora-Mendoza M A, Lincoln E T, Hon G, Nowruzi P, et al. A suite of mobile conversational agents for daily stress management (Popbots): Mixed methods exploratory study. JMIR Formative Research, 2021, 5(9): e25294 doi: 10.2196/25294
    [78] Fu G, Zhao Q, Li J, Luo D, Song C, Zhai W, et al. Enhancing psychological counseling with large language model: A multifaceted decision-support system for non-professionals. arXiv preprint arXiv: 2308.15192, 2023.
    [79] Prochaska J J, Vogel E A, Chieng A, Kendra M, Baiocchi M, Pajarito S, et al. A therapeutic relational agent for reducing problematic substance use (Woebot): Development and usability study. Journal of Medical Internet Research, 2021, 23(3): e24850 doi: 10.2196/24850
    [80] Ogawa M, Oyama G, Morito K, Kobayashi M, Yamada Y, Shinkawa K, et al. Can AI make people happy? The effect of AI-based chatbot on smile and speech in Parkinson's disease. Parkinsonism & Related Disorders, 2022, 99: 43−46 doi: 10.1016/j.parkreldis.2022.04.018
    [81] Ali M R, Razavi S Z, Langevin R, Al Mamun A, Kane B, Rawassizadeh R, et al. A virtual conversational agent for teens with autism spectrum disorder: Experimental results and design lessons. In: Proceedings of the 20th ACM International Conference on Intelligent Virtual Agents. New York, USA: ACM, 2020. 1−8.
    [82] Lee Y C, Yamashita N, Huang Y, Fu W T. “I hear you, I feel you”: Encouraging deep self-disclosure through a chatbot. In: Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. New York, USA: ACM, 2020. 1−12.
    [83] van Waterschoot J, Hendrickx I, Khan A, Klabbers E, de Korte M, Strik H, et al. BLISS: An agent for collecting spoken dialogue data about health and well-being. In: Proceedings of the 12th Conference on Language Resources and Evaluation. Marseille, France: European Language Resources Association, 2020. 449−458.
    [84] Shin J, Yoon H, Lee S, Park S, Liu Y, Choi J D, et al. FedTherapist: Mental health monitoring with user-generated linguistic expressions on smartphones via federated learning. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore: Association for Computational Linguistics, 2023. 11971−11988.
    [85] Ali M R, Sen T, Kane B, Bose S, Carroll T M, Epstein R, et al. Novel computational linguistic measures, dialogue system and the development of SOPHIE: Standardized online patient for healthcare interaction education. IEEE Transactions on Affective Computing, 2021, 14(1): 223−235 doi: 10.1109/taffc.2021.3054717
    [86] Daher K, Casas J, Abou Khaled O, Mugellini E. Empathic chatbot response for medical assistance. In: Proceedings of the 20th ACM International Conference on Intelligent Virtual Agents. New York, USA: ACM, 2020. 1−3.
    [87] Demasi O, Li Y, Yu Z. A multi-persona chatbot for hotline counselor training. In: Proceedings of Findings of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020. 3623−3636.
    [88] Chen T, Shen Y, Chen X, Zhang L. PsyChatbot: A psychological counseling agent towards depressed Chinese population based on cognitive behavioural therapy. ACM Transactions on Asian and Low-Resource Language Information Processing, 2024, 23(3): 1−18 doi: 10.1145/3676962
    [89] Ludin N, Holt-Quick C, Hopkins S, Stasiak K, Hetrick S, Warren J, et al. A chatbot to support young people during the COVID-19 pandemic in New Zealand: Evaluation of the real-world rollout of an open trial. Journal of Medical Internet Research, 2022, 24(11): e38743 doi: 10.2196/38743
    [90] Lee S, Kim S, Kim M, Kang D, Yang D, Kim H, et al. Cactus: Towards psychological counseling conversations using cognitive behavioral theory. In: Proceedings of Findings of the Association for Computational Linguistics. Miami, USA: Association for Computational Linguistics, 2024. 14245−14274.
    [91] Chen K, Sun Z, Wen Y, Lian H, Gao Y, Li Y. Psy-Insight: Explainable multi-turn bilingual dataset for mental health counseling. arXiv preprint arXiv: 2503.03607, 2025.
    [92] Zhu Z, Wang S, Wang Y, Wu J. Integrating visual modalities with large language models for mental health support. In: Proceedings of the 31st International Conference on Computational Linguistics. Abu Dhabi, UAE: Association for Computational Linguistics, 2025. 8939−8954.
    [93] Gounder K K, Tripathy B K, Wagh D A. A hybrid-multimodal mental health chatbot for psychological counselling. In: Proceedings of the International Conference on Biologically Inspired Techniques in Many-Criteria Decision-Making Technologies. Singapore: Springer Nature Singapore, 2024. 223−232.
    [94] De Nieva J O, Joaquin J A, Tan C B, Marc Te R K, Ong E. Investigating students' use of a mental health chatbot to alleviate academic stress. In: Proceedings of the 6th International ACM In-Cooperation HCI and UX Conference. New York, USA: ACM, 2020. 1−10.
    [95] Eagle T, Blau C, Bales S, Desai N, Li V, Whittaker S. “I donť know what you mean by 'I am anxious”': A new method for evaluating conversational agent responses to standardized mental health inputs for anxiety and depression. ACM Transactions on Interactive Intelligent Systems, 2022, 12(2): 1−23 doi: 10.1145/3488057
    [96] Alhamed F, Bendayan R, Ive J, Specia L. Monitoring depression severity and symptoms in user-generated content: An annotation scheme and guidelines. In: Proceedings of the 14th Workshop on Computational Approaches to Subjectivity, Sentiment, and Social Media Analysis. Bangkok, Thailand: Association for Computational Linguistics, 2024. 227−233.
    [97] Yao X, Mikhelson M, Watkins S C, Choi E, Thomaz E, de Barbaro K. Development and evaluation of three chatbots for postpartum mood and anxiety disorders. arXiv preprint arXiv: 2308.07407, 2023.
    [98] Kim T, Bae S, Kim H A, Lee S W, Hong H, Yang C, et al. MindfulDiary: Harnessing large language model to support psychiatric patients' journaling. In: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. New York, USA: ACM, 2024. 1−20.
    [99] Tao Y, Yang M, Shen H, Yang Z, Weng Z, Hu B. Classifying anxiety and depression through LLMs virtual interactions: A case study with ChatGPT. In: Proceedings of the 2023 IEEE International Conference on Bioinformatics and Biomedicine. Istanbul, Turkiye: IEEE, 2023. 2259−2264.
    [100] Alambo A, Gaur M, Lokala U, Kursuncu U, Thirunarayan K, Gyrard A, et al. Question answering for suicide risk assessment using Reddit. In: Proceedings of the 2019 IEEE 13th International Conference on Semantic Computing. Newport Beach, USA: IEEE, 2019. 468−473.
    [101] Gabrielli S, Rizzi S, Bassi G, Carbone S, Maimone R, Marchesoni M, et al. Engagement and effectiveness of a healthy-coping intervention via chatbot for university students during the COVID-19 pandemic: Mixed methods proof-of-concept study. JMIR mHealth and uHealth, 2021, 9(5): e27965 doi: 10.2196/27965
    [102] Kim J, Muhic J, Robert L P, Park S Y. Designing chatbots with black americans with chronic conditions: Overcoming challenges against COVID-19. In: Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. New York, USA: ACM, 2022. 1−17.
    [103] Maharjan R, Doherty K, Rohani D A, Bækgaard P, Bardram J E. Experiences of a speech-enabled conversational agent for the self-report of well-being among people living with affective disorders: An in-the-wild study. ACM Transactions on Interactive Intelligent Systems, 2022, 12(2): 1−29 doi: 10.1145/3484508
    [104] Maharjan R, Rohani D A, Doherty K, Bækgaard P, Bardram J E. What is the difference? Investigating the self-report of wellbeing via conversational agent and web app. IEEE Pervasive Computing, 2022, 21(2): 60−68 doi: 10.1109/MPRV.2022.3147374
    [105] Maeng W, Lee J. Designing and evaluating a chatbot for survivors of image-based sexual abuse. In: Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. New York, USA: ACM, 2022. 1−21.
    [106] Park H, Lee J. Designing a conversational agent for sexual assault survivors: Defining burden of self-disclosure and envisioning survivor-centered solutions. In: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. New York, USA: ACM, 2021. 1−17.
    [107] Beilharz F, Sukunesan S, Rossell S L, Kulkarni J, Sharp G. Development of a positive body image chatbot (KIT) with young people and parents/carers: Qualitative focus group study. Journal of Medical Internet Research, 2021, 23(6): e27807 doi: 10.2196/27807
    [108] Chopra M, Chatterjee A, Dey L, Das P P. Deciphering psycho-social effects of eating disorder: Analysis of Reddit posts using large language models and topic modeling. In: Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities. Miami, USA: Association for Computational Linguistics, 2024. 156−164.
    [109] Han H J, Mendu S, Jaworski B K, Owen J E, Abdullah S. PTSDialogue: Designing a conversational agent to support individuals with post-traumatic stress disorder. In: Proceedings of the 2021 ACM International Joint Conference on Pervasive and Ubiquitous Computing and Proceedings of the 2021 ACM International Symposium on Wearable Computers. New York, USA: ACM, 2021. 198−203.
    [110] Marziali R A, Franceschetti C, Dinculescu A, Nistorescu A, Kristály D M, Mo?oi A A, et al. Reducing loneliness and social isolation of older adults through voice assistants: Literature review and bibliometric analysis. Journal of Medical Internet Research, 2024, 26: e50534 doi: 10.2196/50534
    [111] Razavi S Z, Schubert L K, Van Orden K A, Ali M R. Discourse behavior of older adults interacting with a dialogue agent competent in multiple topics. ACM Transactions on Interactive Intelligent Systems, 2022, 12(2): 1−21 doi: 10.1145/3484510
    [112] Maharjan R, Rohani D A, B?kgaard P, Bardram J E, Doherty K. Can we talk? Design implications for the questionnaire-driven self-report of health and wellbeing via conversational agent. In: Proceedings of the 3rd Conference on Conversational User Interfaces. Bilbao, Spain (Online): ACM, 2021. 1−11.
    [113] Boyd K, Potts C, Bond R R, Mulvenna M, Broderick T, Burns C, et al. Usability testing and trust analysis of a mental health and wellbeing chatbot. In: Proceedings of the 33rd European Conference on Cognitive Ergonomics. Kaiserslautern, Germany: ACM, 2022. 1−8.
    [114] Zhao Y, Liu D, Wan C, Liu X, Nie J Y, Liu J. JMS-QA: A joint hierarchical architecture for mental health question answering. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023, 32: 352−363
    [115] Liu S, Deng N, Sabour S, Jia Y, Huang M, Mihalcea R. Task-adaptive tokenization: Enhancing long-form text generation efficacy in mental health and beyond. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore: Association for Computational Linguistics, 2023. 15264−15281.
    [116] Li J, Sun X. A syntactically constrained bidirectional-asynchronous approach for emotional conversation generation. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Brussels, Belgium: Association for Computational Linguistics, 2018. 678−683.
    [117] Shen L, Feng Y. CDL: Curriculum dual learning for emotion-controllable response generation. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020. 556−566.
    [118] Pei J, Li C. S2SPMN: A simple and effective framework for response generation with relevant information. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Brussels, Belgium: Association for Computational Linguistics, 2018. 745−750.
    [119] Das A, Selek S, Warner A R, Zuo X, Hu Y, Kuttichi Keloth V, et al. Conversational bots for psychotherapy: A study of generative transformer models using domain-specific dialogues. In: Proceedings of the 21st Workshop on Biomedical Language Processing. Dublin, Ireland: Association for Computational Linguistics, 2022. 285−297.
    [120] Young S, Gašić M, Thomson B, Williams J D. Pomdp-based statistical spoken dialog systems: A review. Proceedings of the IEEE, 2013, 101(5): 1160−1179 doi: 10.1109/JPROC.2012.2225812
    [121] Zhou H, Huang M, Zhang T, Zhu X, Liu B. Emotional chatting machine: Emotional conversation generation with internal and external memory. In: Proceedings of the AAAI Conference on Artificial Intelligence. New Orleans, USA: AAAI Press, 2018. 730−739.
    [122] Schatzmann J, Weilhammer K, Stuttle M N, Young S J. A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies. The Knowledge Engineering Review, 2006, 21(2): 97−126 doi: 10.1017/s0269888906000944
    [123] Wen T H, Gasic M, Mrkšić N, Su PH, Vandyke D, Young S. Semantically conditioned LSTM-based natural language generation for spoken dialogue systems. In: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Lisbon, Portugal: Association for Computational Linguistics, 2015. 1711−1721.
    [124] Bosselut A, Rashkin H, Sap M, Malaviya C, Celikyilmaz A, Choi Y. COMET: Commonsense transformers for automatic knowledge graph construction. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics, 2019. 4762−4779.
    [125] Bahdanau D. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv: 1409.0473, 2014.
    [126] Devlin J, Chang M W, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Minneapolis, USA: Association for Computational Linguistics, 2019. 4171−4186.
    [127] Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, et al. RoBERTa: A robustly optimized BERT pretraining approach. In: Proceedings of the 8th International Conference on Learning Representations. Virtual Event: OpenReview.net, 2020.
    [128] Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language models are unsupervised multitask learners[Online], available: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf, February 14, 2019.
    [129] Ji S, Zhang T, Ansari L, Fu J, Tiwari P, Cambria E. MentalBERT: Publicly available pretrained language models for mental healthcare. In: Proceedings of the Thirteenth Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association, 2022. 7184−7190.
    [130] Ji S, Zhang T, Yang K, Ananiadou S, Cambria E, Tiedemann J. Domain-specific continued pretraining of language models for capturing long context in mental health. arXiv preprint arXiv: 2304.10447, 2023.
    [131] Dai Z, Yang Z, Yang Y, Carbonell J, Le Q, Salakhutdinov R. Transformer-XL: Attentive language models beyond a fixed-length context. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics, 2019. 2978−2988.
    [132] Beltagy I, Peters M E, Cohan A. Longformer: The long-document transformer. arXiv preprint arXiv: 2004.05150, 2020.
    [133] Zhai W, Qi H, Zhao Q, Li J, Wang Z, Wang H, et al. Chinese MentalBERT: Domain-adaptive pre-training on social media for Chinese mental health text analysis. In: Proceedings of Findings of the Association for Computational Linguistics. Bangkok, Thailand: Association for Computational Linguistics, 2024. 10574−10585.
    [134] Cui Y, Che W, Liu T, Qin B, Yang Z. Pre-training with whole word masking for Chinese BERT. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29: 3504−3514 doi: 10.1109/TASLP.2021.3124365
    [135] 希尔C E. 助人技术: 探索、领悟、行动三阶段模式(第3版). 胡博, 等译. 北京: 中国人民大学出版社, 2013.

    Hill C E. Helping Skills: Facilitating Exploration, Insight, and Action . Translated by Hu Bo, et al. Beijing: China Renmin University Press, 2013.
    [136] Cheng J, Sabour S, Sun H, Chen Z, Huang M. Pal: Persona-augmented emotional support conversation generation. In: Proceedings of Findings of the 61st Annual Meeting of the Association for Computational Linguistics. Toronto, Canada: Association for Computational Linguistics, 2023. 535−554.
    [137] Ji Z, Lee N, Frieske R, Yu T, Su D, Xu Y, et al. Survey of hallucination in natural language generation. ACM Computing Surveys, 2023, 55(12): 1−38
    [138] Hua Y, Liu F, Yang K, Li Z, Na H, Sheu Y H, et al. Large language models in mental health care: A scoping review. Current Treatment Options in Psychiatry, 2025, 12(1): 1−18 doi: 10.2196/preprints.64088
    [139] Giray L. Prompt engineering with ChatGPT: A guide for academic writers. Annals of Biomedical Engineering, 2023, 51(12): 2629−2633 doi: 10.1007/s10439-023-03272-4
    [140] Short C E, Short J C. The artificially intelligent entrepreneur: ChatGPT, prompt engineering, and entrepreneurial rhetoric creation. Journal of Business Venturing Insights, 2023, 19: e388 doi: 10.1016/j.jbvi.2023.e00388
    [141] Liu P, Yuan W, Fu J, Jiang Z, Hayashi H, Neubig G. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 2023, 55(9): 1−35 doi: 10.1145/3560815
    [142] Reynolds L, Mcdonell K. Prompt programming for large language models: Beyond the few-shot paradigm. In: Proceedings of Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems. New York, USA: ACM, 2021. 1−7.
    [143] Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 2022, 35: 24824−24837 doi: 10.59350/hqv3w-d3q61
    [144] Yang K, Ji S, Zhang T, Xie Q, Kuang Z, Ananiadou S. Towards interpretable mental health analysis with large language models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore: Association for Computational Linguistics, 2023. 6056−6077.
    [145] Liu X, McDuff D, Kovacs G, Galatzer-Levy I, Sunshine J, Zhan J, et al. Large language models are few-shot health learners. arXiv preprint arXiv: 2305.15525, 2023.
    [146] Wang H, Wang R, Mi F, Deng Y, Wang Z, Liang B, et al. Cue-CoT: Chain-of-thought prompting for responding to in-depth dialogue questions with LLMs. In: Proceedings of Findings of the Association for Computational Linguistics. Singapore: Association for Computational Linguistics, 2023. 12047−12064.
    [147] Li Z, Chen G, Shao R, Xie Y, Jiang D, Nie L. Enhancing emotional generation capability of large language models via emotional chain-of-thought. arXiv preprint arXiv: 2401.06836, 2024.
    [148] Goleman D. Emotional intelligence. Why it can matter more than IQ. Learning, 1996, 24(6): 49−50
    [149] Chen Z, Lu Y, Wang W Y. Empowering psychotherapy with large language models: Cognitive distortion detection through diagnosis of thought prompting. In: Proceedings of Findings of the Association for Computational Linguistics. Singapore: Association for Computational Linguistics, 2023. 4295−4304.
    [150] Lee Y K, Lee I, Shin M, Bae S, Hahn S. Chain of empathy: Enhancing empathetic response of large language models based on psychotherapy models. arXiv preprint arXiv: 2311.04915, 2023.
    [151] Han G, Liu W, Huang X, Borsari B. Chain-of-interaction: Enhancing large language models for psychiatric behavior understanding by dyadic contexts. In: Proceedings of the 2024 IEEE 12th International Conference on Healthcare Informatics. Orlando, USA: IEEE, 2024. 392−401.
    [152] Wang B, Wang J, Sun Y, Fu X, Zhao Y, Qin B. Psychological counseling cannot be achieved overnight: Automated psychological counseling through multi-session conversations. In: Proceedings of Findings of the Association for Computational Linguistics. San Diego, California, United States: Association for Computational Linguistics, 2026. 16593−16609.
    [153] Wang M, Wang P, Wu L, Yang X, Wang D, Feng S, et al. AnnaAgent: Dynamic evolution agent system with multi-session memory for realistic seeker simulation. In: Proceedings of Findings of the Association for Computational Linguistics. Vienna, Austria: Association for Computational Linguistics, 2025. 23221−23235.
    [154] Hu Y, Tan M, Zhang C, Li Z, Liang X, Yang M, et al. Aptness: Incorporating appraisal theory and emotion support strategies for empathetic response generation. In: Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. Boise, USA: ACM, 2024. 900−909.
    [155] Xu Z, Chen D, Kuang J, Yi Z, Li Y, Shen Y. Dynamic demonstration retrieval and cognitive understanding for emotional support conversation. In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. Washington, USA: ACM, 2024. 774−784.
    [156] Li C, Liu X, Wang Y, Li D, Lan Y, Shen C. Dialogue for prompting: A policy-gradient-based discrete prompt generation for few-shot learning. In: Proceedings of the AAAI Conference on Artificial Intelligence. Vancouver, Canada: AAAI Press, 2024. 18481−18489.
    [157] Wang J, Xu C, Leong C T, Li W, Li J. Muffin: Mitigating unhelpfulness in emotional support conversations with multifaceted AI feedback. In: Proceedings of Findings of the Association for Computational Linguistics. Bangkok, Thailand: Association for Computational Linguistics, 2024. 567−585.
    [158] Lin B, Bouneffouf D, Cecchi G, Varshney K R. Towards healthy AI: Large language models need therapists too. In: Proceedings of the 4th Workshop on Trustworthy Natural Language Processing. Mexico: Association for Computational Linguistics, 2023. 61−70.
    [159] Yang Z, Ren Z, Wang Y, Peng S, Sun H, Zhu X, et al. Enhancing empathetic response generation by augmenting LLMs with small-scale empathetic models. arXiv preprint arXiv: 2402.11801, 2024.
    [160] Hu Z, Wang L, Lan Y, Xu W, Lim E P, Bing L, et al. LLM-adapters: An adapter family for parameter-efficient fine-tuning of large language models. In: Proceedings of Findings of the Association for Computational Linguistics. Singapore: Association for Computational Linguistics, 2023. 5254−5276.
    [161] Qiu H, Lan Z. PsyDial: A large-scale long-term conversational dataset for mental health support. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Vienna, Austria: Association for Computational Linguistics, 2025. 21624−21655.
    [162] Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv: 2307.09288, 2023.
    [163] Huang L, Yu W, Ma W, Zhong W, Feng Z, Wang H, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 2025, 43(2): 1−55 doi: 10.1145/3703155
    [164] Gao Y, Xiong Y, Gao X, Jia K, Pan J, Bi Y, et al. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv: 2312.10997, 2023.
    [165] Du Y, Li S, Torralba A, Tenenbaum J B, Mordatch I. Improving factuality and reasoning in language models through multiagent debate. In: Proceedings of the Forty-first International Conference on Machine Learning. Vienna, Austria: PMLR, 2024. 1−31.
    [166] Qiu H, He H, Zhang S, Li A, Lan Z. Smile: Single-turn to multi-turn inclusive language expansion via ChatGPT for mental health support. In: Proceedings of Findings of the Association for Computational Linguistics. Miami, USA: Association for Computational Linguistics, 2024. 615−636.
    [167] Li A, Ma L, Mei Y, He H, Zhang S, Qiu H, et al. Understanding client reactions in online mental health counseling. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics. Toronto, Canada: Association for Computational Linguistics, 2023. 10358−10376.
    [168] Chen Y, Li Y, Wen M. Chinese psychological QA database and its research problems. In: Proceedings of the 9th International Conference on Dependable Systems and Their Applications. Wulumuqi, China: IEEE, 2022. 786−792.
    [169] Yao B, Shi C, Zou L, Dai L, Wu M, Chen L, et al. D4: A Chinese dialogue dataset for depression-diagnosis-oriented chat. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, 2022. 2438−2459.
    [170] Zhao J, Zhang T, Hu J, Liu Y, Jin Q, Wang X, et al. M3ED: Multi-modal multi-scene multi-label emotional dialogue database. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Dublin, Ireland: Association for Computational Linguistics, 2022. 5699−5710.
    [171] Sun H, Zhao J, Wang X, Zhao S, Zhou J, Wang H, et al. EmotionTalk: An interactive Chinese multimodal emotion dataset with rich annotations. In: Proceedings of Findings of the Association for Computational Linguistics. San Diego, California, United States: Association for Computational Linguistics, 2026. 9054−9071.
    [172] Chen Y, Fan W, Xing X, Pang J, Huang M, Han W, et al. CPED: A large-scale Chinese personalized and emotional dialogue dataset for conversational AI. arXiv preprint arXiv: 2205.14727, 2022.
    [173] Althoff T, Clark K, Leskovec J. Large-scale analysis of counseling conversations: An application of natural language processing to mental health. Transactions of the Association for Computational Linguistics, 2016, 4: 463−476 doi: 10.1162/tacl_a_00111
    [174] Sharma A, Miner A S, Atkins D C, Althoff T. A computational approach to understanding empathy expressed in text-based mental health support. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. Online: Association for Computational Linguistics, 2020. 5263−5276.
    [175] Chen P C, Rohmatillah M, Lin Y T, Chien J T. ConvCounsel: A conversational dataset for student counseling. In: Proceedings of the 27th Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques. Hsinchu, Taiwan: IEEE, 2024. 1−6.
    [176] Hosseini M, Caragea C. It takes two to empathize: One to seek and one to provide. In: Proceedings of the AAAI Conference on Artificial Intelligence. Online: AAAI Press, 2021. 13018−13026.
    [177] Saha T, Chopra S, Saha S, Bhattacharyya P, Kumar P. A large-scale dataset for motivational dialogue system: An application of natural language generation to mental health. In: Proceedings of the 2021 International Joint Conference on Neural Networks. Shenzhen, China: IEEE, 2021. 1−8.
    [178] Racha S, Joshi P, Raman A, Jangid N, Sharma M, Ramakrishnan G, et al. MHQA: A diverse, knowledge intensive mental health question answering challenge for language models. arXiv preprint arXiv: 2502.15418, 2025.
    [179] Rashkin H, Smith E M, Li M, Boureau Y L. Towards empathetic open-domain conversation models: A new benchmark and dataset. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics, 2019. 5370−5381.
    [180] Zheng C, Sabour S, Wen J, Zhang Z, Huang M. Augesc: Dialogue augmentation with large language models for emotional support conversation. In: Proceedings of Findings of the Association for Computational Linguistics. Singapore: Association for Computational Linguistics, 2023. 1552−1568.
    [181] Zheng Z, Liao L, Deng Y, Nie L. Building emotional support chatbots in the era of LLMs. arXiv preprint arXiv: 2308.11584, 2023.
    [182] Wang M, Smith N A, Mitamura T. What is the Jeopardy model? A quasi-synchronous grammar for QA. In: Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL). Prague, Czech Republic: Association for Computational Linguistics, 2007. 22−32.
    [183] Krithara A, Nentidis A, Bougiatiotis K, Paliouras G. BioASQ-QA: A manually curated corpus for Biomedical Question Answering. Scientific Data, 2023, 10(1): 170 doi: 10.1038/s41597-023-02068-4
    [184] Jin Q, Dhingra B, Liu Z, Cohen W W, Lu X. PubMedQA: A dataset for biomedical research question answering. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing. Hong Kong, China: Association for Computational Linguistics, 2019. 2567−2577.
    [185] Pal A, Umapathi L K, Sankarasubbu M. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In: Proceedings of the Conference on Health, Inference, and Learning. Online: PMLR, 2022. 248−260.
    [186] Busso C, Bulut M, Lee C C, Kazemzadeh A, Mower E, Kim S, et al. IEMOCAP: Interactive emotional dyadic motion capture database. Language Resources and Evaluation, 2008, 42(4): 335−359 doi: 10.1007/s10579-008-9076-6
    [187] Poria S, Hazarika D, Majumder N, Naik G, Cambria E, Mihalcea R. Meld: A multimodal multi-party dataset for emotion recognition in conversations. In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics, 2019. 527−536.
    [188] Saha T, Patra A, Saha S, Bhattacharyya P. Towards emotion-aided multi-modal dialogue act classification. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020. 4361−4372.
    [189] Blache P, Antoine S, De Jong D, Huttner L M, Kerr E, Legou T, et al. The badalona Corpus-An audio, video and Neuro-Physiological conversational dataset. In: Proceedings of the Thirteenth Language Resources and Evaluation Conference. Marseille, France: Association for Computational Linguistics, 2022. 5170−5177.
    [190] Park C Y, Cha N, Kang S, Kim A, Khandoker A H, Hadjileontiadis L, et al. K-EmoCon, a multimodal sensor dataset for continuous emotion recognition in naturalistic conversations. Scientific Data, 2020, 7(1): 293 doi: 10.1038/s41597-020-00630-y
    [191] Bengio Y, Ducharme R, Vincent P, Jauvin C. A neural probabilistic language model. Journal of Machine Learning Research, 2003, 3(Feb): 1137−1155
    [192] Papineni K, Roukos S, Ward T, Zhu W J. Bleu: A method for automatic evaluation of machine translation. In: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. Philadelphia, USA: Association for Computational Linguistics, 2002. 311−318.
    [193] Lin C Y. ROUGE: A package for automatic evaluation of summaries. In: Proceedings of the Text Summarization Branches Out. Barcelona, Spain: Association for Computational Linguistics, 2004. 74−81.
    [194] Li J, Galley M, Brockett C, Gao J, Dolan B. A diversity-promoting objective function for neural conversation models. In: Proceedings of the 54th Annual Meeting of North American Chapter of the Association for Computational Linguistics. San Diego, California: Association for Computational Linguistics, 2016. 110−119.
    [195] Zhang T, Kishore V, Wu F, Weinberger K Q, Artzi Y. BERTScore: Evaluating text generation with BERT. In: Proceedings of the International Conference on Learning Representations. Addis Ababa, Ethiopia: OpenReview.net, 2020. 1−40.
    [196] Liu C W, Lowe R, Serban I V, Noseworthy M, Charlin L, Pineau J. How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. In: Proceedings of the 2016 conference on empirical methods in natural language processing. Austin, Texas: Association for Computational Linguistics, 2016. 2122−2132.
    [197] Sun L, Liu Y, Joseph G, Yu Z, Zhu H, Dow S P. Comparing experts and novices for AI data work: Insights on allocating human intelligence to design a conversational agent. In: Proceedings of the AAAI Conference on Human Computation and Crowdsourcing. Online: AAAI Press, 2022. 195−206.
    [198] Zhao W, Zhao Y, Lu X, Wang S, Tong Y, Qin B. Is ChatGPT equipped with emotional dialogue capabilities? arXiv preprint arXiv: 2304.09582, 2023.
    [199] Peng K, Nisbett R E, Wong N Y. Validity problems comparing values across cultures and possible solutions. Psychological Methods, 1997, 2(4): 329 doi: 10.1037/1082-989x.2.4.329
    [200] Wilson L, Marasoiu M. The development and use of chatbots in public health: Scoping review. JMIR Human Factors, 2022, 9(4): e35882 doi: 10.2196/35882
    [201] Flückiger C, Del Re A C, Wampold B E, Horvath A O. The alliance in adult psychotherapy: A meta-analytic synthesis. Psychotherapy, 2018, 55(4): 316
    [202] Li A, Lu Y, Song N, Zhang S, Ma L, Lan Z. Automatic evaluation for mental health counseling using LLMs. arXiv preprint arXiv: 2402.11958, 2024.
    [203] Arias D, Saxena S, Verguet S. Quantifying the global burden of mental disorders and their economic value. EClinicalMedicine, 2022, 54: 101675 doi: 10.1016/j.eclinm.2022.101675
    [204] Spallek S, Birrell L, Kershaw S, Devine E K, Thornton L. Can we use ChatGPT for mental health and substance use education? Examining its quality and potential harms. JMIR Medical Education, 2023, 9(1): e51243 doi: 10.2196/preprints.51243
    [205] Farhat F. ChatGPT as a complementary mental health resource: A boon or a bane. Annals of Biomedical Engineering, 2024, 52(5): 1111−1114 doi: 10.1007/s10439-023-03326-7
    [206] Wei Y, Guo L, Lian C, Chen J. ChatGPT: Opportunities, risks and priorities for psychiatry. Asian Journal of Psychiatry, 2023, 90: 103808 doi: 10.1016/j.ajp.2023.103808
    [207] Blease C, Worthen A, Torous J. Psychiatrists' experiences and opinions of generative artificial intelligence in mental healthcare: An online mixed methods survey. Psychiatry Research, 2024, 333: 115724 doi: 10.1016/j.psychres.2024.115724
    [208] Bender E M, Gebru T, McMillan-Major A, Shmitchell S. On the dangers of stochastic parrots: Can language models be too big? In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. New York, USA: ACM, 2021. 610−623.
    [209] Elyoseph Z, Gur T, Haber Y, Simon T, Angert T, Navon Y, et al. An ethical perspective on the democratization of mental health with generative AI. JMIR Mental Health, 2024, 11: e58011 doi: 10.2196/58011
    [210] Ma Z, Mei Y, Su Z. Understanding the benefits and challenges of using large language model-based conversational agents for mental well-being support. In: AMIA Annual Symposium Proceedings. New York, USA: American Medical Informatics Association, 2024. 1105−1114.
    [211] Perlis R H, Goldberg J F, Ostacher M J, Schneck C D. Clinical decision support for bipolar depression using large language models. Neuropsychopharmacology, 2024, 49(9): 1412−1416 doi: 10.1038/s41386-024-01841-2
    [212] Filippova K. Controlled hallucinations: Learning to generate faithfully from noisy data. In: Proceedings of Findings of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020. 864−870.
    [213] Harrigian K, Aguirre C, Dredze M. Do models of mental health based on social media data generalize? In: Proceedings of Findings of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020. 3774−3788.
    [214] Hua Y, Jiang H, Lin S, Yang J, Plasek J M, Bates D W, et al. Using Twitter data to understand public perceptions of approved versus off-label use for COVID-19-related medications. Journal of the American Medical Informatics Association, 2022, 29(10): 1668−1678 doi: 10.1093/jamia/ocac114
    [215] Jin H, Chen S, Dilixiati D, Jiang Y, Zhu K Q, Wu M. PsyEval: A comprehensive large language model evaluation benchmark for mental health. npj Mental Health Research, 2026, DOI: 10.1038/s44184-026-00227-0.
    [216] Lamichhane B. Evaluation of ChatGPT for NLP-based mental health applications. arXiv preprint arXiv: 2303.15727, 2023.
    [217] Yongsatianchot N, Torshizi P G, Marsella S. Investigating large language models' perception of emotion using appraisal theory. In: Proceedings of the 11th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos. Cambridge, USA: IEEE, 2023. 1−8.
    [218] Grabb D. The impact of prompt engineering in large language model performance: A psychiatric example. Journal of Medical Artificial Intelligence, 2023, 6: Article No. 20 doi: 10.21037/jmai-23-71
    [219] Vats A, Liu Z, Su P, Paul D, Ma Y, Pang Y, et al. Recovering from privacy-preserving masking with large language models. In: Proceedings of the 2024 IEEE International Conference on Acoustics, Speech and Signal Processing. Seoul, South Korea: IEEE, 2024. 10771−10775.
    [220] Sarwar N, Roy Dipta S. FedMentor: Domain-aware differential privacy for heterogeneous federated LLMs in mental health. arXiv preprint arXiv: 2509.14275, 2025.
    [221] Sarwar S M. Fedmentalcare: Towards privacy-preserving fine-tuned LLMs to analyze mental health status using federated learning framework. arXiv preprint arXiv: 2503.05786, 2025.
    [222] Yao Y, Duan J, Xu K, Cai Y, Sun Z, Zhang Y. A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly. High-Confidence Computing, 2024, 4(2): 100211 doi: 10.1016/j.hcc.2024.100211
    [223] Liao J, Li H, Yang R, Nair S V, Zhu W, Liew J C K, et al. Bias and equity in LLM applications for healthcare: A scoping review. SSRN preprint, DOI: 10.2139/ssrn.6053515,2026.
    [224] Han T, Song W, Ding Z, Li Z, Fang C, Li Y, et al. Debiasing LLMs by masking unfairness-driving attention heads. In: Proceedings of Findings of the Association for Computational Linguistics. San Diego, California, United States: Association for Computational Linguistics, 2026. 37980−37992.
    [225] Cheng X, Chen R, Zan H, Jia Y, Peng M. Biasfilter: An inference-time debiasing framework for large language models. In: Proceedings of Findings of the Association for Computational Linguistics. Suzhou, China: Association for Computational Linguistics, 2025. 15187−15205.
    [226] Rehman M U, Muzammil S. Towards fairer AI: Multi-agent debiasing of LLMs with online evidence retrieval. Proceedings of the AAAI Symposium Series, 2025, 7(1): 103−108 doi: 10.1609/aaaiss.v7i1.36874
    [227] Baidal M, Derner E, Oliver N. Guardians of trust: Risks and opportunities for LLMs in mental health. In: Proceedings of the Fourth Workshop on NLP for Positive Impact. Vienna, Austria: Association for Computational Linguistics, 2025. 11−22.
    [228] Arnaiz-Rodriguez A, Baidal M, Derner E, Annable J L, Ball M, Ince M, et al. Between help and harm: An evaluation study of mental health crisis handling by large language models. JMIR Mental Health, 2026, 13: Article No. e88435 doi: 10.2196/88435
    [229] Weber S, Klebe D, Wolf L, Aeberli C, Homan S, Psathakis N, et al. Suicide- and crisis-risk detection using large language models in mental-health chatbots. medRxiv preprint, DOI: 10.64898/2026.01.12.26343914,2026.
    [230] Li L, Kong S, Zhao H, Li C, Teng Y, Wang Y. Chain of Risks Evaluation (CORE): A framework for safer large language models in public mental health. Psychiatry and Clinical Neurosciences, 2025, 79(6): 299−305 doi: 10.1111/pcn.13781
    [231] Peng B, Galley M, He P, Cheng H, Xie Y, Hu Y, et al. Check your facts and try again: Improving large language models with external knowledge and automated feedback. arXiv preprint arXiv: 2302.12813, 2023.
    [232] Muqtadir A, Bilal H S M, Yousaf A, Ahmed H F, Hussain J. Mitigating hallucinations using ensemble of knowledge graph and vector store in large language models to enhance mental health support. arXiv preprint arXiv: 2410.10853, 2024.
    [233] Kim S, Kim J, Shin S, Chung H, Moon D, Kwon Y, et al. Being kind isn't always being safe: Diagnosing affective hallucination in LLMs. In: Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics. Rabat, Morocco: Association for Computational Linguistics, 2025. 1−39.
    [234] Sha Y, Pan H, Luo G, Shi C, Wang J, Li K. MDD-Thinker: Towards large reasoning models for major depressive disorder diagnosis. arXiv preprint arXiv: 2509.24217, 2025.
    [235] Beurer-Kellner L, Fischer M, Vechev M. Guiding LLMs the right way: Fast, non-invasive constrained generation. In: Proceedings of the International Conference on Machine Learning. Vienna, Austria: PMLR, 2024. 3658−3673.
    [236] Krause B, Gotmare A D, McCann B, Keskar N S, Joty S, Socher R, et al. Gedi: Generative discriminator guided sequence generation. In: Proceedings of Findings of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2021. 4929−4952.
    [237] Mumuni F, Mumuni A. Explainable artificial intelligence (XAI): From inherent explainability to large language models. arXiv preprint arXiv: 2501.09967, 2025.
    [238] Gou Z, Shao Z, Gong Y, Shen Y, Yang Y, Duan N, et al. Critic: Large language models can self-correct with tool-interactive critiquing. In: Proceedings of the 12th International Conference on Learning Representations. Vienna, Austria: OpenReview.net, 2024. 1−18.
    [239] Leiser F, Eckhardt S, Leuthe V, Knaeble M, Maedche A, Schwabe G, et al. Hill: A hallucination identifier for large language models. In: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. Honolulu, USA: ACM, 2024. 1−13.
    [240] Dhuliawala S, Komeili M, Xu J, Raileanu R, Li X, Celikyilmaz A, et al. Chain-of-verification reduces hallucination in large language models. In: Proceedings of Findings of the Association for Computational Linguistics. Bangkok, Thailand: Association for Computational Linguistics, 2024. 3563−3578.
    [241] Ali M, Mabrouk M, Taha Z. Hallucination mitigation techniques in large language models. International Journal of Intelligent Computing & Information Sciences, 2024, 24(4): 73−81
    [242] Hu H, Ma C, Wang Q, Lin L, Zhou Y, Cui L, et al. TheraMind: A strategic and adaptive agent for longitudinal psychological counseling. In: Proceedings of the ACM Web Conference 2026. Dubai, United Arab Emirates: Association for Computing Machinery, 2026. 9136−9147.
    [243] Xiao M, Yang K, Zhao P, Zhang E, Kuang Z, Liu Z, et al. MentraSuite: Post-Training large language models for mental health reasoning and assessment. arXiv preprint arXiv: 2512.09636, 2025.
    [244] Ayers J W, Poliak A, Dredze M, Leas E C, Zhu Z, Kelley J B, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine, 2023, 183(6): 589−596 doi: 10.1001/jamainternmed.2023.1838
    [245] Li Y, Li Z, Zhang K, Dan R, Jiang S, Zhang Y. ChatDoctor: A medical chat model fine-tuned on a large language model Meta-AI (Llama) using medical domain knowledge. Cureus, 2023, 15(6): e40895 doi: 10.7759/cureus.40895
    [246] Paulus A, Zharmagambetov A, Guo C, Amos B, Tian Y. Advprompter: Fast adaptive adversarial prompting for LLMs. In: Proceedings of the 42nd International Conference on Machine Learning. Vancouver, Canada: PMLR, 2025. 48439−48469.
    [247] Wang Y, Zhao Y, Petzold L. Are large language models ready for healthcare? A comparative study on clinical language understanding. In: Proceedings of the machine learning for healthcare conference. New York, USA: PMLR, 2023. 804−823.
    [248] Harrer S. Attention is not all you need: The complicated case of ethically using large language models in healthcare and medicine. EBioMedicine, 2023, 90: 104512 doi: 10.1016/j.ebiom.2023.104512
    [249] Zhang J, Xu X, Zhang N, Liu R, Hooi B, Deng S. Exploring collaboration mechanisms for LLM agents: A social psychology view. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Bangkok, Thailand: Association for Computational Linguistics, 2024. 14544−14607.
  • 加载中
计量
  • 文章访问数:  9
  • HTML全文浏览量:  9
  • 被引次数: 0
出版历程
  • 收稿日期:  2025-12-19
  • 录用日期:  2026-07-11
  • 网络出版日期:  2026-09-23

目录

    /

    返回文章
    返回