面向中药视觉问答的多任务学习方法

Research on multi-task learning methods for traditional chinese medicine visual question answering

  • 摘要: 随着人工智能技术的快速发展,视觉问答(visual question answering, VQA)技术在中医药领域的应用逐渐兴起,给定一幅中药图像,模型即可回答与该中药相关的问题. 借助VQA技术,人们能够更加便捷、直观地认识中药,有助于中医药文化的传播与推广. 然而,目前中药VQA领域面临数据集严重匮乏、现有VQA模型适配度不高等问题. 为此,构建了一个专门的中药VQA数据集,并在此基础上提出一种基于多任务学习的中药VQA模型TCMML. TCMML采用Faster R-CNN与Chinese BERT分别提取图像和文本特征,并通过基于自注意力与交叉注意力的端到端联合注意力网络实现多模态特征融合. 此外,模型引入多任务学习策略:通过任务共享层进行跨模态对齐并学习任务间的关联信息,同时由5个任务专家模块分别应对中药VQA的5个子任务,最终实现高精度的答案预测. 实验结果表明,TCMML在中药VQA任务上取得了较高的准确率,相较于现有主流模型具有明显的优势,验证了多任务学习策略在该任务中的有效性.

     

    Abstract: With the rapid development of artificial intelligence technology, visual question answering (VQA) has gradually emerged in the field of traditional Chinese medicine (TCM). Given a TCM image, the model can answer questions related to the herb. By leveraging VQA technology, people can gain a more convenient and intuitive understanding of TCM, facilitating the dissemination and promotion of TCM culture. However, the current TCM VQA field faces severe dataset shortages and low adaptability of existing VQA models. To address this, a specialized TCM VQA dataset was constructed, and a multi-task learning-based TCM VQA model (TCMML) was proposed. The model employs Faster R-CNN and Chinese BERT to extract image and text features, respectively, and achieves multimodal feature fusion through an end-to-end joint attention network based on self-attention and cross-attention. Additionally, the model incorporates a multi-task learning strategy: a task-shared layer enables cross-modal alignment and learning of task correlations, while five task-specific modules handle five subtasks of TCM VQA, ultimately achieving high-precision answer prediction. Experimental results demonstrate that TCMML achieves high accuracy in TCM VQA tasks, significantly outperforming mainstream models, validating the effectiveness of the multi-task learning strategy in this task.

     

/

返回文章
返回