基于动态链式推理的轨道交通表单问答大语言模型

Rail Transit Table Question Answering Large Language Model Based on Dynamic Chain Reasoning

  • 摘要:
    目的 针对轨道交通领域表格与文本混合数据的智能问答难点问题,提出一种基于动态链式推理框架的大语言模型(以下简称“动态链大语言模型”),旨在解决结构化表格与非结构化文本的协同理解问题,为一线工程师提供设备检修、故障诊断等场景下的智能问答工具。
    方法 所提动态链大语言模型通过构建包含情景分析、信息过滤、逻辑构造、智能验证和引擎执行等多阶段推理机制,实现复杂数据的语义协同与精准解析。采用DeepSeek R1模型进行指令微调,结合轨道交通领域专业数据集TRTDD(轨道交通表格问答数据集)完成领域知识深度融合,并通过LoRA(低秩适配)技术降低训练成本。基于HybridQA(混合表格文本问答)基准数据集、TATQA(金融表格文本问答)数据集和TRTDD,对比分析参数量为32 B和70 B的动态链大语言模型同GLM4-9 B-Chat、Qwen2.5-72 B-Instruct、Llama-3.3-70 B-Instruct等通用大语言模型之间的性能差异。
    结果及结论 试验结果表明,所提参数量为32 B、70 B动态链大语言模型在EM(精确匹配)和F1值指标上显著优于通用大语言模型,验证了动态链式推理框架在复杂业务场景下的优越性,可为轨道交通领域提供高效的多模态数据问答解决方案。

     

    Abstract:
    Objective Addressing the challenges of mixed tabular-textual data question answering in the rail transit field, a large language model based on a dynamic chain reasoning framework (hereinafter referred to as the DCLL (dynamic chain large language) model) is proposed, aiming to tackle the collaborative comprehension of structured tabular data and unstructured textual information and to provide frontline engineers with intelligent question-answering tools for scenarios such as equipment maintenance and fault diagnosis.
    Method The proposed DCLL model achieves semantic collaboration and precise parsing of complex data by constructing a multi-stage reasoning mechanism encompassing scenario analysis, information filtering, logic construction, intelligent verification, and engine execution. Instruction fine-tuning is performed using the DeepSeek R1 model, combining the field-specific professional dataset TRTDD (transportation table domain dataset) to complete the deep integration of field knowledge, while employing LoRA (low-rank adaptation)-based parameter-efficient adaptation technology to reduce training costs. Based on the HybridQA (hybrid table and text question answering) benchmark dataset, the TATQA (financial table and text question answering) dataset, and TRTDD, a comparative analysis is conducted to evaluate the performance differences between the DCLL models with parameter sizes of 32B and 70B and general-purpose large language models such as GLM4-9B-Chat, Qwen2.5-72B-Instruct, and Llama-3.3-70B-Instruct.
    Result & Conclusion  Experimental results demonstrate that the proposed DCLL models with parameter sizes of 32B and 70B significantly outperform general-purpose large language models in both EM (exact math) and F1-score metrics. This approach validates the superiority of the dynamic chain reasoning framework in complex operational scenarios, providing an efficient multimodal data question-answering solution for the rail transit field.

     

/

返回文章
返回