发表机构
Digital Touch Point Co., Ltd.; iApp Technology Co., Ltd.(数字触点有限公司; iApp技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出多语言检索增强型泰国健康顾问TSWAP,基于泰国传统医学知识库,采用混合检索等技术,发布相关基准与日志,发现量化及强制检索的部署要点。
AI 中文摘要
我们提出TSWAP,这是一款已部署的八语言会话式健康顾问,通过检索增强生成技术,基于经过验证的泰国传统医学知识库和认证健康服务提供商。未修改的开源权重大语言模型Qwen3.6-35B-A3B(基于vLLM框架),通过混合密集-稀疏检索器结合交叉编码器重排序,被适配到约3.06万个文本块的泰语索引中;首轮查询分类器强制工具检索以进行实体查找;基于规则的安全层管控医疗范围并执行泰国紧急路由;所有八种语言通过“先翻译后检索”方式零样本提供服务。我们发布首个泰国传统医学/健康检索基准(含50个问题及标准答案文档ID,Recall@5=0.88)、生产级问答日志(259个案例中91.1%的重测通过率),以及71个问题的前沿无检索探针,用于展示各基础支撑模块的贡献:无安全提示时,后端模型生成了完整的药物剂量方案并符合超出范围的请求;无知识库时,模型未生成任何可验证的服务提供商推荐。我们还报告了两项可迁移的部署发现:以英语校准的4位AWQ量化会破坏泰语声调符号,强制检索路由是实现可靠基础支撑的必要条件。
英文摘要
We present TSWAP, a deployed eight-language conversational wellness advisor grounded, via retrieval-augmented generation, in a verified knowledge base of Thai traditional medicine and certified wellness providers. An unmodified open-weight LLM (Qwen3.6-35B-A3B on vLLM) is grounded on a ~30.6K-chunk Thai index by a hybrid dense-sparse retriever with cross-encoder reranking; a first-turn query classifier forces tool-based retrieval for entity lookups; a rule-based safety layer enforces medical scope and Thai emergency routing; and all eight languages are served zero-shot with translate-then-retrieve. We release the first Thai traditional-medicine/wellness retrieval benchmark (50 questions with gold document IDs; Recall@5 = 0.88), production QA logs (91.1% test-retest pass over 259 cases), and a 71-question frontier no-retrieval probe showing what each grounding pillar contributes: without the safety prompt the backend model family produced a full drug-dosing schedule and complied with out-of-scope requests, and without the knowledge base it produced zero verifiable provider recommendations. We further report two transferable deployment findings: English-calibrated 4-bit AWQ quantization corrupts Thai tone marks, and forced-retrieval routing is necessary for reliable grounding.
Comments8 pages, 2 tables. Data and evaluation logs: https://huggingface.co/datasets/iapp/tswap-wellness-benchmark