arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06027cs.CLcs.AIcs.HC

FormBharo:为印度农村设计和评估用于对话式表单填写的语音助手

FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India

Aman Dalmia, Sanskriti Midha, Jigar Doshi

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对印度农村无法读写的人群开发了语音助手FormBharo,结合LLMs与规则控制完成表单填写,发布了相关基准数据集,通过端到端评估确定最优模型配置以平衡准确性、成本与延迟。

中文摘要 AI 辅助

在印度,几乎每项社会福利都以表单申请为起点,但最需要这些福利的人往往不具备读写能力,要触达他们需要通过语音对话。目前这项工作由一线卫生工作者承担,他们每次只能为一名受益人登记,这是对有限人力的低效利用。我们开发了FormBharo(印地语意为“填写表单”),这是一款语音助手,它通过电话通话在严格的延迟和成本预算下完成结构化表单填写,方法是将大型语言模型(LLMs)与确定性的基于规则的验证及流程控制相结合。该助手正在与ARMMAN(印度一家开展大规模母婴移动健康项目的非政府组织)合作进行试点,以登记低收入、说印地语的母亲加入产前和产后护理项目。据我们所知,这是首个为该人群试点的用于填写登记表单的语音助手。我们公开发布了FormVoiceAgentBench,这是一个基准数据集,包含人类录制的印地语音频以及960次模拟通话中的3760次多轮对话测试,用于评估助手的组件(转录、信息提取、回复生成)和在真实声学变化下的端到端表单完成情况。当LLMs接收易出错的真实语音转录本而非参考转录本时,表单完成率会下降约41个百分点。基于规则的控制可纠正许多轮级提取错误,帮助规模更小、成本更低的模型在表单完成方面达到或超过前沿模型的表现。组件性能无法预测端到端性能:GPT-5.5在参考转录本上的轮级提取准确率领先(99.8%),但在表单完成方面排名较低。由于错误会在流程中同时传播和抵消,最优模型组合只能通过端到端评估确定。最后,没有任何单一模型能同时在准确性、成本和延迟三个维度上表现最佳,因此我们采用基于帕累托的加权求和标量化方法来选择兼顾三者的可部署配置。

英文摘要

In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write. Reaching them requires a spoken conversation. Today that work falls to frontline health workers who enroll beneficiaries one at a time, a poor use of stretched capacity. We built FormBharo ("fill the form" in Hindi), a voice agent that fills a structured form over a phone call under tight latency and cost budgets by pairing Large Language Models (LLMs) with deterministic, rule-based validation and flow control. It is being piloted with ARMMAN, an NGO running large-scale maternal and child mobile-health programs in India, to enroll low-income, Hindi-speaking mothers in antenatal and postnatal care. To our knowledge, it is the first voice agent piloted to fill an enrollment form for this population. We openly release FormVoiceAgentBench, a benchmark pairing human-recorded Hindi audio with 3,760 multi-turn conversation tests across 960 simulated calls, to evaluate our agent's components (transcription, extraction, reply generation) and end-to-end form completion under real acoustic variations. Form completion drops by up to ~41 points when LLMs receive error-prone real-speech transcripts instead of reference ones. The rule-based controls recover many turn-level extraction errors, helping smaller, cheaper models match or surpass frontier models on form completion. Component performance does not predict end-to-end performance: GPT-5.5 leads turn-level extraction accuracy on reference transcripts (99.8%) but ranks lower on form completion. Since errors both propagate and cancel across the pipeline, the optimal model choice of models emerges only through end-to-end evaluation. Finally, no single model is best across accuracy, cost, and latency at once, so we use a Pareto-based weighted-sum scalarization to select a deployable configuration balancing the three.

↑