高德地图voice agent算法实习生
任职要求
职位要求: 985/211高校研究生及以上学历或优秀本科生,计算机、人工智能、软件、数学等相关专业,有语音、自然语言处理、多模态等背景; 在语音领域(包括但不限于语音对话 / TTS / ASR)有一线的实践经验; 熟练掌握C/C++,Python,Shell编程语言,对数据结构和算法设计有较好的理解; 熟悉 Pytorch / megatron等深度学习框架,熟悉 Transformer 架构以及大语言模型基础…
工作职责
团队介绍: 高德语音技术部,是负责高德全栈语音技术的综合性团队。团队核心技术能力包括:自研TTS基座大模型、端侧模型、多语种、RTC流式语音、语音内容生成、语音识别、多模态模型、模型服务与推理。业务支撑面向高德全部核心场景,包括语音导航、AI领航员、IP语音定制、国际化、AI语音助手、智能外呼、内容生成等。 团队定位是通过前沿语音技术的研究和落地,赋能下一代AI产品创新。近期部分技术(https://arxiv.org/abs/2507.12197https://arxiv.org/abs/2507.12197)和产品进展介绍(https://mp.weixin.qq.com/s/cCeHbNW0jbC_LNVPZlGeHg)https://mp.weixin.qq.com/s/cCeHbNW0jbC_LNVPZlGeHghttps://arxiv.org/abs/2507.12197)和产品进展介绍(https://mp.weixin.qq.com/s/cCeHbNW0jbC_LNVPZlGeHg) 具体职责: 围绕voice agent/speech language model的研究工作,包括但不限于如下事项: 跟进最领先的语音交互技术,包括但不限于提出新的技术框架、改进现有的算法、持续提升相关技术及业务指标,鼓励撰写论文及申请专利; 结合业务场景,探索跨模态(文字/语音/视觉)混合训练的最佳实践,探索基于speech language model的后训练(SFT+RL)技术,持续优化交互响应、交互内容,结合规划agent/工具调用agent,持续提升voice agent的交互体验,从而反馈到高德agent的整体能力; 探索流式全双工对话中,更加高效且合理的模型架构,包括但不限于COT Reasoning in streaming full-duplex等; 海量的语音数据,尤其是对话数据的处理构建:定性分析、定量评估、参与设计自动评估框架,研发 scalable 的改进方案,持续提升数据质量;
岗位背景 我们正在建设LLM-based Voice Agent。该系统将工具调用与业务系统集成,实现自然语音交互、意图理解、任务推进、问题解决与必要时的真人交互转接。该岗位将负责从语音识别、对话理解、LLM Agent 编排到端到端语音体验优化的核心能力建设。 设计并优化端到端语音交互系统设计,包括:用户打断(Barge-in)、流式语音输入输出 、多轮语音对话管理/轮次切换 、低延迟语音链路优化 、语音识别纠错等。 基于大语言模型构建语音 Agent 能力,包括任务型对话、多轮上下文管理、流程推进、复杂问题拆解、工具调用和异常兜底。 负责售前/售后/客服等业务知识的接入与问答能力建设,包括 FAQ、知识库、RAG、业务规则、工单系统、预约系统等。 优化语音对话体验,包括低延迟流式交互、打断处理、长静音处理、噪声场景鲁棒性、语音播报自然度和对话节奏控制。 建立 Voice Agent 的评测体系,包括 ASR 准确率、意图识别准确率、任务完成率、问题解决率、转人工率、幻觉率、响应延迟和用户满意度等指标。 与产品、业务、后端、电话系统、IVR、呼叫中心平台等团队协作,推动 Voice Agent 从 PoC、灰度测试到正式上线。 持续分析线上语音对话日志,定位失败案例,优化 Prompt、Agent Workflow、知识召回、模型策略和业务规则。 有以下经验优先: WebRTC / SIP / 电话系统 呼叫中心或 IVR 系统 RAG / Agent Workflow / Function Calling 实时语音或实时 LLM 推理优化
OverviewThe CoreAI Voice Agent team brings together talents in the areas of signal processing, speech modeling, statistical modeling and deep learning to develop and deliver robust, natural and scalable speech technologies, across a rich set of scenarios and languages. We welcome Applied Scientists to join our cutting-edge voice agent team. You’ll work at the intersection of deep learning, signal processing, and speech/audio modeling to push the boundaries of natural, expressive, and multilingual speech generation. Your innovations will power next-generation products in conversational AI and accessibility—impacting and enriching the lives of millions of users worldwide. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. In alignment with our Microsoft values, we are committed to cultivating an inclusive work environment for all employees to positively impact our culture every day. Responsibilities• Design and develop novel speech algorithms to advance state-of-the-art technologies, with a focus on real-world applications in speech generation • Tackle scalability challenges by aligning solutions with evolving stakeholder requirements • Leverage large-scale computing frameworks, data analysis platforms, and modeling environments to enhance model performance and efficiency • Deploy models into production environments and iterate based on empirical results and user impact • Conduct rigorous experimentation by evaluating multiple models in live scenarios to assess comparative performance • Continuously monitor deployed algorithms to ensure they meet expected behavior, accuracy thresholds, and performance guardrails
OverviewMicrosoft Azure Voice AI | Microsoft Foundry Agent PlatformJoin our innovative Azure Voice AI team, where we are building voice, avatar, and voice agent technologies as foundational capabilities of the Microsoft Foundry Agent Platform.Our mission is to empower every person and every organization on the planet with human‑like, diverse, and delightful AI voices, avatars, and voice‑first agents, enabling natural and effective interactions across devices, applications, and platforms.We are seeking passionate and talented engineers to join our growing team. In this role, you will work on core voice agent technologies, spanning system design, engineering, integration, and continuous refinement of speech, avatar, and agent components in production environments.You will collaborate closely with experts across AI, platform engineering, and product teams to deliver end‑to‑end voice agent solutions that scale across Microsoft products and partners. Responsibilities• Design and build voice and speech technologies (e.g., ASR, TTS, real‑time interaction pipelines) that power voice agents within the Microsoft Foundry Agent Platform • Implement and evolve voice agent capabilities, including real‑time speech processing, multimodal orchestration, and agent runtime integration • Develop and integrate talking avatar technologies, supporting both zero‑shot and customized experiences • Apply and adapt modern generative AI technologies (including diffusion and language modeling) where appropriate, with an emphasis on engineering robustness and system integration • Collaborate with cross‑functional teams to integrate voice agents, speech services, and avatar components into platforms, SDKs, and end‑to‑end applications • Continuously improve system latency, quality, scalability, and reliability in real‑world deployments • Stay current with industry and research advancements in voice AI, agentic systems, and generative technologies, and translate them into practical solutions • Contribute to engineering excellence through design reviews, code reviews, and shared best practices
1、负责 Voice Agent 中控编排系统(Orchestrator)的设计与落地 2、构建 ASR → NL → LLM → TTS 端到端语音链路并做工程化优化(并发控制、低延迟优化、分片流式处理、错误恢复机制) 3、设计与优化 Prompt 工程、Function Calling、工具调用、Agent 状态机 4、构建 Voice Agent 的“中断/打断”检测体系与决策引擎,包括不限于:音频能量检测、ASR 级别中断词库、LLM-based interrupt classifier、优先级调度。 5、推动 Agent 系统的质量体系建设,包括不限于:自动化评测、回放系统、Agent Trace、模型响应审计、Latency Profiling。 6、深度参与 Voice Agent 的性能优化,如:Token 成本优化、缓存策略、向量库优化、ASR/TTS 服务吞吐提升、服务并发治理。吞吐、并发、缓存、Token 成本 7、跨团队协作,与产品/算法/SRE 共同推进 Voice Agent 场景落地,包括不限于:新功能快速落地、A/B 实验、Agent 行为修正。 8、跟踪 Voice/LLM/Agent 前沿技术,例如:语音大模型(Whisper/Salmonn)、MCP、Multi-Agent、上下文压缩、OpenAI Realtime API