logo of tesla

特斯拉Voice AI / LLM-based Voice Agent Engineer

社招全职AI与数据分析地点:上海状态:招聘

任职要求


计算机、人工智能、语音识别自然语言处理、软件工程等相关专业背景。
熟悉 LLM 应用开发,理解 Prompt Engineering、RAG、Function Calling、Agent Workflow、多轮对话管理等核心技术。
熟悉语音对话机器人链路,了解 ASRTTS、VAD、IVR、呼叫中心、实时音频流处理等相关技术。
有客服机器人、电话机器人、Voicebot、智能客服、语音助手或任务型对话系统相关经验者优先。
熟悉 Python,具备良好的工程实现能力,能够独立完成算法原型、服务化接口、日志分析和效果评估。
理解大模型在业务场景中的风险,包括幻觉、误答、流程偏离、上下文丢失、工具调用错误等,并具备相应治理思路。
具备较强的问题拆解能力和业务理…
登录查看完整任职要求
微信扫码,1秒登录

工作职责


岗位背景
我们正在建设LLM-based Voice Agent。该系统将工具调用与业务系统集成,实现自然语音交互、意图理解、任务推进、问题解决与必要时的真人交互转接。该岗位将负责从语音识别、对话理解、LLM Agent 编排到端到端语音体验优化的核心能力建设。



设计并优化端到端语音交互系统设计,包括:用户打断(Barge-in)、流式语音输入输出 、多轮语音对话管理/轮次切换 、低延迟语音链路优化 、语音识别纠错等。
基于大语言模型构建语音 Agent 能力,包括任务型对话、多轮上下文管理、流程推进、复杂问题拆解、工具调用和异常兜底。
负责售前/售后/客服等业务知识的接入与问答能力建设,包括 FAQ、知识库、RAG、业务规则、工单系统、预约系统等。
优化语音对话体验,包括低延迟流式交互、打断处理、长静音处理、噪声场景鲁棒性、语音播报自然度和对话节奏控制。
建立 Voice Agent 的评测体系,包括 ASR 准确率、意图识别准确率、任务完成率、问题解决率、转人工率、幻觉率、响应延迟和用户满意度等指标。
与产品、业务、后端、电话系统、IVR、呼叫中心平台等团队协作,推动 Voice Agent 从 PoC、灰度测试到正式上线。
持续分析线上语音对话日志,定位失败案例,优化 Prompt、Agent Workflow、知识召回、模型策略和业务规则。
有以下经验优先:
WebRTC / SIP / 电话系统
呼叫中心或 IVR 系统
RAG / Agent Workflow / Function Calling
实时语音或实时 LLM 推理优化
包括英文材料
语音识别+
NLP+
大模型+
Prompt+
RAG+
AI agent+
语音合成+
还有更多 •••
相关职位

logo of amap
实习高德研究型实习生

团队介绍: 高德语音技术部,是负责高德全栈语音技术的综合性团队。团队核心技术能力包括:自研TTS基座大模型、端侧模型、多语种、RTC流式语音、语音内容生成、语音识别、多模态模型、模型服务与推理。业务支撑面向高德全部核心场景,包括语音导航、AI领航员、IP语音定制、国际化、AI语音助手、智能外呼、内容生成等。 团队定位是通过前沿语音技术的研究和落地,赋能下一代AI产品创新。近期部分技术(https://arxiv.org/abs/2507.12197https://arxiv.org/abs/2507.12197)和产品进展介绍(https://mp.weixin.qq.com/s/cCeHbNW0jbC_LNVPZlGeHg)https://mp.weixin.qq.com/s/cCeHbNW0jbC_LNVPZlGeHghttps://arxiv.org/abs/2507.12197)和产品进展介绍(https://mp.weixin.qq.com/s/cCeHbNW0jbC_LNVPZlGeHg) 具体职责: 围绕voice agent/speech language model的研究工作,包括但不限于如下事项: 跟进最领先的语音交互技术,包括但不限于提出新的技术框架、改进现有的算法、持续提升相关技术及业务指标,鼓励撰写论文及申请专利; 结合业务场景,探索跨模态(文字/语音/视觉)混合训练的最佳实践,探索基于speech language model的后训练(SFT+RL)技术,持续优化交互响应、交互内容,结合规划agent/工具调用agent,持续提升voice agent的交互体验,从而反馈到高德agent的整体能力; 探索流式全双工对话中,更加高效且合理的模型架构,包括但不限于COT Reasoning in streaming full-duplex等; 海量的语音数据,尤其是对话数据的处理构建:定性分析、定量评估、参与设计自动评估框架,研发 scalable 的改进方案,持续提升数据质量;

更新于 2026-02-04北京
logo of lilith
社招本地化

1. Prepare and verify the voiceover materials to ensure that no linguistic or cultural issues slip into the voiced materials. 2. Independently oversee voiceover recording sessions and timely correct any issues in performance and script delivery. 3. Assist in the communication between the local development team and English-speaking voice-over directors during monitored sessions. 4. Direct and guide voice-over actors during live sessions, helping them deliver their best performance. 5. Conduct verification and in-game testing of voiceover materials. 6. Create characters' voice-over designs independently. 7. Champion voice-over best practices in all aspects of development and operations.

更新于 2026-07-07上海
logo of microsoft
社招Applied

OverviewThe CoreAI Voice Agent team brings together talents in the areas of signal processing, speech modeling, statistical modeling and deep learning to develop and deliver robust, natural and scalable speech technologies, across a rich set of scenarios and languages.  We welcome Applied Scientists to join our cutting-edge voice agent team. You’ll work at the intersection of deep learning, signal processing, and speech/audio modeling to push the boundaries of natural, expressive, and multilingual speech generation. Your innovations will power next-generation products in conversational AI and accessibility—impacting and enriching the lives of millions of users worldwide.   Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.   In alignment with our Microsoft values, we are committed to cultivating an inclusive work environment for all employees to positively impact our culture every day. Responsibilities• Design and develop novel speech algorithms to advance state-of-the-art technologies, with a focus on real-world applications in speech generation • Tackle scalability challenges by aligning solutions with evolving stakeholder requirements • Leverage large-scale computing frameworks, data analysis platforms, and modeling environments to enhance model performance and efficiency • Deploy models into production environments and iterate based on empirical results and user impact • Conduct rigorous experimentation by evaluating multiple models in live scenarios to assess comparative performance • Continuously monitor deployed algorithms to ensure they meet expected behavior, accuracy thresholds, and performance guardrails

更新于 2026-06-30北京|苏州
logo of microsoft
社招Software

OverviewMicrosoft Azure Voice AI | Microsoft Foundry Agent PlatformJoin our innovative Azure Voice AI team, where we are building voice, avatar, and voice agent technologies as foundational capabilities of the Microsoft Foundry Agent Platform.Our mission is to empower every person and every organization on the planet with human‑like, diverse, and delightful AI voices, avatars, and voice‑first agents, enabling natural and effective interactions across devices, applications, and platforms.We are seeking passionate and talented engineers to join our growing team. In this role, you will work on core voice agent technologies, spanning system design, engineering, integration, and continuous refinement of speech, avatar, and agent components in production environments.You will collaborate closely with experts across AI, platform engineering, and product teams to deliver end‑to‑end voice agent solutions that scale across Microsoft products and partners. Responsibilities• Design and build voice and speech technologies (e.g., ASR, TTS, real‑time interaction pipelines) that power voice agents within the Microsoft Foundry Agent Platform • Implement and evolve voice agent capabilities, including real‑time speech processing, multimodal orchestration, and agent runtime integration • Develop and integrate talking avatar technologies, supporting both zero‑shot and customized experiences • Apply and adapt modern generative AI technologies (including diffusion and language modeling) where appropriate, with an emphasis on engineering robustness and system integration • Collaborate with cross‑functional teams to integrate voice agents, speech services, and avatar components into platforms, SDKs, and end‑to‑end applications • Continuously improve system latency, quality, scalability, and reliability in real‑world deployments • Stay current with industry and research advancements in voice AI, agentic systems, and generative technologies, and translate them into practical solutions • Contribute to engineering excellence through design reviews, code reviews, and shared best practices

更新于 2026-07-17北京