阿里云ATH事业群-语音交互前沿算法工程师-语音交互
社招全职3年以上地点:杭州状态:招聘
工作描述
任职要求 1、计算机、人工智能、语音信号处理等相关专业硕士/博士学位,在对话系统、语音交互等领域具备扎实的理论基础。 2、实战经验与工程能力: ● 掌握Megatron-LM/DeepSpeed等训练框架,具有大模型分布式训练经验;掌握 vLLM/vLLM-Omni/SGLang/TensorRT/Triton等推理框架,了解 KV Cache 管理、Continuous Batching、量化、投机解码等技术。 ● 具有丰富的对话系统/语音对话/全双工交互方向算法经验。 ● 扎实的大模型基础,熟悉LLM的训练与微调全流程,在 SFT/RL/OPD 等方向有深入的实践经验。 ● 熟悉主流语音端到端模型如Moshi及全模态大模型(如 Qwen-Omni等)架构。 ● 有对话策略(谈判、共情、施压等)或强化学习对齐(RLHF / RLVR / AgenticRL)的实战经验。 ● 对 AI 云产品和行业智能化升级有强烈兴趣,能将算法创新与客…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
Megatron+
https://www.youtube.com/watch?v=hc0u4avAkuM
DeepSpeed+
https://www.youtube.com/watch?v=pDGI668pNg0
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
Triton Inference Server+
https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html
Triton Inference Server is an open source inference serving software that streamlines AI inferencing.
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
还有更多 •••
相关职位
社招3年以上
1. 计算机、人工智能、信号处理等相关专业硕士及以上,3 年以上语音、音视频多模态算法工程经验。 2. 扎实的语音、音视频算法和落地经验,熟悉主流模型架构,有微调(SFT/RL)及线上效果调优的完整项
更新于 2026-08-28北京|杭州
社招5年以上技术类-算法
1. 硕士及以上学历,计算机、人工智能、软件工程、数学、自动化等相关专业优先; 2. 深入理解 Transformer 架构及大语言模型基础知识,熟悉模型评测方案(Evaluation)或具有后训练(
更新于 2026-08-14杭州
社招5年以上技术类-算法
1. 硕士及以上学历,计算机、人工智能、软件工程、数学、自动化等相关专业优先; 2. 深入理解 Transformer 架构及大语言模型基础知识,熟悉模型评测方案(Evaluation)或具有后训练(
更新于 2026-08-14杭州