阿里云阿里云智能-无影大模型AI系统专家-上海/杭州
社招全职3年以上云智能集团地点:杭州 | 上海状态:招聘
工作描述
任职要求 1. 技术技能 1) 训练加速:精通 DeepSpeed、Megatron-LM 等分布式训练框架,熟悉 3D 并行、ZeRO 优化、混合精度训练等技术。 2) 推理优化:掌握 TensorRT、ONNX Runtime、vLLM 等推理引擎,熟悉 PD 分离架构、KV Cache 管理、投机采样等技术。 3) 硬件适配:熟悉 GPU 架构,精通 CUDA/CUDNN 编程,有算子优化经验者优先。 4) 系统设计:具备分布式系统开发经验,熟悉 Kubernetes、Docker 等容器化技术,有 GPU 集群管理经验者加分。 2. 项目经验 1) 至少主导过 1 个千亿参数模型的训练加速项目,或实现推理延迟降低 50% 以上的工程案例。 3. 其他 1) 具备系统思维与工程化落…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
ONNX+
https://github.com/onnx/tutorials
Open Neural Network Exchange (ONNX) is an open standard format for representing machine learning models.
[英文] Introduction to ONNX
https://onnx.ai/onnx/intro/
This documentation describes the ONNX concepts (Open Neural Network Exchange).
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
推理引擎+
https://www.youtube.com/watch?v=_dvk75LEJ34
https://www.youtube.com/watch?v=XtT5i0ZeHHE
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
CUDA+
https://developer.nvidia.com/blog/even-easier-introduction-cuda/
This post is a super simple introduction to CUDA, the popular parallel computing platform and programming model from NVIDIA.
https://www.youtube.com/watch?v=86FAWCzIe_4
Lean how to program with Nvidia CUDA and leverage GPUs for high-performance computing and deep learning.
系统设计+
https://roadmap.sh/system-design
Everything you need to know about designing large scale systems.
https://www.youtube.com/watch?v=F2FmTdLtb_4
This complete system design tutorial covers scalability, reliability, data handling, and high-level architecture with clear explanations, real-world examples, and practical strategies.
还有更多 •••
相关职位
实习阿里巴巴2027
1. 学历背景:2027 届海内外应届毕业生,博士学历优先,或硕士期间有突出科研成果者;计算机、人工智能、数学、统计学等相关专业; 2. 科研能力: 1)在 AI 顶级会议(NeurIPS/ICML/
更新于 2026-03-17杭州|上海
社招3年以上云智能集团
1. 本科及以上学历,计算机科学、电子工程等相关专业背景; 2. 3 年以上产品管理经验,有云计算、EUC(End User Computing)、终端设备或相关领域经验者优先; 3. 深入了解 AW
更新于 2026-08-21杭州
社招2年以上
1. 2年以上IT、互联网、云计算开发相关工作经验; 2. Java、Go等编程语言基础扎实,熟悉常用的开源框架、中间件原理和机制;熟悉分布式系统的设计和应用,能对分布式常用技术进行合理应用,解决问题
更新于 2026-09-04北京|杭州