阿里巴巴AI推理平台-大模型服务性能与可观测性技术专家-杭州/北京
社招全职3年以上技术类-开发地点:北京 | 杭州状态:招聘
工作描述
任职要求 1.计算机或相关专业,基础扎实;熟练使用 Python、C++ 中至少一种语言,能够阅读和修改大型工程代码。 2.熟悉 Transformer 推理、Prefill/Decode、动态批处理和 KV Cache;使用过 vLLM、SGLang 或 TensorRT-LLM。 3.理解 CPU/GPU 执行、存储层次和分布式通信,能够使用性能分析工具定位计算、带宽或通信瓶颈。 4.具备性能工程和实验分析能力,能够定义指标、控制变量,并用多类数据验证结论。 5.具备问题拆解和跨团队协作能力;在推理引擎、GPU 性能或性能平台任一方向有深入经验。 6.能够使用 AI 编程工具提升效率,…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
C+++
https://www.learncpp.com/
LearnCpp.com is a free website devoted to teaching you how to program in modern C++.
https://www.youtube.com/watch?v=ZzaPdXTrSb8
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
还有更多 •••
相关职位

社招4年以上
1. 扎实的系统编程能力,可选精通 C++/Python/Rust,熟悉高性能并发编程与内存管理。 2. 深入理解 KVCache 相关技术,PagedAttention / vAttention 等
更新于 2026-06-16北京|杭州

社招4年以上
1. 扎实的后端系统开发能力,熟悉 C++/Go/Python 中至少一种,熟悉高并发服务、分布式系统、RPC、异步编程、缓存系统与资源调度系统设计。 2. 熟悉分布式资源管控或调度系统,理解多租户隔
更新于 2026-06-16北京|杭州
社招4年以上
1. 扎实的系统编程能力,可选精通 C++/Python/Rust,熟悉高性能并发编程与内存管理。 2. 深入理解 KVCache 相关技术,PagedAttention / vAttention 等
更新于 2026-06-16北京|杭州