阿里巴巴KV Cache分布式存储的网络优化-阿里星
实习兼职阿里巴巴2027届实习生地点:杭州状态:招聘
工作描述
任职要求 1、有GPU计算系统优化经验,熟悉CUDA编程和GPU系统架构; 2、深入理解PyTorch/vLLM推理框架的架构和工作原理,有框架优化经验; 3、掌握CUDA、NCCL等并行编程技术,了解GPU间的通信机制; 4、有大规模分布式系统的设计与实现经验,理解分布式系统的核心问题; 5、良好的工程能力和代码质量意识,能够编写高性能、可维护的代码。 加分项: 1、有顶级开源社区(Linux Kernel、PyTo…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
CUDA+
https://developer.nvidia.com/blog/even-easier-introduction-cuda/
This post is a super simple introduction to CUDA, the popular parallel computing platform and programming model from NVIDIA.
https://www.youtube.com/watch?v=86FAWCzIe_4
Lean how to program with Nvidia CUDA and leverage GPUs for high-performance computing and deep learning.
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
NCCL+
https://developer.nvidia.com/nccl
The NVIDIA Collective Communication Library (NCCL) implements multi-GPU and multi-node communication primitives optimized for NVIDIA GPUs and networking.
分布式系统+
https://www.distributedsystemscourse.com/
The home page of a free online class in distributed systems.
https://www.youtube.com/watch?v=7VbL89mKK3M&list=PLOE1GTZ5ouRPbpTnrZ3Wqjamfwn_Q5Y9A
还有更多 •••
相关职位
校招后端开发
1、计算机相关专业,熟练使用AI coding类工具,具备扎实的工程能力并有成型的作品; 2、有Redis、KV数据库、表格数据库、全闪存高性能分布式存储等产品相关经验者优先; 3、具备良好的分析和解
更新于 2026-08-01北京|上海|杭州
社招2年以上技术类-开发
1. 扎实的系统编程能力,精通 C++/Python,熟悉高性能并发编程与内存管理。 2. 深入理解 KVCache 相关技术,PagedAttention / vAttention 等显存分页管理,
更新于 2026-08-03杭州