阿里巴巴大规模分布式训练推理加速及核心基础设施软硬协同优化-阿里星
校招全职阿里控股2026届秋季应届生招聘地点:北京 | 杭州状态:招聘
工作描述
任职要求 1. 分布式系统、计算机体系结构、编译优化或通信与计算协同设计方向的硕/博士研究生。 2. 具备AI训推计算性能分析与优化的经验,能深入分析AI模型在GPU平台上的性能瓶颈,提出并实施优化方案。针对分布式训练和推理系统,进行性能调优,提升系统的吞吐量和效率。 3. 熟悉业界常见的优化栈(cuda/rocm/cutlass/ck/triton等),在高效的内存管理、通信优化(NvLink/Infiniband/RoCEv2等)关键技术上有实操经验。 4. 分布式系统研发经验是加分项:设计和实现高效的分布式训练和推理框架,解决大规模分布式系统中的通信、同步和负载均衡…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
分布式系统+
https://www.distributedsystemscourse.com/
The home page of a free online class in distributed systems.
https://www.youtube.com/watch?v=7VbL89mKK3M&list=PLOE1GTZ5ouRPbpTnrZ3Wqjamfwn_Q5Y9A
性能调优+
https://goperf.dev/
The Go App Optimization Guide is a series of in-depth, technical articles for developers who want to get more performance out of their Go code without relying on guesswork or cargo cult patterns.
https://web.dev/learn/performance
This course is designed for those new to web performance, a vital aspect of the user experience.
https://www.ibm.com/think/insights/application-performance-optimization
Application performance is not just a simple concern for most organizations; it’s a critical factor in their business’s success.
https://www.oreilly.com/library/view/optimizing-java/9781492039259/
Performance tuning is an experimental science, but that doesn’t mean engineers should resort to guesswork and folklore to get the job done.
还有更多 •••
相关职位
实习阿里巴巴2027
1. 分布式系统、计算机体系结构、编译优化或通信与计算协同设计方向的硕/博士研究生。 2. 具备AI训推计算性能分析与优化的经验,能深入分析AI模型在GPU平台上的性能瓶颈,提出并实施优化方案。针对分
更新于 2026-03-17北京|杭州
实习阿里巴巴2027
1、分布式系统、计算机体系结构、编译优化或通信与计算协同设计方向的博士研究生; 2、具备AI训推计算性能分析与优化的经验,能深入分析AI模型在GPU平台上的性能瓶颈,提出并实施优化方案。针对分布式训练
更新于 2026-03-12杭州
实习阿里巴巴2027
1.工程与系统基础:计算机相关专业背景,具备优秀的工程实现能力,精通 C/C++、Go 或 Python;具备扎实的数据结构、操作系统、分布式系统与存储系统基础,熟悉性能分析与故障定位方法。 2.大规
更新于 2026-06-16北京|杭州