阿里巴巴AI推理平台-大模型异构芯片适配与推理优化工程师-杭州/北京
社招全职2年以上地点:北京 | 杭州状态:招聘
工作描述
任职要求 1. 负责 Qwen、DeepSeek、GLM、Kimi等主流模型在 TPU、NPU 等各类AI 加速芯片上的适配、验证和性能优化。 2. 分析模型结构与硬件特性,完成 attention、KV cache、MoE、多模态模块、DiT/扩散模型等关键链路的推理打通与优化。 3. 基于 vLLM、SGLang 或自研推理框架,推进模型部署、算子替换、图…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
还有更多 •••
相关职位

社招4年以上
1. 扎实的系统编程能力,可选精通 C++/Python/Rust,熟悉高性能并发编程与内存管理。 2. 深入理解 KVCache 相关技术,PagedAttention / vAttention 等
更新于 2026-06-16北京|杭州

社招4年以上
1. 扎实的后端系统开发能力,熟悉 C++/Go/Python 中至少一种,熟悉高并发服务、分布式系统、RPC、异步编程、缓存系统与资源调度系统设计。 2. 熟悉分布式资源管控或调度系统,理解多租户隔
更新于 2026-06-16北京|杭州
社招4年以上
Qualifications 1. 扎实的系统编程能力,可选精通 C++/Python/Rust,熟悉高性能并发编程与内存管理。 2. 深入理解 KVCache 相关技术,PagedAttention
更新于 2026-09-17北京|杭州