京东大模型推理工程师
社招全职算法开发岗地点:上海状态:招聘
工作描述
任职要求 1. 具备大模型推理架构设计能力,可统筹模型切分、并行策略与通信拓扑,完成性能建模与TCO评估; 2. 拥有高性能算子与编译器优化经验,能开发GPU/国产NPU核心算子并深度优化TVM/MLIR中间表示,突破算子融合与指令调度瓶颈; 3. 具备系统瓶颈洞察力,能从高吞吐低延迟服务中定位计算、通信与内存协同的优化关键点; 4. 具有技术路径突破能力,能针对精度与性能矛盾设计新型优化方案,推动Speculative Decoding与KV Cache压缩等技术落地; 5. 致力于AI Infra技术领先,长期投入大模型高效服…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
系统设计+
https://roadmap.sh/system-design
Everything you need to know about designing large scale systems.
https://www.youtube.com/watch?v=F2FmTdLtb_4
This complete system design tutorial covers scalability, reliability, data handling, and high-level architecture with clear explanations, real-world examples, and practical strategies.
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
还有更多 •••
相关职位
社招1-3年ACG
-本科及以上学历,计算机、软件工程、人工智能等相关专业,1-3年大模型推理工程落地经验,熟悉LLM推理原理 -熟练掌握 Python、Linux、Shell,熟悉网络、多进程/多线程、异步并发编程 -
更新于 2026-06-24北京
社招3-5年J0012
1、本科及以上学历,计算机/软件相关专业; 2、熟练掌握 C++/Python,具备高性能系统研发能力; 3、熟悉 Transformer 推理原理,理解 Attention、KV Cache、采样策
更新于 2026-05-29北京|上海|深圳
社招2年以上元宝技术
1.熟练掌握 C++/Python/Go语言,有2年以上llm大模型推理优化经验; 2.具备基础的GPU编程能力,包括但不限于Cuda、OpenCL;熟悉至少一种GPU加速库,如cublas、cud
更新于 2026-08-08北京