快手大模型推理引擎研发工程师
社招全职J0011地点:北京状态:招聘
任职要求
1、本科及以上学历,电子、自动化、计算机等相关专业优先; 2、了解分布式系统或高性能计算(HPC)相关知识,具备扎实的系统编程、数据结构、算法基础及系统设计能力; 3、熟悉 Linux 开发环境,熟练使用 PyTorch 训练框架,掌握 C++ 与 Python 编程语言; 4、具有良好的团队合作精神与沟通能力,热爱技术钻研,善于分析和解…
登录查看完整任职要求
微信扫码,1秒登录
工作职责
1、参与 LLM 推理框架 的研发与性能优化,支持主流开源模型及自研模型的高效推理; 2、负责底层算子优化,包括但不限于访存模式优化、计算流水线编排、低精度量化(如 INT8/FP8)、编译优化等,持续提升硬件资源利用效率与模型推理 MFU(Model FLOPs Utilization); 3、优化推理框架上层调度策略,通过节点内及节点间的计算任务调度与通信优化,提升整体引擎性能; 4、建设与优化 LLM 模型服务相关工具与平台,提升模型部署的易用性及在线服务的稳定性。
包括英文材料
学历+
分布式系统+
https://www.distributedsystemscourse.com/
The home page of a free online class in distributed systems.
https://www.youtube.com/watch?v=7VbL89mKK3M&list=PLOE1GTZ5ouRPbpTnrZ3Wqjamfwn_Q5Y9A
HPC+
https://www.ibm.com/think/topics/hpc
HPC is a technology that uses clusters of powerful processors that work in parallel to process massive, multidimensional data sets and solve complex problems at extremely high speeds.
数据结构+
https://www.youtube.com/watch?v=8hly31xKli0
In this course you will learn about algorithms and data structures, two of the fundamental topics in computer science.
https://www.youtube.com/watch?v=B31LgI4Y4DQ
Learn about data structures in this comprehensive course. We will be implementing these data structures in C or C++.
https://www.youtube.com/watch?v=CBYHwZcbD-s
Data Structures and Algorithms full course tutorial java
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
系统设计+
https://roadmap.sh/system-design
Everything you need to know about designing large scale systems.
https://www.youtube.com/watch?v=F2FmTdLtb_4
This complete system design tutorial covers scalability, reliability, data handling, and high-level architecture with clear explanations, real-world examples, and practical strategies.
Linux+
https://ryanstutorials.net/linuxtutorial/
Ok, so you want to learn how to use the Bash command line interface (terminal) on Unix/Linux.
https://ubuntu.com/tutorials/command-line-for-beginners
The Linux command line is a text interface to your computer.
https://www.youtube.com/watch?v=6WatcfENsOU
In this Linux crash course, you will learn the fundamental skills and tools you need to become a proficient Linux system administrator.
https://www.youtube.com/watch?v=v392lEyM29A
Never fear the command line again, make it fear you.
https://www.youtube.com/watch?v=ZtqBQ68cfJc
还有更多 •••
相关职位
社招3-5年J0012
参与快手大模型推理引擎研发,工作内容包括: 1、参与大模型推理引擎的设计和研发,支撑快手自研以及开源模型的快速部署和高性能推理; 2、通过各种技术手段持续优化性能,降低推理成本,包括但不限于:算子/编译优化、异构推理、模型量化&蒸馏、分布式并行等; 3、支持RL中的多样化采样、generation性能优化等。
更新于 2026-05-24北京
实习J1014
1、参与快手大规模深度学习推理引擎、大模型训练解决方案的研发与优化,包括大模型推理、模型训练框架、微调平台等; 2、参与底层算子的优化、通过优化访存pattern、计算提升推理性能,与算法部门合作,为公司大模型定制训练方案,探索RLHF、MoE、多模态、longcontext等前沿方向,提升训练性能; 3、优化推理框架上层调度策略,通过机内、机间的计算任务调度和通讯优化提升引擎性能;优化现有大语言模型相关工具和平台,提高模型训练、维护效率,降低成本,提升训练服务稳定性。
更新于 2026-03-13北京
社招2年以上
面向大模型在线推理场景,建设高性能、低成本、高可靠的推理加速基础设施,支撑高并发、大规模模型服务。 你将参与以下工作: 1. 设计和研发 KV Cache 存储、复用、调度与加速系统,优化 Prefill/Decode 阶段的缓存管理与资源利用率。 2. 协同 GPU 显存、主存、SSD 及远端存储等多级资源,优化 KV Cache 换入换出、跨节点迁移、共享复用和生命周期管理。 3. 优化首 Token 延迟、端到端时延、推理吞吐和集群资源成本,提升大模型服务的稳定性与弹性能力。 4. 分析并解决推理服务在线上的性能瓶颈、显存碎片、长尾延迟和稳定性问题,持续推动系统架构演进。
更新于 2026-07-17北京