快手【留用实习】训推框架编译优化工程师
实习兼职J1020地点:北京状态:招聘
任职要求
1、硕士及以上学历,专业不限,计算机相关专业优先; 2、了解AI infra 整体技术栈需求,有训练框架或推理框架实战经验、熟悉Tensorflow 或 PyTorch 的使用、有二次开发能力或开源社区贡献经历更佳; 加分项: 1、有大模型相关训练或推理优化经验或GPU 高性能算子开发经验;有vLLM、TensorRT-LLM、MLC-LLM 等框架之一的实践经验; 2、熟悉深度学习编译优化或异构硬件,有 XLA/ TVM /MLIR 开发、优化经验,熟悉pass编写或代码生成原理和实践;或有传统编译器开发经验,熟悉LLVM原理和使用;?4、实习时长3个月及以上, 优先长期实习。
工作职责
1、参与研发业界领先的深度学习编译技术,落地计算优化、显存优化及分布式优化技术到训练框架和推理框架中,赋能深度学习算法落地; 2、XLA 相关编译优化功能开发; 3、结合pytorch/tensorflow等上下游框架适配与集成; 4、异构大模型推理引擎优化,负责调研NV 上各种推理引擎的优化技术,并支持大模型推理各种优化技术在异构硬件上的落地。
包括英文材料
学历+
TensorFlow+
https://www.youtube.com/watch?v=tpCFfeUEGs8
Ready to learn the fundamentals of TensorFlow and deep learning with Python? Well, you’ve come to the right place.
https://www.youtube.com/watch?v=ZUKz4125WNI
This part continues right where part one left off so get that Google Colab window open and get ready to write plenty more TensorFlow code.
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
LLVM+
https://llvm.org/docs/GettingStarted.html
Welcome to the LLVM project!
https://llvm.org/docs/tutorial/
This is the “Kaleidoscope” Language tutorial, showing how to implement a simple language using LLVM components in C++.
https://mcyoung.xyz/2023/08/01/llvm-ir/
“LLVM” is an umbrella name for a number of software components that can be used to build compilers.
https://www.youtube.com/watch?v=Lvc8qx8ukOI
This is the first lecture from the "Programming Language with LLVM" course where we build a full programming language similar to JavaScript from scratch, using LLVM compiler infrastructure.
相关职位
社招J1020
1、参与大模型推理/训练优化。通过研发业界领先的AI Compiler 技术,支撑搜推场景在GPU上的训练计算性能优化;支持大模型推理优化技术在异构硬件上的落地; 2、参与各种大模型推理所需的功能性开发任务;相关编译优化功能开发,以图优化、算子融合、GPU高性能算子开发及自动Codegen等技术手段不断推高在不同卡型上的计算性能极限; 3、参与支持日常的大模型推理服务部署,参与内部日常提效工具的研发。
更新于 2025-05-26
社招gamePlan
1、根据项目方向,协助主策进行游戏系统规划、设计,对现有游戏系统进行优化、调整,增强游戏可玩性; 2、跟进游戏开发进度,配合程序、美术工作,协调完成系统功能的设计制作,进行结果验收与优化迭代; 3、参与游戏系统功能文档攥写,整理系统设计思路,建立系统流程和规范。
更新于 2025-06-12
社招J1001
1、参与大规模的短视频推荐和补贴裂变等业务的模型策略优化,提升用户消费体验(如时长、点击率、互动)和用户留存率、DAU等核心指标; 2、参与机器学习,如深度学习、因果推断领域的技术研发工作,包括但不限于DNN模型优化、因果推断模型、迁移学习、元学习、强化学习和图模型等的算法和系统研发等; 3、参与前沿问题的探索与研究,结合实际应用场景,提供技术优化和落地建议。
更新于 2025-08-05