米哈游平台研发-AI 推理系统工程师
校招全职程序&技术类地点:上海状态:招聘
工作描述
Qualifications 1.2027届本科及以上,有推理部署或 AI 性能优化经验; 2.熟悉至少 2 种主流推理引擎(TensorRT / vLLM / Triton / SGlang 等)的原理与调优手段; 3.熟悉 NVIDIA GPU 生态(CUDA、cuDNN、TensorRT、NCCL),了解其架构演进(A100 → H100 → B200 等); 4.了解 AMD ROCm 或国产 NPU 至少其一的演进路径、算子支持与生态现状; 5.有开源大模型(LLM / 扩散模型 / 多模态) 部署优化实战经验; 6.扎实的性能建模能力:能基于 FLOPs、带宽、显存、Batch Size、Sequence Length 等参数进行数学推导与方案设计; 7.熟练使用 Linux、容器化(Docker / K…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
推理引擎+
https://www.youtube.com/watch?v=_dvk75LEJ34
https://www.youtube.com/watch?v=XtT5i0ZeHHE
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
Triton Inference Server+
https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html
Triton Inference Server is an open source inference serving software that streamlines AI inferencing.
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
CUDA+
https://developer.nvidia.com/blog/even-easier-introduction-cuda/
This post is a super simple introduction to CUDA, the popular parallel computing platform and programming model from NVIDIA.
https://www.youtube.com/watch?v=86FAWCzIe_4
Lean how to program with Nvidia CUDA and leverage GPUs for high-performance computing and deep learning.
NCCL+
https://developer.nvidia.com/nccl
The NVIDIA Collective Communication Library (NCCL) implements multi-GPU and multi-node communication primitives optimized for NVIDIA GPUs and networking.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
Linux+
https://ryanstutorials.net/linuxtutorial/
Ok, so you want to learn how to use the Bash command line interface (terminal) on Unix/Linux.
https://ubuntu.com/tutorials/command-line-for-beginners
The Linux command line is a text interface to your computer.
https://www.youtube.com/watch?v=6WatcfENsOU
In this Linux crash course, you will learn the fundamental skills and tools you need to become a proficient Linux system administrator.
https://www.youtube.com/watch?v=v392lEyM29A
Never fear the command line again, make it fear you.
https://www.youtube.com/watch?v=ZtqBQ68cfJc
还有更多 •••
相关职位
社招3年以上云智能集团
1. 计算机相关专业,具有五年以上Golang、Python或C++开发经验 2. 熟悉vLLM、SGLang等开源推理引擎的架构与实现,了解Continuous Batching、PagedAtte
更新于 2026-07-09北京|杭州

社招5年以上
1、5 年以上分布式系统架构设计与开发经验,具备复杂分布式系统架构设计及开发经验; 2、对分布式系统架构、数据库、Linux操作系统等有深入理解,具备一定的 Linux 系统应用运维经验; 3、有
更新于 2026-03-30北京|杭州
社招5年以上云智能集团
1、5 年以上分布式系统架构设计与开发经验,具备复杂分布式系统架构设计及开发经验; 2、对分布式系统架构、数据库、Linux操作系统等有深入理解,具备一定的 Linux 系统应用运维经验; 3、有
更新于 2026-02-04北京|杭州