米哈游AI 推理系统工程师
社招全职2年以上程序&技术类地点:上海状态:招聘
工作描述
任职要求 1.有2年以上推理部署或 AI 性能优化经验。 2.熟悉至少 2 种主流推理引擎(TensorRT / vLLM / Triton / SGlang 等)的原理与调优手段。 3.熟悉 NVIDIA GPU 生态(CUDA、cuDNN、TensorRT、NCCL),了解其架构演进(A100 → H100 → B200 等)。 4.了解 AMD ROCm 或国产 NPU 至少其一的演进路径、算子支持与生态现状。 5.有 开源大模型(LLM / 扩散模型 / 多模态) 部署优化实战经验。 6.扎实的 性能建模能力:能基于 FLOPs、带宽、显存、Batch Size、Sequence Length 等参数进行数学推导与方案设计。 7.熟练使用 Linux、容器化(Docker / K8s)、网络、性能分…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
推理引擎+
https://www.youtube.com/watch?v=_dvk75LEJ34
https://www.youtube.com/watch?v=XtT5i0ZeHHE
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
Triton Inference Server+
https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html
Triton Inference Server is an open source inference serving software that streamlines AI inferencing.
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
CUDA+
https://developer.nvidia.com/blog/even-easier-introduction-cuda/
This post is a super simple introduction to CUDA, the popular parallel computing platform and programming model from NVIDIA.
https://www.youtube.com/watch?v=86FAWCzIe_4
Lean how to program with Nvidia CUDA and leverage GPUs for high-performance computing and deep learning.
NCCL+
https://developer.nvidia.com/nccl
The NVIDIA Collective Communication Library (NCCL) implements multi-GPU and multi-node communication primitives optimized for NVIDIA GPUs and networking.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
还有更多 •••
相关职位
社招3年以上技术类-开发
1. 对 AI 算法和 AI 系统工程(如迭代模式、端到端系统设计、工程框架、性能建模等)有比较深刻的理解,至少熟练掌握一种常见深度学习框架。 2. 理解异构计算和软硬件结合优化,在性能优化方面有一定
更新于 2026-07-28北京|杭州
校招程序&技术类
Qualifications 1.2027届本科及以上,有推理部署或 AI 性能优化经验; 2.熟悉至少 2 种主流推理引擎(TensorRT / vLLM / Triton / SGlang 等)的
上海
社招3年以上云智能集团
1. 计算机相关专业,具有五年以上Golang、Python或C++开发经验 2. 熟悉vLLM、SGLang等开源推理引擎的架构与实现,了解Continuous Batching、PagedAtte
更新于 2026-07-09北京|杭州