米哈游AI 推理性能优化工程师
社招全职2年以上程序&技术类地点:上海状态:招聘
工作描述
任职要求 1. 有 2 年以上 AI 推理部署、模型性能优化、GPU/NPU 适配或高性能计算相关经验。 2. 熟悉至少 2 种推理引擎或框架,如 TensorRT、vLLM、Triton、SGLang、TensorRT-LLM、ONNX Runtime 等。 3. 熟悉 NVIDIA GPU 生态,包括 CUDA、cuDNN、TensorRT、NCCL,了解 A100、H100、B200 等架构差异。 4. 了解 AMD ROCm 或国产 NPU/GPU 生态,如昇腾、寒武纪、燧原、摩尔线程等,有实际适配经验优先。 5. 熟悉大模型、扩散模型、多模态模型的推理链路,有模型部署、量化、并行推理或性能调优经验。 6. 熟练使用 nsi…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
推理引擎+
https://www.youtube.com/watch?v=_dvk75LEJ34
https://www.youtube.com/watch?v=XtT5i0ZeHHE
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
Triton Inference Server+
https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html
Triton Inference Server is an open source inference serving software that streamlines AI inferencing.
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
ONNX+
https://github.com/onnx/tutorials
Open Neural Network Exchange (ONNX) is an open standard format for representing machine learning models.
[英文] Introduction to ONNX
https://onnx.ai/onnx/intro/
This documentation describes the ONNX concepts (Open Neural Network Exchange).
CUDA+
https://developer.nvidia.com/blog/even-easier-introduction-cuda/
This post is a super simple introduction to CUDA, the popular parallel computing platform and programming model from NVIDIA.
https://www.youtube.com/watch?v=86FAWCzIe_4
Lean how to program with Nvidia CUDA and leverage GPUs for high-performance computing and deep learning.
还有更多 •••
相关职位
校招操作系统及嵌入式
1. 编程能力: 精通C/C++编程,熟悉Python,具备扎实的数据结构与算法基础; 2. AI框架:深入理解至少一种主流AI框架(TensorFlow/PyTorch/MindSpore等)的底层
上海
校招操作系统及嵌入式
1. 编程能力: 精通C/C++编程,熟悉Python,具备扎实的数据结构与算法基础; 2. AI框架:深入理解至少一种主流AI框架(TensorFlow/PyTorch/MindSpore等)的底层
杭州
社招1年以上
● 计算机及相关专业背景,扎实的计算机基础知识,精通Python/Java中至少一门语言。 ● 具有1年以上分布式系统或后端服务系统相关工作经验,具备复杂系统软件的设计和调试能力。 ● 熟悉主流深度学
更新于 2026-07-28北京|杭州