米哈游基础研究-AI 推理优化工程师
校招全职程序&技术类地点:上海状态:招聘
工作描述
任职要求 1.2027届本科及以上,有推理部署或 AI 性能优化经验; 2. 熟悉至少 2 种推理引擎或框架,如 TensorRT、vLLM、Triton、SGLang、TensorRT-LLM、ONNX Runtime 等。 3. 熟悉 NVIDIA GPU 生态,包括 CUDA、cuDNN、TensorRT、NCCL,了解 A100、H100、B200 等架构差异。 4. 了解 AMD ROCm 或国产 NPU/GPU 生态,如昇腾、寒武纪、燧原、摩尔线程等,有实际适配经验优先。 5. 熟悉大模型、扩散模型、多模态模型的推理链路,有模型部署、量化、并行推理或性能调优经验。 6. 熟练使用 nsight、nvprof、…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
推理引擎+
https://www.youtube.com/watch?v=_dvk75LEJ34
https://www.youtube.com/watch?v=XtT5i0ZeHHE
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
Triton Inference Server+
https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html
Triton Inference Server is an open source inference serving software that streamlines AI inferencing.
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
ONNX+
https://github.com/onnx/tutorials
Open Neural Network Exchange (ONNX) is an open standard format for representing machine learning models.
[英文] Introduction to ONNX
https://onnx.ai/onnx/intro/
This documentation describes the ONNX concepts (Open Neural Network Exchange).
CUDA+
https://developer.nvidia.com/blog/even-easier-introduction-cuda/
This post is a super simple introduction to CUDA, the popular parallel computing platform and programming model from NVIDIA.
https://www.youtube.com/watch?v=86FAWCzIe_4
Lean how to program with Nvidia CUDA and leverage GPUs for high-performance computing and deep learning.
还有更多 •••
相关职位
校招程序&技术类
1.熟悉大语言模型和Agent技术的原理、最新进展、能力边界和应用场景 2.具备扎实的编程和工程能力,能够借助 AI 编程工具提升开发效率,完成原型开发、系统实现和持续迭代 3.了解游戏引擎和游戏开发
上海
校招程序&技术类
1)计算机科学、软件工程、数学等相关专业在校学生 2)具备图形学基础知识,对骨骼动画、几何处理、绑定、物理模拟等细分领域之一有深入理解 3)熟练掌握C++编程,具备良好的代码规范和调试能力,有较强的工
上海
校招程序&技术类
1.基于物理的骨骼动画生成,与实际的游戏项目组共同迭代开发在游戏场景下的骨骼动画: a.熟练掌握骨骼动画 / 刚体模拟相关技术(例如,FK&IK、Ragdoll等); b.有基于深度学习的一种或多种
上海