携程AI Infra 研发工程师(GPU 推理方向)(MJ034962)
社招全职1年以上住宿业务AI & BI地点:上海状态:招聘
工作描述
任职要求 计算机及相关专业本科及以上学历,1 年以上 GPU / 高性能计算 / 深度学习推理相关研发经验。熟悉 CUDA 编程基础,了解 GPU 体系架构(SM、Warp、Memory Hierarchy、Tensor Core),具备编写和优化 CUDA kernel 的实践经验。熟悉至少一种主流深度学习推理框架(TensorRT / ONNX Runtime / TVM / Triton Inference Server),了解图优化、算子融合、量化等基本原理。了解主流推荐模型(DLRM、DIN、生成式推荐等)或 Transformer 类模型的结构与推理特点,对模型性能瓶颈有一定认识。熟练使用 c++/python/java等至少一种语言,熟悉 Linux 开发环境,代码规范、工程意识良好。有一定的问题定位与性能分析能力,能使用 Nsight Sys…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
CUDA+
https://developer.nvidia.com/blog/even-easier-introduction-cuda/
This post is a super simple introduction to CUDA, the popular parallel computing platform and programming model from NVIDIA.
https://www.youtube.com/watch?v=86FAWCzIe_4
Lean how to program with Nvidia CUDA and leverage GPUs for high-performance computing and deep learning.
内核+
https://www.youtube.com/watch?v=C43VxGZ_ugU
I rummage around the Linux kernel source and try to understand what makes computers do what they do.
https://www.youtube.com/watch?v=HNIg3TXfdX8&list=PLrGN1Qi7t67V-9uXzj4VSQCffntfvn42v
Learn how to develop your very own kernel from scratch in this programming series!
https://www.youtube.com/watch?v=JDfo2Lc7iLU
Denshi goes over a simple explanation of what computer kernels are and how they work, alonside what makes the Linux kernel any special.
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
ONNX+
https://github.com/onnx/tutorials
Open Neural Network Exchange (ONNX) is an open standard format for representing machine learning models.
[英文] Introduction to ONNX
https://onnx.ai/onnx/intro/
This documentation describes the ONNX concepts (Open Neural Network Exchange).
Triton Inference Server+
https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html
Triton Inference Server is an open source inference serving software that streamlines AI inferencing.
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
C+++
https://www.learncpp.com/
LearnCpp.com is a free website devoted to teaching you how to program in modern C++.
https://www.youtube.com/watch?v=ZzaPdXTrSb8
还有更多 •••
相关职位
社招软件开发岗
1. 具备AI基础设施整体架构设计能力,能够设计支撑多场景、可扩展的智能化系统并推动技术方案执行; 2. 拥有Agent框架与工具链开发经验,熟悉智能化任务编排与自主决策系统构建,能持续优化系统运行效
更新于 2026-07-29北京
实习阿里巴巴2027
我们期待这样的你 ● 本科及以上学历,计算机、软件工程、人工智能 等计算机相关专业。 ● 热爱编程,熟练掌握C/C++、Python 等编程语言,具备扎实的编程功底。掌握常用数据结构与算法,熟悉 网络
更新于 2026-03-24北京
社招3年以上腾讯云技术
1.熟悉主流的大模型推理框架及其加速技术,如vLLM、SGlang、TensorRT-LLM等,熟练分析单机及分布式情况下的性能热点和优化手段; 2.熟悉业界主流开源模型结构,例如DeepSeek、
更新于 2026-07-01上海