腾讯腾讯视频-大模型推理加速工程师-(深圳)(杭州)
社招全职5年以上腾讯视频技术地点:北京状态:招聘
工作描述
任职要求 1.熟悉CPU/GPU异构开发,深入理解CUDA编程模型,能独立完成生成类模型的推理加速或性能调优项目; 2.理解生成类模型的核心架构(如扩散模型UNet/Dit结构),熟悉推理过程中的关键性能卡点; 3.熟悉C++ 、Py…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
CUDA+
https://developer.nvidia.com/blog/even-easier-introduction-cuda/
This post is a super simple introduction to CUDA, the popular parallel computing platform and programming model from NVIDIA.
https://www.youtube.com/watch?v=86FAWCzIe_4
Lean how to program with Nvidia CUDA and leverage GPUs for high-performance computing and deep learning.
性能调优+
https://goperf.dev/
The Go App Optimization Guide is a series of in-depth, technical articles for developers who want to get more performance out of their Go code without relying on guesswork or cargo cult patterns.
https://web.dev/learn/performance
This course is designed for those new to web performance, a vital aspect of the user experience.
https://www.ibm.com/think/insights/application-performance-optimization
Application performance is not just a simple concern for most organizations; it’s a critical factor in their business’s success.
https://www.oreilly.com/library/view/optimizing-java/9781492039259/
Performance tuning is an experimental science, but that doesn’t mean engineers should resort to guesswork and folklore to get the job done.
C+++
https://www.learncpp.com/
LearnCpp.com is a free website devoted to teaching you how to program in modern C++.
https://www.youtube.com/watch?v=ZzaPdXTrSb8
还有更多 •••
相关职位
社招2年以上TEG公共技术
1.了解AI基础设施、机器学习系统或高性能计算相关领域经验, 具有 vllm/sglang/TensorRT/FasterTransformer 等推理引擎实践经验; 2.精通主流多模态或全模态大模
更新于 2026-06-08深圳
社招5年以上TEG公共技术
1.熟练掌握 C/C++、Python语言,有计算机体系结构背景或软件开发背景,熟悉系统性能调优的方式; 2.具备基础的GPU编程能力,包括但不限于Cuda、OpenCL;熟悉至少一种GPU加速库,如
更新于 2026-07-15深圳
社招技术类
1、计算机、高性能计算(HPC)、软件工程、人工智能等相关专业硕士及以上学历,具备大模型推理优化或 ML Sys 实际生产环境落地经验 2、具有较强的编程能力,熟练使用Python等编程语言,熟悉C
更新于 2026-07-21上海