腾讯边缘AI推理研发工程师(深圳)
社招全职2年以上腾讯云(TEG)技术地点:北京状态:招聘
工作描述
任职要求 推理方向: 1.有 AI 推理在工业上大规模落地的经验,熟悉 LLM 的模型架构和常用的推理加速方法; 2.熟悉 GPU/TPU 架构,并且能根据硬件架构合理设计上层算力调度、推理加速等相关软件栈,充分发挥硬件算力; 3.熟悉 sglang、vLLM、TensorRT-LLM 等常用推理框架,理解框架的设计和推理加速方法,了解异构芯片的适配过程; 缓存方向: 1.有 KVCache 优化或 AI 推理在工业上大规模落地的经验,熟悉 LLM 的模型架构和常用的推理…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
还有更多 •••
相关职位
社招5年以上技术-芯片
1. 电子工程,计算机等相关专业硕士及以上学历。 2. 具备3年以上AI推理优化相关工作经验,深刻理解并行计算和CUDA编程,熟悉TensorRT和TensorRT-LLM的模型部署和优化。 3. 熟
更新于 2026-07-29上海
社招3年以上云智能集团
1、深入了解Transformer架构,熟悉Pytorch和CUDA,具备主流推理引擎vLLM/SGLang研发和优化能力; 2、精通Go/Java/C++/Python中的至少一门语言,并有生产级系
更新于 2026-01-28杭州

社招3年以上
1、深入了解Transformer架构,熟悉Pytorch和CUDA,具备主流推理引擎vLLM/SGLang研发和优化能力; 2、精通Go/Java/C++/Python中的至少一门语言,并有生产级系
更新于 2026-04-03杭州