
得物【技术保障】模型推理优化专家
社招全职技术类地点:杭州 | 上海状态:招聘
工作描述
任职要求 1. 具备扎实的 LLM 推理优化实战经验,能独立分析并优化TTFT、TPOT、吞吐等核心指标,有可量化的优化案例优先。 2. 深入理解至少一种主流推理框架(vLLM/SGLang/TensorRT-LLM/Triton等)的原理与调优方法。 3. 熟悉 Kubernetes 核心机制:调度器、资源模型、DevicePlugin机制、节点亲和性、HPA/KEDA等,有基于 Kubernetes 的 GPU 工作负载研发和运维经验。 4. 具备系统级性能分析能力,能使用 nsys、nvtop、perf、eBPF 等…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
还有更多 •••
相关职位

社招8年以上技术类
1.8年以上仓储或数据中心领域建设和管理经验,熟悉仓储或数据中心的规划建设和规模化运维; 2.优秀的分析问题和解决问题的能力,具备良好的跨团队协作和推动能力。 工作职责 1.负责公司仓储和机房IT基
更新于 2026-01-26上海
社招3年以上技术类-运维
1. 具备3年以上技术风险、SRE或技术保障相关工作经验,熟练掌握故障管理、变更防控、监控保障、大促保障体系等; 2. 具备一定的故障定位与较强的跨团队协同能力,能在高压环境下快速响应、判断并推动问题
更新于 2026-01-29上海
实习J1020
1. 本科及以上学历,广播电视工程、电子信息工程、计算机等相关专业优先; 2. 对互联网会议及直播技术有浓厚兴趣,学习能力强,动手能力强,工作细致严谨、责任心强,具备良好的沟通表达能力和抗压能力。
更新于 2026-03-31北京