网易AI研发工程师(国产化芯片适配)
社招全职网易智企地点:杭州状态:招聘
工作描述
任职要求 1、计算机科学、人工智能、电子信息等相关专业,本科及以上学历。 2、熟悉tensorrt、triton server、vLLM、TensorRT-LLM、RTP-LLM、SGLang等模型推理部署框架,并在至少一个框架上有深度实践经验。 3、熟悉剪枝、量化、蒸馏、KV Cache、PagedAttention等模型部署与性能优化相关技术。能针对具体场景的数据与业务特点,优化模型部署性能。有大模型线上部署与性能优化相关经验优先。 4、熟悉昇腾、寒武纪、海光等国产AI芯片的模型适配,同时熟悉端侧模…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
Triton Inference Server+
https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html
Triton Inference Server is an open source inference serving software that streamlines AI inferencing.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
还有更多 •••
相关职位
社招3-5年数据仓库
1. 具备优秀的编程功底,熟练掌握Java/Python中的至少一种;熟悉前端技术栈,能独立完成“从数据清洗到后端服务,再到前端展示”的全链路开发。 2. 有使用大语言模型API、Agent框架或其他
更新于 2026-03-24北京|上海|杭州
实习阿里巴巴2027
1. 本科及以上学历,计算机科学、软件工程、人工智能、信息安全、网络安全、通信等相关专业优先; 2. 熟练Java/PHP/C++/Go/Python中的至少一种技术语言,具备良好的软件工程规范和代码
更新于 2026-03-17北京|杭州
社招8年以上研发类
岗位职责 1、负责企业级的AI Agent平台产品分析、架构设计、研发与持续迭代演进; 2、负责Agent技术研究应用及产品能力建设,包括不限于Multi-Agent、Workflow等,构建和完善多
更新于 2026-08-17杭州