美团【LongCat大模型人才校招】基座大模型推理引擎工程师
校招全职核心本地商业-基础研发平台地点:北京 | 上海状态:招聘
工作描述
任职要求 1.理论基础,深入理解Transformer架构核心机制(Attention/MoE/Memory等),熟悉大模型训练流程及推理流程。 2.工程能力,熟悉主流推理框架(SGLang/vLLM)源码,对PD分离、模型量化 、投机推理、调度重叠、前缀缓存 等关键技术有实战落地经验。精通C++/CUDA/AscendC,具备复杂算子(如FlashAttention、量化GEMM等)的开发与调优经验者优化。掌握RDMA网络编程及分布式系统理论,有MoonCake/LMCache/Dynamo等分布式KV缓存系统实践经验者优先。 3.系统经验,具备大模型推理系统的一线工程经验,熟悉大规模PD分离集群的运维、监控及性能调优者优先。 4.工程素养,代码能力强,具备优秀的性能 profiling、瓶颈…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
还有更多 •••
相关职位
校招核心本地商业-基
1.具备良好的计算机基础素养和分析解决问题的能力,熟练掌握C++或Python。 2.学习能力强,对机器学习系统优化有技术热情,富有极客精神。 3.熟悉PyTorch框架和TVM/MLIR等编译优化技
更新于 2026-06-03北京|上海
校招核心本地商业-基
1.硕士及以上学历,计算机或相关专业,博士优先。 2.在 ML / NLP / RL / CV / Speech 等相关方向有扎实的研究基础,在 ACL / EMNLP / NAACL / Neur
更新于 2026-06-03北京|上海
校招核心本地商业-基
1. 硕士及以上学历,计算机、数学、统计学或相关专业。 2. 熟悉Java/Python/C++等编程语言,良好的编码习惯和一定的工程能力 。 3. 具有深度学习和大模型原理的基础知识,具有多模态大模
更新于 2026-06-03北京|上海