月之暗面RL infra
社招全职技术类/Technical地点:北京状态:招聘
工作描述
职位概览 在 Moonshot,我们相信强化学习是通往 AGI 的关键路径。我们正在寻找深谙 GPU 性能极限的 RL Infra 工程师,打造支撑下一代大模型自我进化的基础设施。 你将直面 RL 训练的核心矛盾:训练与采样的异构负载、频繁权重复制带来的通信开销、多轮超长Agentic Coding和KimiClaw 场景的rollout效率。通过极致的工程优化,让 RL 训练跑得更快、更稳——每一次参数更新,都让 Kimi 的推理能力更进一步。 核心职责 RL 训练架构:针对大规模 Agentic RL 场景,设计训练与采样的混合调度策略,优化多模型(Policy、Reference、Reward、Value)的并行协同与显存共享 Rollout 效率优化:深度定制…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
Megatron+
https://www.youtube.com/watch?v=hc0u4avAkuM
还有更多 •••
相关职位
社招算法类/AI
职位描述: 主要负责维护和开发Moonshot内部的强化学习后训练框架,支持万亿参数模型reasoning、agentic等方向的文本&多模态RL后训练; 与训练推理引擎方向的团队合作,探索算法、框架
更新于 2026-01-19北京
社招程序&技术类
1. 8+ years of professional experience in software engineering, machine learning, or related technic
上海
社招算法类/AI
Research Scientist / Engineer – Agentic RL Locations: Beijing · Shanghai · Shenzhen · Singapore · Si
更新于 2026-04-03北京