
MiniMax大模型工程师-RL框架&RL推理-2027届
校招全职算法地点:北京 | 上海状态:招聘
工作描述
我们希望你在某个领域有真正的深度——推理加速、GPU 性能优化、分布式系统、RL 工程,都行——同时对算法前沿保持真实的好奇心。你不需要什么都懂,但你应该知道自己不懂的东西在哪,并且想去弄懂它。 1. RL Rollout:推理是 RL 循环的发动机,我们使用Forge作为我们的RL框架。Rollout占据整个 RL 循环大部分的 wall-clock 时间,并且约束RL算法和Rollout策略的多样性,能多快、多灵活地生成轨迹,直接决定迭代速度。难点包括:Agentic(多轮、环境交互、外部状态机感知的推理系统优化、工具调用和环境的速度不稳定性)、长尾(深入的调度策略和latency优化)、KV Cache 管理、极致的吞吐优化、RL算法Co-design的Ro…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
分布式系统+
https://www.distributedsystemscourse.com/
The home page of a free online class in distributed systems.
https://www.youtube.com/watch?v=7VbL89mKK3M&list=PLOE1GTZ5ouRPbpTnrZ3Wqjamfwn_Q5Y9A
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
推理引擎+
https://www.youtube.com/watch?v=_dvk75LEJ34
https://www.youtube.com/watch?v=XtT5i0ZeHHE
系统设计+
https://roadmap.sh/system-design
Everything you need to know about designing large scale systems.
https://www.youtube.com/watch?v=F2FmTdLtb_4
This complete system design tutorial covers scalability, reliability, data handling, and high-level architecture with clear explanations, real-world examples, and practical strategies.
还有更多 •••
相关职位

社招
职位描述 RL 正在成为提升大模型能力的重要方向。大规模 RL 的效率和上限,不仅取决于算法设计,也高度依赖底层系统能力,包括 RL 框架、Rollout、推理 serving、调度策略、环境交互和系
更新于 2026-09-04北京|上海

实习算法
RL 正在成为提升大模型能力的重要方向。大规模 RL 的效率和上限,不仅取决于算法设计,也高度依赖底层系统能力,包括 RL 框架、Rollout、推理 serving、调度策略、环境交互和系统稳定性。
更新于 2026-09-08北京|上海

实习算法
RL 正在成为提升大模型能力的重要方向。大规模 RL 的效率和上限,不仅取决于算法设计,也高度依赖底层系统能力,包括 RL 框架、Rollout、推理 serving、调度策略、环境交互和系统稳定性。
更新于 2026-09-07北京|上海