月之暗面RL Infra 研究工程师
社招全职算法类/AI地点:北京状态:招聘
工作描述
职位描述: 主要负责维护和开发Moonshot内部的强化学习后训练框架,支持万亿参数模型reasoning、agentic等方向的文本&多模态RL后训练; 与训练推理引擎方向的团队合作,探索算法、框架、硬件的协同设计,提升大规模强化学习训练的稳定性和效率; 职位要求: 有扎实的工程…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
推理引擎+
https://www.youtube.com/watch?v=_dvk75LEJ34
https://www.youtube.com/watch?v=XtT5i0ZeHHE
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
还有更多 •••
相关职位
社招技术类/Tech
职位概览 在 Moonshot,我们相信强化学习是通往 AGI 的关键路径。我们正在寻找深谙 GPU 性能极限的 RL Infra 工程师,打造支撑下一代大模型自我进化的基础设施。 你将直面 RL 训
更新于 2026-05-11北京
社招程序&技术类
1. 8+ years of professional experience in software engineering, machine learning, or related technic
上海
社招算法类/AI
Research Scientist / Engineer – Agentic RL Locations: Beijing · Shanghai · Shenzhen · Singapore · Si
更新于 2026-04-03北京