月之暗面大模型后训练算法实习生(RL & Agent方向)
实习兼职算法类/AI地点:北京状态:招聘
工作描述
做什么(说点实际的!):
挖数据:自动化找模型 worse cases,定位问题,提升模型能力
调偏好:用RLHF做preference优化,提升模型对齐效果
搞基建:搭后训练数据 pipelin…登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
相关职位

实习算法序列
-计算机、人工智能、模式识别等相关专业在读硕士及以上学历 -扎实的编程基础,熟练掌握Python或C++ -熟悉至少一种深度学习框架,例如PyTorch,了解Transformers工具库 -熟悉Li
更新于 2026-03-16北京|上海|香港

校招算法序列
-计算机、人工智能、模式识别等相关专业在读硕士及以上学历 -扎实的编程基础,熟练掌握Python或C++ -熟悉至少一种深度学习框架,例如PyTorch,了解Transformers工具库 -熟悉Li
更新于 2026-03-16北京|上海|香港
实习日常实习
1. 博士/硕士生在读,计算机、自动化、数学或相关人工智能专业; 2. 具备大模型研发经验,在 Post-Training(如SFT、RL、Model Merge、OPD等)某一方向上有研究成果或实战
更新于 2026-06-24北京|上海|杭州