蚂蚁金服蚂蚁集团-Code Agent 后训练算法专家-健康事业群
社招全职5年以上技术类-算法地点:上海 | 杭州状态:招聘
工作描述
任职要求 1. 学历背景 硕士/博士学位,计算机科学、人工智能、软件工程、数学、自动 化或相关专业优先。 2. 编程与工程能力 具备扎实的编程功底与软件工程素养,精通 Python,熟悉 C/C++;对代码生成、程序分析、编译原理、软件测试等领域有系统性理解;能够独立构建和维护复杂的实验pipeline,具备良好的代码规范与工程落地能力。 3. 强化学习理论与实践 深入理解强化学习核心理论(PPO、GRPO、DPO 及其变体),能够从数学原理层面分析策略梯度估计的方差控制、KL散度约束、价值函数基线设计等关键问题;有将 RL 算法应用于语言模型后训练的实际经验,了解 RLHF/RLAIF 的完整技术栈。 4. 大语言模型后训练经验 熟悉大语言模型的 SFT → Reward Modeling → RL 全链路后训练流程,理解 Alignment、Instruction Following、Chain-of-Thought 等核心概念;有独立完成过至少一个完整后训练实验的经验,对训练过程中的常见问题(奖励坍缩、模式坍缩、过度优化等)有实际应对经验。 5. 分布式训练与系统优化 熟练掌握至少一种主流分布式训练框架(DeepSpeed、Megatron-LM、FSDP、veRL/OpenRLHF 等),理解数据并行、模型并行、流水线并行的原理与工程实现;具备在大规模 GPU 集群上进行算法调优与性能排障的实战经验。 6. 研究能力与学术素养 具备优秀的文献调研与前沿追踪能力,能够快速消化最新研究成果并转化为可落地的技术方案;在 NeurIPS、ICML、ICLR、ACL、EMNLP、ICSE、FSE 等顶级会议或期刊发表过相关高质量论文者优先。 …
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
C+
https://www.freecodecamp.org/chinese/news/the-c-beginners-handbook/
本手册遵循二八定律。你将在 20% 的时间内学习 80% 的 C 编程语言。
https://www.youtube.com/watch?v=87SH2Cn0s9A
https://www.youtube.com/watch?v=KJgsSFOSQv0
This course will give you a full introduction into all of the core concepts in the C programming language.
https://www.youtube.com/watch?v=PaPN51Mm5qQ
In this complete C programming course, Dr. Charles Severance (aka Dr. Chuck) will help you understand computer architecture and low-level programming with the help of the classic C Programming language book written by Brian Kernighan and Dennis Ritchie.
C+++
https://www.learncpp.com/
LearnCpp.com is a free website devoted to teaching you how to program in modern C++.
https://www.youtube.com/watch?v=ZzaPdXTrSb8
编译原理+
https://compilerbook.com/
We're picking up right where we left off and write a compiler and a virtual machine for Monkey.
https://craftinginterpreters.com/
Crafting Interpreters contains everything you need to implement a full-featured, efficient scripting language.
https://interpreterbook.com/
In this book we will create a programming language together.
https://nostarch.com/writing-c-compiler
Build a Real Programming Language from Scratch
https://www.youtube.com/watch?v=5ZmFlxrNaN8&list=PLBlnK6fEyqRjT3oJxFXRgjPNzeS-LFY-q
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
FSDP+
https://docs.pytorch.org/tutorials/intermediate/FSDP_tutorial.html
In DistributedDataParallel (DDP) training, each rank owns a model replica and processes a batch of data, finally it uses all-reduce to sync gradients across ranks.
https://www.youtube.com/watch?v=PjEwLgyzuzQ
FSDP provides a comprehensive framework for large model training in PyTorch.
还有更多 •••
相关职位
社招2年以上微信支付技术
1.计算机、人工智能等相关专业硕士以上学历; 2.有大规模强化学习、大模型Code/Agent研发相关经验者优先; 3.具有扎实的深度学习算法基础,熟悉深度学习框架和分布式训练推理加速,有实操经验
更新于 2026-09-08上海
校招核心本地商业-基
【岗位要求】 1、在Agentic RL、PRM或复杂代码推理等方向有深入研究 2、顶会论文发表(ACL/EMNLP/NeurIPS/ICLR/KDD等)者优先 3、GitHub高Star AI原生项
更新于 2026-06-03北京|上海
实习核心本地商业-基
1、有好奇心,敢想敢做,学习能力强,能在复杂问题的深度思考与拆解能力; 2、在 Agentic RL、过程奖励(PRM)或复杂代码推理等方向有深入研究及顶会论文发表(ACL/EMNLP/NeurIPS
更新于 2026-04-03北京|上海