蚂蚁金服蚂蚁集团-垂类模型后训练科学家-北京/杭州/上海【AGI专项】
社招全职3年以上技术类-算法地点:北京 | 上海 | 杭州状态:招聘
工作描述
任职要求 ● 计算机科学、人工智能或相关专业背景,具备大模型后训练实战经验。 ● 精通 SFT、RLHF(PPO/DPO/GRPO)及对齐算法,具备构建复杂奖励模型(Reward Model)的实战经验。 ● 理解Agentic技术栈,有仿真环境构建、工具调用(Tool Use)及多轮决策轨迹合成的相关研发经验。 ● 具备扎实的算法工程实现能力,熟悉 PyTorch、Megatron、vLLM 等主流训练推理框架,能够解决从数据合成到模型落地的全链路工程问题,有handson进行修改的能力 ● 具备良好的定义、分析和解决问题能力,具备敏锐的数据洞察力。 ● 具备较强的团队合作和沟通能力,能够与工程团队、产品团队或其他相关团队紧密配…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
还有更多 •••
相关职位
社招3年以上技术类-算法
1. 计算机、人工智能、金融经济相关专业背景。 2. 熟练掌握机器学习、自然语言处理、大语言模型等相关领域的基本理论和算法。 3. 熟练掌握Python编程语言,熟悉主流深度学习框架(如PyTorch
更新于 2026-07-02北京|上海|杭州
社招3年以上技术-投研
1. 硕士及以上学历; 2. 3年以上头部买方、券商投研经验,医药方向; 3. 扎实的基本面研究功底,清晰的逻辑思维和解决问题能力,快速学习能力,能适应全新领域以及不断演进的变化趋势; 4. 优秀的团
更新于 2026-01-21北京
社招3年以上技术-投研
1. 硕士及以上学历; 2. 3年以上头部买方、券商投研经验,TMT方向; 3. 扎实的基本面研究功底,清晰的逻辑思维和解决问题能力,快速学习能力,能适应全新领域以及不断演进的变化趋势; 4. 优秀的
更新于 2026-01-28北京