蚂蚁金服【Plan A】大模型Agentic RL算法工程师-百灵-27届
校招全职2027届蚂蚁星- Plan A人才计划地点:上海 | 杭州状态:招聘
工作描述
任职要求 1. 理论基础扎实: 拥有深厚的深度学习与非凸优化理论功底,具备从第一性原理(FirstPrinciples)对复杂问题进行数学建模和拆解的能力; 2. 架构深度理解: 深入掌握 Transformer 及其变体架构,熟悉 MoE(混合专家模型)机制及主流 LLM 的底层实现细节; 3. 后训练范式: 精通 SFT(监督微调)、RLHF(人类反馈强化学习)、DPO/PPO 等对齐算法,深刻理解Alignment(对齐)背后的数学原理与优化挑战; 4. Agentic 实战经验: 具备复杂 Agentic 系统(如 Tool-use, Reasoning,Planning)的训练与全链路搭建经验,理解 Agent 在复杂环境下的交互逻辑; 5. 科学实验素养: 具备独立设计实验、进行严谨消融实验(Ablation Study)及归因分析的能力,能在复杂工程约束下做出最优算法决策。 加分项 1. 在顶会 /…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
还有更多 •••
相关职位
实习蚂蚁星- Pla
1. 计算机科学、人工智能、数学等相关专业硕士及以上学历; 2. 深入理解 Transformer 架构及 SFT / RLHF / DPO / PPO / GRPO 等核心算法;熟悉 AI Agen
更新于 2026-05-12北京|上海|杭州
校招2027届蚂蚁星
1. 计算机科学、人工智能、数学等相关专业硕士及以上学历; 2. 深入理解 Transformer 架构及 SFT / RLHF / DPO / PPO / GRPO 等核心算法;熟悉 AI Agen
更新于 2026-05-14北京|上海|杭州
校招2027届蚂蚁星
1. 热爱人工智能领域,对探索新事物充满热情; 2. 硕士及以上学历,计算机科学、人工智能或相关专业背景; 3. 熟练掌握机器学习、自然语言处理、大语言模型等相关领域的基本理论和算法,具
更新于 2026-05-14北京|杭州|成都