蚂蚁金服蚂蚁集团-算法工程师-阿福模型安全
社招全职1年以上技术类-算法地点:杭州状态:招聘
工作描述
任职要求 1、计算机/数学/统计/AI 相关专业本科及以上,硕士/博士优先; 2、1年以上大模型、多模态、自然语言处理等相关工作经验; 3、深刻理解 LLM 及后训练:SFT/DPO/GRPO等;理解 Agent/RAG 的典型失败模式与评测要点; 4、扎实的算法与工程能力:熟练 Python,能独立完成算法开发、数据处…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
NLP+
https://www.youtube.com/watch?v=fNxaJsNG3-s&list=PLQY2H8rRoyvzDbLUZkbudP-MFQZwNmU4S
Welcome to Zero to Hero for Natural Language Processing using TensorFlow!
https://www.youtube.com/watch?v=R-AG4-qZs1A&list=PLeo1K3hjS3uuvuAXhYjV2lMEShq2UYSwX
Natural Language Processing tutorial for beginners series in Python.
https://www.youtube.com/watch?v=rmVRLeJRkl4&list=PLoROMvodv4rMFqRtEuo6SGjY4XbRIVRd4
The foundations of the effective modern methods for deep learning applied to NLP.
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
GRPO+
https://cameronrwolfe.substack.com/p/grpo
Most early work on RL for LLMs used Proximal Policy Optimization (PPO) as the default RL optimizer, but recent reasoning research relies upon Group Relative Policy Optimization (GRPO).
还有更多 •••
相关职位
社招1年以上技术类-算法
1、计算机/数学/统计/AI 相关专业本科及以上,硕士/博士优先; 2、1年以上大模型、多模态、自然语言处理等相关工作经验; 3、深刻理解 LLM 及后训练:SFT/DPO/GRPO等;理解 Agen
更新于 2026-06-16杭州
社招1年以上技术类-算法
1. 硕士及以上学历,计算机科学或相关专业背景。 2. 具备大模型研发经验,在 Post-Training(如SFT、RL、Model Merge、OPD等)某一方向上有深入积累。 3. 算法与工程兼
更新于 2026-07-15北京|上海|杭州
社招算法开发岗
1.硕士及以上学历,计算机相关专业; 2.具有优秀的编程基础,熟练使用Python/C++等至少一种编程语言; 3.熟悉NLP、CV、ML等相关的技术,深入理解大模型相关技术栈(如Reward Mod
更新于 2026-06-17北京
