蚂蚁金服【蚂蚁星】大模型后训练算法工程师- AI金融场景-27届
校招全职2027届蚂蚁星- Plan A人才计划地点:杭州状态:招聘
工作描述
任职要求 1. 国内外顶尖高校计算机、人工智能、数学、统计等相关专业硕士/博士(应届或毕业两年内)。在AI顶级会议(NeurIPS, ICML, ICLR, ACL, EMNLP, KDD等)以第一作者发表过高水平论文者优先; 2. 具备LLM/NLP相关科研或项目经验,有大模型后训练(SFT/RL)相关研究或落地经验者优先;有小微信贷风控、金融AI相关研究背景者加分; 3. 深入理解LLM预训练及后训练原理,对RLHF、GRPO、OPD等RL算法有深刻理解及实战经验。精通PyTorch深度学习框架,熟悉DeepSpeed、Megatron-LM等分布式训练框架;具备大规模集群训练调试与性能优化经验; 4. 熟悉vLLM、TensorRT-LLM等推理引擎,有模型量化、剪枝或推理加速实际落地经验者优先。具备优秀的数据构建与分析能力,能够从业务视角抽象出数据问题,并设计相应的算法解决方案; 5. 具备极强的问题拆解、实验设计与创新能力,良好的跨团队协作意识,追求技术落地与业务价值。认同网商银行普惠金融价值观,愿意与团队共同探索AI时代小微信贷风控新范式。 工作职责 部门介绍 网商银行是由蚂蚁集团发起设立的中国首批民营银行之一,也是全球领先的互联网银行。我们始终坚守“普惠金融”初心…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
NeurIPS+
https://neurips.cc/
ICML+
https://icml.cc/
ICLR+
https://iclr.cc/
ACL+
https://www.aclweb.org/portal/
Computational linguistics is the scientific study of language from a computational perspective.
EMNLP+
SIGKDD+
https://www.kdd.org/
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
NLP+
https://www.youtube.com/watch?v=fNxaJsNG3-s&list=PLQY2H8rRoyvzDbLUZkbudP-MFQZwNmU4S
Welcome to Zero to Hero for Natural Language Processing using TensorFlow!
https://www.youtube.com/watch?v=R-AG4-qZs1A&list=PLeo1K3hjS3uuvuAXhYjV2lMEShq2UYSwX
Natural Language Processing tutorial for beginners series in Python.
https://www.youtube.com/watch?v=rmVRLeJRkl4&list=PLoROMvodv4rMFqRtEuo6SGjY4XbRIVRd4
The foundations of the effective modern methods for deep learning applied to NLP.
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
还有更多 •••
相关职位
实习蚂蚁星- Pla
1. 国内外顶尖高校计算机、人工智能、数学、统计等相关专业硕士/博士(应届或毕业两年内)。在AI顶级会议(NeurIPS, ICML, ICLR, ACL, EMNLP, KDD等)以第一作者发表过高
更新于 2026-04-17杭州
实习蚂蚁星- Pla
1.学历背景: 计算机科学、人工智能、数学等相关专业博士(或优秀硕士),26/27届毕业生 。编程能力: 2.熟练掌握Python,精通PyTorch深度学习框架。 加分项: 1.有分布式训练经验(
更新于 2026-04-17杭州
实习蚂蚁星- Pla
1. 人工智能、网络安全、计算机科学或相关领域的硕士及以上学位,博士优先; 2. 深厚的机器学习与深度学习理论基础,熟悉Transformer架构及大型语言模型内部机制; 3. 卓越的算法设计与系统架
北京