蚂蚁金服蚂蚁集团-大模型后训练专家-CPO线
社招全职3年以上技术类-算法地点:杭州状态:招聘
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
相关职位
社招1-3年J0011
1、在 NLP、LLM、深度学习、强化学习方面有一定研究基础,熟悉主流大模型和算法,并有丰富的实践经验; 2、较强的工程实现能力,熟练掌握 pytorch,熟悉DeepSpeed、Megatron等
更新于 2026-07-03北京
社招3年以上WXG公共技术
1.计算机科学、数学、人工智能等相关专业硕士及以上学历; 2.具备良好的数理基础和 NLP 技术基础,能够熟练使用 Megatron,HuggingFace,DeepSpeed,PyTorch 等框
更新于 2026-06-16北京
社招J0011
1、硕士及以上学历,计算机科学、人工智能、自动化、数学等相关专业优先; 2、精通多模态任务设计范式(如视觉思维链、跨模态推理链),具备CoT提示工程、Reward Model设计经验,掌握合成数据
更新于 2026-06-23北京