阿里巴巴阿里国际站-Agentic RL环境与模型训练工程师-杭州/北京
社招全职2年以上技术类-算法地点:北京 | 杭州状态:招聘♡ 收藏
工作描述
任职要求
我们希望你具备:
1. 3年以上机器学习 / 强化学习 / 大模型训练 / 算法工程经验
2. 熟练 Python,有训练 pipeline 或大规模系统经验
3. 了解 RL、trajectory、reward、policy optimization、post-training 等基本方法
4. 对 Agent、Tool Use、Browser/Computer Use、多步任务有兴趣或经验
5. 有 rollout system / simulator / RLHF / agent training 经验优先
工作职责
我们正在围绕 阿里巴巴国际站 + Accio Work 打造…登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
机器学习+
https://www.youtube.com/watch?v=0oyDqO8PjIg
Learn about machine learning and AI with this comprehensive 11-hour course from @LunarTech_ai.
https://www.youtube.com/watch?v=i_LwzRVP7bg
Learn Machine Learning in a way that is accessible to absolute beginners.
https://www.youtube.com/watch?v=NWONeJKn6kc
Learn the theory and practical application of machine learning concepts in this comprehensive course for beginners.
https://www.youtube.com/watch?v=PcbuKRNtCUc
Learn about all the most important concepts and terms related to machine learning and AI.
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
还有更多 •••
相关职位
校招A161076A
1、2027届毕业,获得博士学位,人工智能、计算机科学、数学相关专业优先; 2、技术能力: 1)具备出色的编程能力,精通Python/C++,熟悉TensorFlow/PyTorch等框架,有自然
更新于 2026-04-14深圳

社招2年以上
1. 全日制硕士及以上学历,2 年及以上大模型算法相关工作经验,具备强化学习或 Agent 项目实践经验。 2. 熟悉大模型后训练技术,理解 SFT、偏好优化及 PPO、GRPO 等强化学习方法,能够
更新于 2026-09-23北京

社招2年以上
1. 全日制硕士及以上学历,2 年及以上大模型算法相关工作经验,具备强化学习或 Agent 项目实践经验。 2. 熟悉大模型后训练技术,理解 SFT、偏好优化及 PPO、GRPO 等强化学习方法,能够
更新于 2026-09-20上海