阿里巴巴算法研究员-Agentic RL
实习兼职阿里巴巴研究型实习生地点:杭州状态:招聘
工作描述
任职要求 1、研究经历:具备大模型 post-training / 强化学习方向的研究经验,至少一篇顶会一作论文(NeurIPS / ICML / ICLR / ACL 等)。 2、系统理解:熟悉 Agent 系统设计,对 tool use、multi-step reasoning、skills 等能力有深入理解。 3、工程能力:熟练使用 Python / PyTorch,熟悉主流 RL 训练框架(VeRL 等)及 Agent 开发框架。 4、加分项:有长程规划、异步任务调度、multi-agent 系统方向的研究发表;参与过知名 Agentic RL 项目;具备跨境电商或 B2B 贸易场景经验。 …
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
NeurIPS+
https://neurips.cc/
ICML+
https://icml.cc/
ICLR+
https://iclr.cc/
还有更多 •••
相关职位
实习阿里巴巴日常实习
1、核心负责商家侧 Agent 框架研发,设计可编排的 Skill/Tool/Memory 机制,支撑多商家个性化配置与建模。重点探索 Agentic RL(PPO/GRPO 等),开展奖励设计、信用
更新于 2026-07-08杭州
实习阿里巴巴日常实习
1. 学历背景:计算机、人工智能、数学、电子信息等相关专业在读硕士或博士,具备扎实的深度学习、强化学习及概率统计基础。 2. 熟练掌握 Python 及主流深度学习框架(PyTorch/TensorF
更新于 2026-07-07杭州
校招程序&技术类
1.具备大语言模型数据研发相关工作经验; 2.熟悉大模型完整训练链路,包含预训练、中段训练、后训练阶段; 3.具备agent模型训练与评测实操经验; 4.有 agentic, reasoning, c
上海|北京