阿里巴巴代码场景的Agentic RL与数据合成-阿里星
实习兼职阿里巴巴2027届实习生地点:北京 | 杭州状态:招聘
工作描述
任职要求 1. 计算机、信息科学、数学或相关专业博士学历,深入理解LLM/MLLM/RL/Agent领域,熟悉DPO/PPO/GRPO等RL前沿算法; 2. 有大规模数据合成、Coding Agent训练、大规模RL训练等项目经验者优先; 3. 有较强的自驱力、学习力、创新力、沟通能力和抗压能力,能够跟上正在快速变化的AI时代; 4. 加分项:在ICLR、ICML…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
AI agent+
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
ICLR+
https://iclr.cc/
还有更多 •••
相关职位
校招端点防护
1、本科及以上学历,计算机、软件工程等相关专业优先; 2、熟悉常见等逆向工具比如frida、xposed、Ida Pro; 3、熟悉一门编程语言,可编写漏洞利用POC; 4、有App逆向经验优先,有动
更新于 2026-09-20上海|北京|杭州
实习A196606A
1、2027届硕士及以上学位在读,计算机、软件工程、人工智能等相关专业; 2、熟悉NLP、CV、ML等相关的技术,深入理解大模型相关技术栈(如Reward Model、GRPO/PPO/DPO、SFT
更新于 2026-01-14上海
社招5年以上A159895
1、计算机、软件工程相关专业本科及以上学历,5年以上算法研究与开发经验; 2、具备扎实的算法基础,包括但不限于LLM、CodeLLM、强化学习等领域的认知和实践经验; 3、精通LLM微调、预训练、推理
更新于 2025-11-26杭州