百度Summer Camp - Agentic RL /大模型平台策略推理优化实习生(J100476)
实习兼职ACG地点:北京状态:招聘
工作描述
任职要求 -课题名称二:多步长程复杂任务的Agentic RL训练研究 -课题说明 -本课题旨在针对真实业务agent场景(如代码执行、工具调用、多轮对话决策等),构建agentic RL训练闭环,显著提升模型在复杂任务上的成功率与鲁棒性。 -岗位方向:Agentic RL 算法实习生 -希望同学有扎实的强化学习理论基础,熟悉 PPO、GRPO等主流算法,熟悉大语言模型训练流程(SFT → RM → RL),有 RL实践经验,对agentic有系统性理解(工具调用、多步推理、…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
AI agent+
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
还有更多 •••
相关职位
实习ACG
-热爱技术,追求极致,对多模态大模型、强化学习及Multi-agent等有强烈的好奇心与探索欲 -具备扎实的算法实现能力,熟练掌握 Python/C++ 及 PyTorch/PaddlePaddle
更新于 2026-06-05北京

校招产品&运营类
If you’re a digital native living on the front lines of every trend, have a hound-like nose for hot
更新于 2026-05-20北京|深圳

实习职能类
Position Summary We are seeking summer legal interns from students enrolled in a credit bearing lega
更新于 2026-04-14香港