
Soul APPAgentic-RL 算法工程师
社招全职2年以上地点:上海状态:招聘
工作描述
任职要求 1. 全日制硕士及以上学历,2 年及以上大模型算法相关工作经验,具备强化学习或 Agent 项目实践经验。 2. 熟悉大模型后训练技术,理解 SFT、偏好优化及 PPO、GRPO 等强化学习方法,能够独立开展数据构建、奖励设计、训练调优和效果分析。 3. 熟悉 Agent 的规划、工具调用、记忆管理与多轮交互机制,具备任务环境或评测基准的设计经验,能够分析长任务中的错误传播与奖励归因问题。 4. 熟练使用 PyTorch,具备大模型分布式训练经验;能够处理交互轨迹数据…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
还有更多 •••
相关职位

社招2年以上
1. 全日制硕士及以上学历,2 年及以上大模型算法相关工作经验,具备强化学习或 Agent 项目实践经验。 2. 熟悉大模型后训练技术,理解 SFT、偏好优化及 PPO、GRPO 等强化学习方法,能够
更新于 2026-09-23北京
校招2027届蚂蚁星
1. 理论基础扎实: 拥有深厚的深度学习与非凸优化理论功底,具备从第一性原理(FirstPrinciples)对复杂问题进行数学建模和拆解的能力; 2. 架构深度理解: 深入掌握 Transf
更新于 2026-05-14上海|杭州
社招技术类
1、硕士及以上学历优先,计算机、人工智能、数学等相关专业; 2、扎实的机器学习和深度学习基础,熟悉 Transformer、LLM 原理; 3、熟悉至少一种智能体框架(LangChain、AutoGP
更新于 2026-08-16上海