月之暗面强化学习-大模型算法工程师/研究员(RL方向 校招)
校招全职算法类/AI地点:北京状态:招聘
工作描述
研究基于Long CoT的大模型强化学习相关技术,包括算法或系统,实现技术突破,涉及: 方向一:推理能力Reasoning 方向二:智能体Agent 同时研究其他通往AGI/ASI的前沿技术 任职要求: 985/211高校研…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
智能体+
https://learn.microsoft.com/en-us/shows/ai-agents-for-beginners/
In this 10-lesson course we take you from concept to code while covering the fundamentals of building AI agents.
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
AI agent+
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
学历+
还有更多 •••
相关职位
社招技术
任职要求 1、计算机、人工智能等相关专业,具备扎实的数据结构与算法基础; 2、具备扎实的 Python 编程能力,熟练掌握 PyTorch 等深度学习框架,有优秀的代码规范与工程素养; 3、熟悉 LL
更新于 2026-01-06
社招5年以上核心本地商业-美
必要条件 学历与专业背景 硕士及以上学历,计算机科学、人工智能、机器学习或相关专业 具有 5 年以上强化学习方向的研究或工程经验 RL 深厚积累 扎实的 RL 理论基础,熟悉分层强化学习(
更新于 2025-11-24北京

社招2年以上算法
1. 计算机科学、数学、统计学或相关领域的硕士或博士; 2. 至少 2 年大语言模型或多模态大模型训练相关研究或工作经验; 3. 具有百亿及以上参数级模型后训练经验,具备使用大规模数据集开展分布式
更新于 2026-07-23广州