
地平线【地瓜机器人】Agent算法工程师(后训练方向)
社招全职1-3年软件序列地点:北京状态:招聘♡ 收藏
工作描述
任职要求 任职要求 必要条件 - 计算机、人工智能、数学等相关专业,本科及以上学历,**硕士优先** - 1-3 年大模型 / NLP / RL 相关研发经验 - 扎实的 **Python** 编程能力,熟练使用 **PyTorch** 深度学习框架 - 熟悉大模型训练流程:SFT 微调、偏好对齐(DPO/RLHF)、分布式训练(DeepSpeed/Megatron) - 熟悉主流大语言模型(Qwen、DeepSeek、Llama 等)的使用和微调 - 了解 AI Agent 基本原理:ReAct 范式、工具调用、记忆机制、多智能体协作 加分项 - 有大模型后训练的一线实战经验(SFT 数据构建 + 训练 + 评估完整闭环) - 熟悉 RL 训练框架(**verl / openR…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
NLP+
https://www.youtube.com/watch?v=fNxaJsNG3-s&list=PLQY2H8rRoyvzDbLUZkbudP-MFQZwNmU4S
Welcome to Zero to Hero for Natural Language Processing using TensorFlow!
https://www.youtube.com/watch?v=R-AG4-qZs1A&list=PLeo1K3hjS3uuvuAXhYjV2lMEShq2UYSwX
Natural Language Processing tutorial for beginners series in Python.
https://www.youtube.com/watch?v=rmVRLeJRkl4&list=PLoROMvodv4rMFqRtEuo6SGjY4XbRIVRd4
The foundations of the effective modern methods for deep learning applied to NLP.
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
DeepSpeed+
[英文] Getting Started
https://www.deepspeed.ai/getting-started/
DeepSpeed model training is accomplished using the DeepSpeed engine.
https://www.youtube.com/watch?v=pDGI668pNg0
还有更多 •••
相关职位

实习软件序列
1、学历与专业:计算机科学与技术、软件工程、人工智能等相关领域本科及以上学历; 2、编程与框架:精通 Python/Go/TypeScript,具备良好的代码风格;熟悉 Langchain、LLama
更新于 2026-03-18北京

实习软件序列
1、学历与专业:计算机科学与技术、软件工程、人工智能等相关领域本科及以上学历; 2、编程与框架:精通 Python/Go/TypeScript,具备良好的代码风格;熟悉 Langchain、LLama
更新于 2026-03-18北京

实习产品序列
岗位描述 1、学历背景:专业不限,本科及以上学历(大三、研一/研二在读优先); 2、沟通表达: ① 社牛潜质:具备极强的沟通能力和亲和力,不惧怕与陌生人打交道,能搞定跨部门的协作; ② 文案高手:能把
更新于 2026-09-18北京