通义Token Foundry-RLHF 算法专家-Qwen
社招全职3年以上技术类-算法地点:北京 | 杭州状态:招聘
工作描述
任职要求 1. 计算机、机器学习、人工智能、自然语言处理等相关专业,硕士及以上学历。 2. 具有 RLHF / Reward Modeling / 偏好对齐方向的实践经验,熟悉 PPO、DPO、GRPO 等主流算法的原理与工程实现。 3. 精通 Python 及 PyTorch 等深度学习框架,具备扎实的工程能力,能够独立完成从数据处理、模型训练到评测部署的全链路工作。 4. 对数据质量高度敏感,具备优秀的数据分析与问题诊断能力,能够从复杂反馈信号中提炼有效的优化方向。 加分项 1. 熟悉主流 RL 训练框架(如 veRL )及大规模分布式训练方案(DeepSpeed、Megatron-LM、FSDP)。 2. 在 Reward Model 泛化性、Reward Hacking 缓解、偏好数据合成等方向有深入研究或实际落地…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
机器学习+
https://www.youtube.com/watch?v=0oyDqO8PjIg
Learn about machine learning and AI with this comprehensive 11-hour course from @LunarTech_ai.
https://www.youtube.com/watch?v=i_LwzRVP7bg
Learn Machine Learning in a way that is accessible to absolute beginners.
https://www.youtube.com/watch?v=NWONeJKn6kc
Learn the theory and practical application of machine learning concepts in this comprehensive course for beginners.
https://www.youtube.com/watch?v=PcbuKRNtCUc
Learn about all the most important concepts and terms related to machine learning and AI.
NLP+
https://www.youtube.com/watch?v=fNxaJsNG3-s&list=PLQY2H8rRoyvzDbLUZkbudP-MFQZwNmU4S
Welcome to Zero to Hero for Natural Language Processing using TensorFlow!
https://www.youtube.com/watch?v=R-AG4-qZs1A&list=PLeo1K3hjS3uuvuAXhYjV2lMEShq2UYSwX
Natural Language Processing tutorial for beginners series in Python.
https://www.youtube.com/watch?v=rmVRLeJRkl4&list=PLoROMvodv4rMFqRtEuo6SGjY4XbRIVRd4
The foundations of the effective modern methods for deep learning applied to NLP.
学历+
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
还有更多 •••
相关职位

社招3年以上技术类-开发
1. 计算机科学、软件工程或相关专业硕士及以上学历,3年以上后端系统或平台型产品研发经验。 2. 熟悉大模型驱动的Agent架构,具备Task Planning、Function Calling、
更新于 2026-06-15北京|杭州
社招1年以上技术类-开发
1. 硕士及以上学历,机械工程、自动化、计算机等相关专业。 2. 1 年以上机器人相关工作经验,具备真实机器人系统开发与调试经验,能够独立完成机器人平台的部署,有定位并解决系统问题(如控制抖动、通信延
更新于 2026-06-16杭州
社招3年以上技术类-算法
1. 计算机科学、人工智能、机器学习或相关领域的硕士或博士学位。 2. 在扩散模型、自回归模型、多模态生成理解、计算机视觉、NLP、AIGC、计算机图形学、机器学习等一个或多个领域有较深入的研究。 3
更新于 2026-09-14北京|杭州