阿里巴巴日常实习生-RLHF算法工程师-Qwen基础模型
实习兼职阿里巴巴日常实习生地点:北京 | 杭州 | 上海状态:招聘
工作描述
任职要求 1. 计算机、机器学习、人工智能、自然语言处理等相关专业,硕士及以上学历。 2. 具有 RLHF / Reward Modeling / 偏好对齐方向的实践经验,熟悉 PPO、DPO、GRPO 等主流算法的原理与工程实现。 3. 精通 Python 及 PyTorch 等深度学习框架,具备扎实的工程能力,能够独立完成从数据处理、模型训练到评测部署的全链路工作。 4. 对数据质量高度敏感,具备优秀的数据分析与问题诊断能力,能够从复杂反馈信号中提炼有效的优化方向。 加分项 1. 熟悉主流 RL 训练框架(如 veRL )及大规模分布式训练方案(DeepSpeed、Megatron-LM、FSDP)。 2. 在 Reward Model 泛化性、Reward Hacking 缓解、偏好数据合成等方向有深入研究或实际落地经验。 3. 曾在 NeurIPS、ICLR、ICML、ACL 等顶级会议发表对齐、强化学习或 NLP 相关论文,具有一定学术影响力。 4. 拥有知名开源项目贡献经历,在开…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
机器学习+
https://www.youtube.com/watch?v=0oyDqO8PjIg
Learn about machine learning and AI with this comprehensive 11-hour course from @LunarTech_ai.
https://www.youtube.com/watch?v=i_LwzRVP7bg
Learn Machine Learning in a way that is accessible to absolute beginners.
https://www.youtube.com/watch?v=NWONeJKn6kc
Learn the theory and practical application of machine learning concepts in this comprehensive course for beginners.
https://www.youtube.com/watch?v=PcbuKRNtCUc
Learn about all the most important concepts and terms related to machine learning and AI.
NLP+
https://www.youtube.com/watch?v=fNxaJsNG3-s&list=PLQY2H8rRoyvzDbLUZkbudP-MFQZwNmU4S
Welcome to Zero to Hero for Natural Language Processing using TensorFlow!
https://www.youtube.com/watch?v=R-AG4-qZs1A&list=PLeo1K3hjS3uuvuAXhYjV2lMEShq2UYSwX
Natural Language Processing tutorial for beginners series in Python.
https://www.youtube.com/watch?v=rmVRLeJRkl4&list=PLoROMvodv4rMFqRtEuo6SGjY4XbRIVRd4
The foundations of the effective modern methods for deep learning applied to NLP.
学历+
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
还有更多 •••
相关职位
实习阿里巴巴日常实习
1、人工智能、计算机及相关相关专业博士或硕士在读,在视觉生成、计算机视觉、多模态等领域基础扎实 2、代码能力扎实 ,熟练掌握PyTorch开发,有PyTorch分布式训练经验 3、熟悉生成模型(VAE
更新于 2026-08-14北京
实习阿里巴巴日常实习
1、本科及以上学历在读,计算机、人工智能、设计、统计学等相关专业优先,可实习3个月以上,每周至少4天; 2、AI 工具深度使用者——日常高频使用 Claude Code / Cursor / Qod
更新于 2026-08-14上海
实习阿里巴巴日常实习
1、深度参与AIGC产品的需求分析、功能设计与迭代规划,从用户视角出发打磨核心体验 2、调研AIGC工具及创作者生态的竞品动态,输出有深度的市场与产品分析报告 3、参与用户访谈与行为数据分析,将真实反
更新于 2026-07-29北京|杭州