千问千问事业部-千问/夸克-Post-Training 高级算法专家-北京/杭州
社招全职3年以上技术类-算法地点:北京 | 杭州状态:招聘♡ 收藏
工作描述
任职要求 ● 计算机科学、人工智能、电子工程或相关领域的硕士或博士学位。 ● 在顶级学术会议 (NeurIPS, ICML, ICLR, ACL, EMNLP 等) 发表过相关高质量论文。 ● 在自然语言处理 (NLP) 或大模型 (LLM/VLM) 领域拥有 3 年以上的研发经验,对 Post-training 技术(SFT, RLHF, DPO, PPO、RLVR 等)方向拥有深厚的理论功底和业界公认的成功实践案例。 ● 对深度学习和机器学习有精深的理解,尤其熟悉 Transformer、MoE 等前沿模型架构。在强化学习 (RL) 领域有扎实的理论基础,并主导过其在 LLM 中的创新应用。 ● 具备卓越的工程实现能力,精通 Python 及主流深度学习框架 (PyTorch/TensorFlow)。 ● 具备出色的算法设计与分析能力,能够独立设计、执行和分析复杂的模型实验。 ● 具备优秀的领导力、项目管理能力和团队协作精神,能够带领团队攻克技术难关,并与算…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
NeurIPS+
https://neurips.cc/
ICML+
https://icml.cc/
ICLR+
https://iclr.cc/
ACL+
https://www.aclweb.org/portal/
Computational linguistics is the scientific study of language from a computational perspective.
EMNLP+
NLP+
https://www.youtube.com/watch?v=fNxaJsNG3-s&list=PLQY2H8rRoyvzDbLUZkbudP-MFQZwNmU4S
Welcome to Zero to Hero for Natural Language Processing using TensorFlow!
https://www.youtube.com/watch?v=R-AG4-qZs1A&list=PLeo1K3hjS3uuvuAXhYjV2lMEShq2UYSwX
Natural Language Processing tutorial for beginners series in Python.
https://www.youtube.com/watch?v=rmVRLeJRkl4&list=PLoROMvodv4rMFqRtEuo6SGjY4XbRIVRd4
The foundations of the effective modern methods for deep learning applied to NLP.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
还有更多 •••
相关职位

社招3年以上技术类-算法
● 计算机科学、人工智能、电子工程或相关领域的硕士或博士学位。 ● 在顶级学术会议 (NeurIPS, ICML, ICLR, ACL, EMNLP 等) 发表过相关高质量论文。 ● 在自然语言处理
更新于 2026-04-07北京|杭州

社招3年以上产品类-用户型
基本要求: 1、本科及以上学历,专业不限(计算机、电子、心理学、认知科学、人机交互优先); 2、3-7年AI产品经验(有独立项目落地能力) 3、逻辑清晰、表达能力强,具备良好的跨团队沟通能力与推动力,
更新于 2026-06-18北京|杭州
社招3年以上技术类-算法
1. 硕士及以上学历,数学、强化学习、自然语言处理等相关专业; 2. 在强化学习方面具有丰富的专业知识,熟练掌握深度强化学习算法在大语言模型中的应用及前沿知识; 3. 熟悉大模型相关深度学习框架,如T
更新于 2026-07-21北京|杭州