腾讯元宝-大模型后训练算法工程师
社招全职3年以上元宝技术地点:北京状态:招聘
工作描述
任职要求 1.研究生及以上学历,计算机、人工智能、数学等相关专业(有数学、编程竞赛加分); 2.多年NLP/深度学习研发经验,至少1年大模型应用相关实战经验; 3.深入理解LLM技术栈(如SFT、RM、RLHF、数据合成等); 4.熟悉…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
NLP+
https://www.youtube.com/watch?v=fNxaJsNG3-s&list=PLQY2H8rRoyvzDbLUZkbudP-MFQZwNmU4S
Welcome to Zero to Hero for Natural Language Processing using TensorFlow!
https://www.youtube.com/watch?v=R-AG4-qZs1A&list=PLeo1K3hjS3uuvuAXhYjV2lMEShq2UYSwX
Natural Language Processing tutorial for beginners series in Python.
https://www.youtube.com/watch?v=rmVRLeJRkl4&list=PLoROMvodv4rMFqRtEuo6SGjY4XbRIVRd4
The foundations of the effective modern methods for deep learning applied to NLP.
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
还有更多 •••
相关职位
社招3年以上元宝技术
1.研究生及以上学历,计算机、人工智能、数学等相关专业(有数学、编程竞赛加分); 2.多年NLP/深度学习研发经验,至少1年大模型应用相关实战经验; 3.深入理解LLM技术栈(如SFT、RM、RL
更新于 2026-06-05上海
社招1年以上元宝技术
1.Prompt 工程能力:熟练掌握 System/User Prompt 设计,了解 Few-shot、Chain-of-Thought 等技术,能基于验证集快速迭代优化; 2.大模型应用经验:有
更新于 2026-06-12深圳
社招1年以上元宝产品
1.背景要求: 具备 1 年以上策略产品或 AI 相关产品经验; 2.策略功底: 对生成模型底层原理有深入理解,熟悉常见的 Prompt Engineering、模型评估方法及数据标注逻辑; 3.
更新于 2026-07-06北京