
同花顺大模型后训练算法工程师(博士专项)
校招全职ai 算法类地点:杭州状态:招聘
工作描述
任职要求 1.计算机、人工智能、机器学习、自然语言处理、强化学习等相关专业,博士优先。 2.熟悉大模型训练和后训练流程,理解 SFT、RLHF、DPO、PPO、GRPO、Reward Model、RLVR 等方法。 3.具备扎实的机器学习和强化学习基础,熟悉 PyTorch、Transformers、DeepSpeed、Megatron、TRL 等框架。 4.有数据构造、模型评测、实验分析和系统性优化能力。 5.对模型可靠性、幻觉控制、金融合规、领域专业性有较强意识。 加分项 •有大模型后训练、偏好对齐、R…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
机器学习+
https://www.youtube.com/watch?v=0oyDqO8PjIg
Learn about machine learning and AI with this comprehensive 11-hour course from @LunarTech_ai.
https://www.youtube.com/watch?v=i_LwzRVP7bg
Learn Machine Learning in a way that is accessible to absolute beginners.
https://www.youtube.com/watch?v=NWONeJKn6kc
Learn the theory and practical application of machine learning concepts in this comprehensive course for beginners.
https://www.youtube.com/watch?v=PcbuKRNtCUc
Learn about all the most important concepts and terms related to machine learning and AI.
NLP+
https://www.youtube.com/watch?v=fNxaJsNG3-s&list=PLQY2H8rRoyvzDbLUZkbudP-MFQZwNmU4S
Welcome to Zero to Hero for Natural Language Processing using TensorFlow!
https://www.youtube.com/watch?v=R-AG4-qZs1A&list=PLeo1K3hjS3uuvuAXhYjV2lMEShq2UYSwX
Natural Language Processing tutorial for beginners series in Python.
https://www.youtube.com/watch?v=rmVRLeJRkl4&list=PLoROMvodv4rMFqRtEuo6SGjY4XbRIVRd4
The foundations of the effective modern methods for deep learning applied to NLP.
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
还有更多 •••
相关职位
社招算法研究
职位描述 1、负责模型在 Coding Agent 场景的优化; 2、研究大规模数据合成和强化学习方案,提升模型在各类 Coding 框架下的性能; 3、设计和实现评测方案,全方位衡量模型在真实场景中
更新于 2026-05-12北京
社招3年以上元宝技术
1.研究生及以上学历,计算机、人工智能、数学等相关专业(有数学、编程竞赛加分); 2.多年NLP/深度学习研发经验,至少1年大模型应用相关实战经验; 3.深入理解LLM技术栈(如SFT、RM、RL
更新于 2026-06-22北京
社招3年以上元宝技术
1.研究生及以上学历,计算机、人工智能、数学等相关专业(有数学、编程竞赛加分); 2.多年NLP/深度学习研发经验,至少1年大模型应用相关实战经验; 3.深入理解LLM技术栈(如SFT、RM、RL
更新于 2026-06-05上海