
智能互联千问事业部-大模型应用算法专家-北京
社招全职3年以上技术类-算法地点:北京状态:招聘
工作描述
任职要求 1.本科及以上学历,计算机/软件/数学等相关专业,1 年以上算法相关工作经验; 2.深入理解 RLHF 机制,具备 DPO、PPO 等至少一种主流对齐算法的大规模调优与落地经验;有复杂 Reward Model 训练经验者优先; 3.熟悉推理加速与部署,了解大模型推理底层的显存管理与算子优化(如 KV Cache、PagedAttention、FlashAtten…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
还有更多 •••
相关职位
社招3年以上技术类-算法
1.本科及以上学历,计算机/软件/数学等相关专业,1 年以上算法相关工作经验; 2.深入理解 RLHF 机制,具备 DPO、PPO 等至少一种主流对齐算法的大规模调优与落地经验;有复杂 Reward
更新于 2026-06-30北京
社招3年以上技术类-算法
1.计算机或数学相关专业本科及以上学历,1年以上互联网行业工作经验; 2.扎实的C++/Java基础,熟悉python,掌握SQL查询语言,具有优秀的编程能力,熟练使用tensorflow/cuda/
更新于 2026-06-05北京

社招3年以上技术类-算法
1.计算机或数学相关专业本科及以上学历,1年以上互联网行业工作经验; 2.扎实的C++/Java基础,熟悉python,掌握SQL查询语言,具有优秀的编程能力,熟练使用tensorflow/cuda/
更新于 2026-04-03北京