快手多模态视频生成数据与算法工程师(RLHF方向)-【可灵AI】
社招全职3-5年J0011地点:北京 | 深圳状态:招聘
工作描述
任职要求 1、硕士及以上学历,具备扎实的工程能力与算法背景,对多模态大模型、视频生成和模型后训练有强烈兴趣; 2、以核心成员身份参与过 RLHF/偏好数据构建或后训练项目,熟悉数据采集、送标、质检、评测、训练、回归的完整链路; 3、在图像或视频生成任务中,具备 DPO / GRPO / ReFL / PPO 等方法的实际工程经验,理解算法效果、稳定性与成本之间的权衡; 4、熟悉计算机视觉基础与生成建模(图像/视频生成、时序建模、VLM 等),具备VLM微调经验; 5、具备良好的跨团队协作与问题拆解能力,能够在复杂业务环境中推动方案落地,对结果负责。 加分…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
还有更多 •••
相关职位