千问千问事业部-多模态大模型Agentic算法专家-北京、广州
社招全职1年以上技术类-算法地点:北京 | 广州状态:招聘
工作描述
任职要求 1. 算法功底: 精通PPO、GRPO、DPO等强化学习算法,有大规模模型RLHF实战经验。 2. 专业深度: 深刻理解CoT、自反思及工具学习,熟悉分布式训练框架。 3. 创新能力: 具备极强的算法创新与工程落地能力,能将前沿综述思想转化为生产力 4. 学术背…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
NeurIPS+
https://neurips.cc/
还有更多 •••
相关职位
社招2年以上
1. 具备VLM/LLM/RL/Reasoning/Agent相关背景知识,熟悉主流VLM大模型算法架构,了解VLM的alignment常见方法,包括但不限于SFT/DPO/PPO/GRPO等; 2.
更新于 2025-11-25杭州
实习A181864
1、2027届硕士及以上学位在读,计算机/人工智能/软件工程相关专业优先; 2、实习时间6个月以上,具备优秀的编程能力,扎实的数据结构和基础算法功底,熟练掌握Python,熟悉深度学习框架(如PyTo
更新于 2026-04-17杭州
校招多模态大模型与应
1. 获得本科及以上学历,计算机、人工智能、自动化、数学、物理等相关专业; 2. 在CVPR、ICCV、NeurIPS、KDD、SIGIR、WWW、RecSys、ICLR、ICML、ACL等顶会有高质
更新于 2026-06-16北京
