腾讯微信搜索-Agent算法专家
社招全职1年以上搜一搜技术地点:北京状态:招聘
工作描述
任职要求 1.算法功底:精通PPO、GRPO、DPO等强化学习算法,有大模型RLHF实战经验; 2.专业深度:深刻理解CoT、自反思及工具学习,熟悉分布式训练框架; 3.学术背景:在NeurIPS、ICLR、ICML等顶会以一…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
还有更多 •••
相关职位
社招1年以上搜一搜技术
1.具有优秀的基础算法、代码能力,熟练掌握C/C++或Python编程语言,有ACM/ICPC、等比赛获奖者优先; 2.具有扎实的计算机视觉、机器学习基础,熟悉CV、AIGC、NLP、RL等技术领域
更新于 2026-09-01北京
社招1年以上搜一搜技术
1.岗位要求:; 2.熟悉AI基础硬件设置,有真实的大规模推理系统的设计开发部署经验; 3.熟悉各种主流LLM/VLM的模型结构,具有 vllm/sglang/TRT-llm等推理引擎优化实践经验
更新于 2026-06-11北京
社招1年以上搜一搜技术
1.熟悉AI基础硬件设置,有真实的大规模推理系统的设计开发部署经验; 2.熟悉各种主流LLM/VLM的模型结构,具有 vllm/sglang/TRT-llm等推理引擎优化实践经验; 3.熟悉LLM
更新于 2026-06-11北京