腾讯微信搜索-Agent算法专家
社招全职1年以上搜一搜技术地点:北京状态:招聘
任职要求
1.算法功底:精通PPO、GRPO、DPO等强化学习算法,有大模型RLHF实战经验; 2.专业深度:深刻理解CoT、自反思及工…
登录查看完整任职要求
微信扫码,1秒登录
工作职责
1.参与微信搜索Agent能力优化,包括Search Agent(DeepSearch/DeepResearch)和真实世界复杂任务上的Agentic能力; 2.跟进前沿技术:Mid-Train、SFT、GRM、PRM、RLVR、Agentic RL、Agent自进化、Context管理/Memory等; 3.探索稳定高效的Agentic RL方案,探索下一代大模型结合Agent结合搜索的技术和产品范式。
包括英文材料
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
还有更多 •••
相关职位
社招3年以上搜一搜产品
1.深入分析用户需求,探索大模型回答开放域复杂问题的更优解,细化到对模型效果的理想态定义; 2.搭建并完善开放域问题的评估体系,保证Agent效果可控、稳定; 3.深入垂直场景,制定并推动Agent产品架构及效果优化方向,如模型改进、检索增强等; 4.设计和迭代Agent执行过程,包括需求理解、规划任务拆解、搜索调用等; 5.跨团队协作,推动评估与优化方案落地。
更新于 2025-09-23北京
社招1年以上搜一搜技术
1.负责多模态大模型的研发和应用,研究相关技术在应用领域的全新解决方案,包括而不限于多模态理解生成,视觉Agent等能力; 2.数据建设、指令微调、偏好对齐、模型优化; 3.相关应用落地,包括搜索、工具应用、对话等。
更新于 2026-06-11北京
社招1年以上搜一搜技术
1.工作职责:; 2.负责开发和优化LLM,VLM等大模型的推理引擎,构建适合AI Search,智能 Agent相关领域大规落地应用中的推理基础架构; 3.紧跟 LLM Infra 领域的前沿技术演进突破,将合适成果落地于实际应用; 4.与搜索算法同学深度合作,联合优化,设计实现能够给大型搜索系统带来代际更迭的大模型。
更新于 2026-06-11北京