腾讯大模型Code/Agent后训练算法研究员-(深圳)or(北京)or
社招全职2年以上微信支付技术地点:上海状态:招聘
工作描述
任职要求 1.计算机、人工智能等相关专业硕士以上学历; 2.有大规模强化学习、大模型Code/Agent研发相关经验者优先; 3.具有扎实的深度学习算法基础,熟悉深度学习框架和分布式训练推理加速,有实操经验者优先; 4.在多模态/CV/NLP等领域顶级会议(期刊)发表过论文、主导/参与…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
AI agent+
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
还有更多 •••
相关职位
实习核心本地商业-基
1、有好奇心,敢想敢做,学习能力强,能在复杂问题的深度思考与拆解能力; 2、在 Agentic RL、过程奖励(PRM)或复杂代码推理等方向有深入研究及顶会论文发表(ACL/EMNLP/NeurIPS
更新于 2026-04-03北京|上海
社招A182748
1、优秀的代码能力、数据结构和基础算法功底,熟练使用PyTorch、TensorFlow、JAX等任一深度学习框架; 2、熟悉大模型或RL算法和技术,在相关领域有过良好研究记录者优先; 3、在大模型领
更新于 2026-04-21北京
校招程序&技术类
1、计算机科学、人工智能、机器学习、软件工程或相关领域硕士 / 博士学历。 2、熟悉 Transformer 架构和大模型训练流程,对 SFT、RLHF、DPO、PPO、GRPO 等 Post-tra
北京