智谱大模型后训练算法工程师(Coding Agent 方向)
社招全职算法研究地点:北京状态:招聘
工作描述
职位描述 1、负责模型在 Coding Agent 场景的优化; 2、研究大规模数据合成和强化学习方案,提升模型在各类 Coding 框架下的性能; 3、设计和实现评测方案,全方位衡量模型在真实场景中的 Coding Agent 相关能力; 职位要求: 1、本科及以上学历,计算机、软件工程、人工智能等相关专业; 2、熟悉 LLM 相关技术,具有数据构建 / 指令微调 / 强化学习经验,具备优秀的代码能力和基础算法功底,有丰富的工程经验和良好的编程习惯…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
AI agent+
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
学历+
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
还有更多 •••
相关职位

校招ai 算法类
1.计算机、人工智能、机器学习、自然语言处理、强化学习等相关专业,博士优先。 2.熟悉大模型训练和后训练流程,理解 SFT、RLHF、DPO、PPO、GRPO、Reward Model、RLVR 等方
杭州
社招3年以上元宝技术
1.研究生及以上学历,计算机、人工智能、数学等相关专业(有数学、编程竞赛加分); 2.多年NLP/深度学习研发经验,至少1年大模型应用相关实战经验; 3.深入理解LLM技术栈(如SFT、RM、RL
更新于 2026-06-22北京
社招3年以上元宝技术
1.研究生及以上学历,计算机、人工智能、数学等相关专业(有数学、编程竞赛加分); 2.多年NLP/深度学习研发经验,至少1年大模型应用相关实战经验; 3.深入理解LLM技术栈(如SFT、RM、RL
更新于 2026-06-05上海