腾讯混元Agent强化学习框架工程师(深圳/北京/上海)
社招全职1年以上混元-模型算法技术地点:北京状态:招聘
工作描述
任职要求 1.具备扎实的 Python 编程能力,熟悉异步编程(Asyncio)、并发处理和工程化最佳实践; 2.熟悉大模型与 Agent 相关应用技术,理解模型调用、工具调用、上下文管理、任务执行、日志 Trace 和结果评估等核心链路; 3.熟悉 Kubernetes 和容器化技术,具备在集群环境下进行开发、部署、排障或性能优化的经验; 4.了解大模型训练流程和基本原理,包括预训练、SFT、RLHF、强化学习训练或自动化评估中的至少一类; 5.具备良好的软件工程能力,重视模块化设计、测试、日志、性能和稳定性治理;…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
asyncio+
https://realpython.com/async-io-python/
Python’s asyncio library enables you to write concurrent code using the async and await keywords.
https://www.youtube.com/watch?v=oAkLSJNr5zY
In this video, we'll be learning all about AsyncIO in Python and how to write asynchronous code using the async/await syntax.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
AI agent+
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
Kubernetes+
https://kubernetes.io/docs/tutorials/kubernetes-basics/
This tutorial provides a walkthrough of the basics of the Kubernetes cluster orchestration system.
https://kubernetes.io/zh-cn/docs/tutorials/kubernetes-basics/
本教程介绍 Kubernetes 集群编排系统的基础知识。每个模块包含关于 Kubernetes 主要特性和概念的一些背景信息,还包括一个在线教程供你学习。
https://www.youtube.com/watch?v=s_o8dwzRlu4
Hands-On Kubernetes Tutorial | Learn Kubernetes in 1 Hour - Kubernetes Course for Beginners
https://www.youtube.com/watch?v=X48VuDVv0do
Full Kubernetes Tutorial | Kubernetes Course | Hands-on course with a lot of demos
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
还有更多 •••
相关职位
社招3年以上混元-模型算法技
1.计算机相关专业本科及以上学历,3年及以上后端 / 平台 / Infra 研发经验; 2.精通至少一门主流后端语言(Python / Go / Java 等),主导过中大型平台或系统的设计与落地,
更新于 2026-07-31北京
社招1年以上混元-模型算法技
1.精通 Python,具备扎实的软件工程能力与系统设计能力,能在复杂系统中推进架构落地; 2.熟练掌握 Docker 容器化交付体系,理解镜像构建优化、依赖隔离、网络/存储、制品/镜像仓库等能力,
更新于 2026-07-31北京