
同花顺基座后训练算法工程师(强化学习方向)
校招全职AI算法类地点:杭州状态:招聘
工作描述
任职要求 【工作内容】 1、负责智能体基座模型后训练算法的设计与落地,覆盖 SFT、DPO、RLHF、RLVR 等训练范式,持续提升模型在金融任务上的能力表现。 2、负责奖励体系设计,包括规则奖励、可验证奖励、Reward Model 与 Judge Model 的训练与校准,解决奖励作弊与目标错配问题。 3、负责 Agentic RL 训练链路建设,包括多轮轨迹采样、工具调用环境接入、长程信用分配与训练/采样效率优化。 4、负责后训练过程的稳定性工程,包括熵坍缩、能力遗忘、KL 失控等问题的诊断与治理,保障多阶段训练的可控与可复现。 5、负责能力归因与迭代闭环,与数据评测、金融研究团队协作,将评测暴露的能力短板转化为下一轮训练目标。 【职位要求】 1、本科及以上学历,计算机、人工智能、数学、统计学等相关专业优先。 2、熟悉大语言模型与智能体基本原理,深入理解 SFT、偏好学习与强化学习的训练流程与失效模式。 3、熟悉 PPO、GRPO 等策略优化算法,熟练使用 verl、OpenRLHF、TRL 等主流后训练框架。 4、具备扎实的 Python 与 PyTor…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
智能体+
https://learn.microsoft.com/en-us/shows/ai-agents-for-beginners/
In this 10-lesson course we take you from concept to code while covering the fundamentals of building AI agents.
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
学历+
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
Ray+
https://github.com/ray-project/ray
Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
https://www.youtube.com/watch?v=FhXfEXUUQp0
In this video, I'll teach you everything you need to know about Apache Ray!
https://www.youtube.com/watch?v=fMiAyj2kgac
Using powerful machine learning algorithms is easy using Ray.io and Python.
https://www.youtube.com/watch?v=q_aTbb7XeL4
Parallel and Distributed computing sounds scary until you try this fantastic Python library.
Kubernetes+
https://kubernetes.io/docs/tutorials/kubernetes-basics/
This tutorial provides a walkthrough of the basics of the Kubernetes cluster orchestration system.
https://kubernetes.io/zh-cn/docs/tutorials/kubernetes-basics/
本教程介绍 Kubernetes 集群编排系统的基础知识。每个模块包含关于 Kubernetes 主要特性和概念的一些背景信息,还包括一个在线教程供你学习。
https://www.youtube.com/watch?v=s_o8dwzRlu4
Hands-On Kubernetes Tutorial | Learn Kubernetes in 1 Hour - Kubernetes Course for Beginners
https://www.youtube.com/watch?v=X48VuDVv0do
Full Kubernetes Tutorial | Kubernetes Course | Hands-on course with a lot of demos
还有更多 •••
相关职位

校招AI算法类
【职位要求】 1、本科及以上学历,计算机、人工智能、数学、统计学等相关专业优先。 2、熟悉大语言模型和智能体基本原理,理解 SFT、强化学习等后训练流程。 3、具备优秀的问题分析和实验研究能力,能够从
杭州

校招AI算法类
【职位要求】 1、本科及以上学历,计算机、人工智能、数学、统计学等相关专业优先。 2、熟悉大语言模型与智能体基本原理,深入理解 SFT、偏好学习与强化学习的训练流程与失效模式。 3、熟悉 PPO、GR
杭州
社招3年以上技术类-算法
学历与专业背景:计算机/机器人/AI 相关专业硕士及以上学历,3 年以上多模态大模型或具身智能核心研发经验。 项目经历(满足其一): •主导过 VLA/VLM/WAM 至少一个方向从 0 到 1 的模
更新于 2026-08-14北京