千问智能信息-大模型训练优化专家-强化学习
社招全职1年以上地点:北京 | 杭州 | 广州状态:停招♡ 收藏
工作描述
任职要求 1. 3年及以上大模型训练工程经验,有扎实的深度学习算法基础,精通各类大模型常用训练框架,熟练掌握各种编译、调试、性能分析工具; 2. 熟悉强化学习算法PPO、DPO、GRPO、DAPO等以及相应的高效工程实现,有大模型强化学习工程支持经验和效果优化经验; 3. 精通ray分布式计算框架开发实现,掌握一种或多种分布式训练框架(v…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
还有更多 •••
相关职位

社招5年以上技术类-算法
1、5年以上,计算机、电子信息工程、自动化控制、数学、信息安全等相关专业背景,硕士及以上学历; 2、在机器学习或深度学习领域有实习或者项目经历,具备以下一个或多个方向的研究和应用经验,如多模态数据处理
更新于 2026-03-31北京

社招3年以上技术类-开发
1、3年以上JAVA开发经验,理解io、多线程、集合等基础框架,了解JVM原理,有良好的编程习惯; 2、熟悉分布式系统设计,熟悉分布式、缓存、消息等机制;能对分布式常用技术进行合理应用,解决问题; 3
更新于 2026-04-03北京