小米agent后训练算法实习生-2027届
实习兼职地点:北京状态:招聘
工作描述
任职要求 基本要求 - 计算机、人工智能、数学、统计等相关专业在读硕士/博士(优秀本科生亦可); - 扎实的机器学习与深度学习基础,深入理解 Transformer 架构与 LLM 训练/推理原理; - 熟悉强化学习基础算法(PPO、DQN、Policy Gradient、Actor-Critic 等),了解 LLM 对齐算法(RLHF、DPO、GRPO、KTO 等)的原理与实现细节; - 熟练使用 Python 与 PyTorch,有实际的大模型训练/微调经验(SFT、LoRA、RLHF 任一); - 了解分布式训练相关技术(数据并行、张量并行、流水线并行、ZeRO 等),具备多卡/多机训练实操经验; - 数学功底扎实,能独立推导优化目标、分析 loss 设计与梯度行为。 工作职责 工作内容 1. Agent 后训练全流程建设: 负责基于大语言模型的通用 Agent 的后训练(Post-training)工作,覆盖 SFT、偏好对齐(DPO/ORPO/KTO 等)、RLHF/RLAIF、RLVR 等阶段,针…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
机器学习+
https://www.youtube.com/watch?v=0oyDqO8PjIg
Learn about machine learning and AI with this comprehensive 11-hour course from @LunarTech_ai.
https://www.youtube.com/watch?v=i_LwzRVP7bg
Learn Machine Learning in a way that is accessible to absolute beginners.
https://www.youtube.com/watch?v=NWONeJKn6kc
Learn the theory and practical application of machine learning concepts in this comprehensive course for beginners.
https://www.youtube.com/watch?v=PcbuKRNtCUc
Learn about all the most important concepts and terms related to machine learning and AI.
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
GRPO+
https://cameronrwolfe.substack.com/p/grpo
Most early work on RL for LLMs used Proximal Policy Optimization (PPO) as the default RL optimizer, but recent reasoning research relies upon Group Relative Policy Optimization (GRPO).
还有更多 •••
相关职位
实习算法类/AI
做什么(说点实际的!): 挖数据:自动化找模型 worse cases,定位问题,提升模型能力 调偏好:用RLHF做preference优化,提升模型对齐效果 搞基建:搭后训练数据 pipeline,
更新于 2025-12-23北京
校招
岗位职责 构建面向量化研究、因子挖掘和策略开发的 Agent 环境、Harness、任务与评测体系; 探索 AutoResearch 范式,使 Agent 能够自主提出研究假设、编写代码、运行实验、完
更新于 2026-08-14北京
社招
岗位职责 构建面向量化研究、因子挖掘和策略开发的 Agent 环境、Harness、任务与评测体系; 探索 AutoResearch 范式,使 Agent 能够自主提出研究假设、编写代码、运行实验、完
更新于 2026-07-28北京