
MiniMax大模型训练系统专家-实验平台 / 稳定性与容错 / Checkpoint / 训练可观测
社招全职研发地点:北京 | 上海状态:招聘
工作描述
- 我们正在建设支撑大模型研发的下一代 AI 基础设施。 - 这里关注的不只是“把训练任务跑起来”,更需要为训练任务跑得稳(ETTR)、效率高(MFU)负责,让大规模、长周期的模型训练在复杂的软硬件环境中具备更好的可观测性、稳定性、可恢复性与可复现性。 - 你将工作在算法、训练框架与底层算力基础设施的交界处,负责大模型实验平台、训练稳定性、故障诊断与容错恢复、Checkpoint 及训练效率评估等关键系统,直接影响预训练与后训练的研发体验、迭代速度和算力产出。 - 本岗位不以算子开发或并行训练框架研发为核心,重点解决训练全生命周期中的系统基础设施问题。 职位描述 1. 负责大模型实验与训练平台建设,覆盖训练数据管理、实验配置、任务编排、指标观测、产物管理、模型评测、结果追踪与实验复现,持续提升模型研发体验和迭代效率; 2. 建设大规模分布式训练稳定性体系,围绕异常感知、故障诊断、根因定位、故障隔离、自动重试与断点恢复,形成“发现—诊断—恢复—治理”的完整闭环,提升长周期训练任务的成功率; 3. 设计并优化高性能 Checkpoint 与训练状态管理系统,解决大规模并行训练中的…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
分布式系统+
https://www.distributedsystemscourse.com/
The home page of a free online class in distributed systems.
https://www.youtube.com/watch?v=7VbL89mKK3M&list=PLOE1GTZ5ouRPbpTnrZ3Wqjamfwn_Q5Y9A
Kubernetes+
https://kubernetes.io/docs/tutorials/kubernetes-basics/
This tutorial provides a walkthrough of the basics of the Kubernetes cluster orchestration system.
https://kubernetes.io/zh-cn/docs/tutorials/kubernetes-basics/
本教程介绍 Kubernetes 集群编排系统的基础知识。每个模块包含关于 Kubernetes 主要特性和概念的一些背景信息,还包括一个在线教程供你学习。
https://www.youtube.com/watch?v=s_o8dwzRlu4
Hands-On Kubernetes Tutorial | Learn Kubernetes in 1 Hour - Kubernetes Course for Beginners
https://www.youtube.com/watch?v=X48VuDVv0do
Full Kubernetes Tutorial | Kubernetes Course | Hands-on course with a lot of demos
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
还有更多 •••
相关职位
社招5年以上
1. 计算机相关专业本科及以上学历,5 年以上系统软件或基础设施开发经验,3 年以上架构设计经验, 有容器平台、Serverless、安全沙箱领域的架构主导经历优先。 2. 具备全局性的复杂系统架构能
更新于 2026-06-22北京|深圳|杭州
校招程序&技术类
1、硕士及以上学历,计算机、软件工程、人工智能等相关专业在读优先 2、熟练掌握Linux环境下的C/C++与Python语言 3、精通以下至少一项的背景知识或经验:分布式训练框架、高性能计算与通信、G
上海|北京

实习算法
1.本科及以上学历、计算机、软件工程等相关专业优先; 2.有扎实的计算机科学知识,掌握Pytorch,具备良好的编程能力和代码风格。 3. 对AI大模型相关核心技术感兴趣, 对megatron dee
更新于 2026-07-09上海