
MiniMax大模型训练工程专家-稳定性方向
社招全职5年以上基础架构地点:北京 | 上海状态:招聘
工作描述
一、一句话定位 为大模型「预训练 + 后训练」的基础设施稳定性与效率负总责的资深工程师——横向拉通 GPU 算力、RDMA 网络、通信库、高性能存储、集群调度、训练任务生命周期等多个技术域,确保大规模训练不中断 / 少中断,并持续把有效训练算力(MFU / ETTR)做上去。 二、为什么这个岗位重要 大模型训练是公司最核心、最烧算力的战役。训练稳定性直接决定旗舰模型能否按期炼成。 这是一个端到端 Owner 岗,不是分工到某个组件的螺丝钉——你对"训练能不能稳、跑得快不快"这个结果负责,向上直面算法 / 训练团队的真实诉求。 你将有从 0 到 1 搭建训练稳定性体系的空间:可观测、故障自愈、预案、效率工程,都还有大片可建设的地带。 三、你将负责什么(职责) 1. 训练全栈基础设施保障(横向拉通,核心) GPU 算力:GPU 机器交付与健康管理、坏卡 / 掉卡 / 降频的检测与自愈、节点池管理与快速置换。 RDMA / 高性能网络:IB / RoCE 网络的稳定性与拓扑感知;网络抖动 / 丢包 / 降速的定位与消除。 网络通信库:NCCL(及同类)通信的调优与故障定位——hang、慢 all-reduce、拓扑不亲和等问题的根因排查。 高性能存储:训练数据与 checkpoint 的高吞吐存储链路(并行文件系统 / 对象存…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
NCCL+
https://developer.nvidia.com/nccl
The NVIDIA Collective Communication Library (NCCL) implements multi-GPU and multi-node communication primitives optimized for NVIDIA GPUs and networking.
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
Megatron+
https://www.youtube.com/watch?v=hc0u4avAkuM
还有更多 •••
相关职位
社招2年以上
1.熟悉Megatron/PyTorch等框架的基本的训练流程; 2.掌握GPU/NPU等工作原理、常见操作命令,至少熟练掌握一种编程语言:CUDA或Python中的一种或多种,具备扎实的工程能力和调
更新于 2025-12-15杭州
社招5年以上技术类-算法
● 学历背景:硕士研究生及以上学历,计算机、人工智能或相关专业,1-3 年算法相关工作经验。 ● 编程基础:扎实的Python编程能力,代码风格良好;熟练使用PyTorch等框架,有从数据处理到模型训
更新于 2026-08-29北京|杭州
社招5-10年引擎
任职资格: 精通 PyTorch,具备训练框架源码级阅读与修改能力,有实际性能优化经验。 熟练掌握 Megatron、DeepSpeed、veRL、OpenRLHF、TRL、Llama-Factory
更新于 2026-08-02上海|北京|杭州