
MiniMaxK8s 架构师
社招全职研发地点:北京 | 上海状态:招聘
工作描述
- 我们正在构建面向下一代大模型的 AI 基础设施。 - 你将深入超大规模 GPU 集群的核心调度链路,解决异构算力管理、复杂工作负载编排、资源效率与系统稳定性等关键问题,为模型训练、推理和持续迭代提供高性能、高弹性、高可用的算力底座。 - 这不是一个以 Kubernetes 运维为主的岗位。我们期待你能够深入调度系统,定义架构、设计机制,并将复杂的基础设施能力沉淀为简单、可靠的工程系统。 职位描述 1. 主导 MiniMax 大模型推理平台与机器学习平台的核心架构设计及演进,构建统一、高效、可扩展的 AI 计算基础设施; 2. 设计超大规模 GPU 集群的资源调度系统,解决异构资源建模、拓扑感知调度、队列与配额管理、优先级与抢占、Gang Scheduling、弹性伸缩等复杂问题; 3. 面向训练、推理、评测及数据处理等多类型工作负载,持续优化资源流转效率、任务吞吐、调度时延和 GPU 有效利用率; 4. 推动在线与离线业务的统一资源管理,在业务稳定性、服务质量与集群利用率之…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
高可用+
https://redis.io/blog/high-availability-architecture/
A high available architecture is when there are a number of different components, modules, or services that work together to maintain optimal performance, irrespective of peak-time loads.
https://www.ibm.com/think/topics/high-availability
High availability (HA) is a term that refers to a system’s ability to be accessible and reliable close to 100% of the time.
Kubernetes+
https://kubernetes.io/docs/tutorials/kubernetes-basics/
This tutorial provides a walkthrough of the basics of the Kubernetes cluster orchestration system.
https://kubernetes.io/zh-cn/docs/tutorials/kubernetes-basics/
本教程介绍 Kubernetes 集群编排系统的基础知识。每个模块包含关于 Kubernetes 主要特性和概念的一些背景信息,还包括一个在线教程供你学习。
https://www.youtube.com/watch?v=s_o8dwzRlu4
Hands-On Kubernetes Tutorial | Learn Kubernetes in 1 Hour - Kubernetes Course for Beginners
https://www.youtube.com/watch?v=X48VuDVv0do
Full Kubernetes Tutorial | Kubernetes Course | Hands-on course with a lot of demos
机器学习+
https://www.youtube.com/watch?v=0oyDqO8PjIg
Learn about machine learning and AI with this comprehensive 11-hour course from @LunarTech_ai.
https://www.youtube.com/watch?v=i_LwzRVP7bg
Learn Machine Learning in a way that is accessible to absolute beginners.
https://www.youtube.com/watch?v=NWONeJKn6kc
Learn the theory and practical application of machine learning concepts in this comprehensive course for beginners.
https://www.youtube.com/watch?v=PcbuKRNtCUc
Learn about all the most important concepts and terms related to machine learning and AI.
系统设计+
https://roadmap.sh/system-design
Everything you need to know about designing large scale systems.
https://www.youtube.com/watch?v=F2FmTdLtb_4
This complete system design tutorial covers scalability, reliability, data handling, and high-level architecture with clear explanations, real-world examples, and practical strategies.
AI agent+
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
分布式系统+
https://www.distributedsystemscourse.com/
The home page of a free online class in distributed systems.
https://www.youtube.com/watch?v=7VbL89mKK3M&list=PLOE1GTZ5ouRPbpTnrZ3Wqjamfwn_Q5Y9A
还有更多 •••
相关职位

社招研发
- 我们正在构建面向下一代大模型的 AI 基础设施。 - 你将深入超大规模 GPU 集群的核心调度链路,解决异构算力管理、复杂工作负载编排、资源效率与系统稳定性等关键问题,为模型训练、推理和持续迭代提
更新于 2026-02-08北京|上海
社招3年以上程序&技术类
1.本科及以上学历,计算机/电子/通信等相关专业,3 年以上 K8s 生产环境工作经验。 2.深入理解 K8s 核心组件(kubelet、kube-scheduler、controller-manag
上海
社招3年以上程序&技术类
1. 本科及以上学历,计算机/电子/通信等相关专业,3 年以上 K8s 生产环境工作经验。 2. 深入理解 K8s 核心组件(kubelet、kube-scheduler、controller-man
上海