月之暗面Kubernetes 调度器开发工程师
社招全职3年以上技术类/Technical地点:深圳 | 北京状态:招聘
工作描述
岗位职责
负责 Kubernetes 调度器及调度插件的深度定制,设计面向 AI 工作负载的调度策略(GPU 拓扑感知、NUMA 亲和、网络亲和、RDMA 域感知);
攻克超大规模集群(万卡级)调度性能瓶颈,优化调度吞吐、调度延迟与决策质量,支持每秒数百 Pod 的调度并发;
构建异构资源调度体系,实现 GPU/CPU/内存/高速互联网络的多维资源建模与在离线混部,提升集群整体利用率;
研发抢占、回填、 gang-schedulin…登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
Kubernetes+
https://kubernetes.io/docs/tutorials/kubernetes-basics/
This tutorial provides a walkthrough of the basics of the Kubernetes cluster orchestration system.
https://kubernetes.io/zh-cn/docs/tutorials/kubernetes-basics/
本教程介绍 Kubernetes 集群编排系统的基础知识。每个模块包含关于 Kubernetes 主要特性和概念的一些背景信息,还包括一个在线教程供你学习。
https://www.youtube.com/watch?v=s_o8dwzRlu4
Hands-On Kubernetes Tutorial | Learn Kubernetes in 1 Hour - Kubernetes Course for Beginners
https://www.youtube.com/watch?v=X48VuDVv0do
Full Kubernetes Tutorial | Kubernetes Course | Hands-on course with a lot of demos
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
Go+
https://www.youtube.com/watch?v=8uiZC0l4Ajw
学习Golang的完整教程!从开始到结束不到一个小时,包括如何在Go中构建API的完整演示。没有多余的内容,只有你需要知道的知识。
分布式系统+
https://www.distributedsystemscourse.com/
The home page of a free online class in distributed systems.
https://www.youtube.com/watch?v=7VbL89mKK3M&list=PLOE1GTZ5ouRPbpTnrZ3Wqjamfwn_Q5Y9A
还有更多 •••
相关职位
社招5年以上LB技术-数据
1. 5年以上Kubernetes调度器核心开发经验 2. 精通Go语言,有大型Go项目架构设计经验 3. 深入理解kube-scheduler完整架构、调度周期和绑定周期 4. 有CPU密集型工作负
更新于 2025-12-23上海
社招3年以上程序&技术类
1. 本科及以上学历,3年以上大规模 K8S 集群运维经验; 2. 熟悉至少一种主流公有云/私有云,包括但不限于Aliyun/腾讯云/AWS/GCP; 3. 深入理解K8S核心架构及主要组件,熟悉 A
上海
社招3年以上技术类/Tech
岗位职责 设计并维护支撑大模型训练/推理的 Kubernetes 平台,负责集群生命周期管理、节点治理、网络/存储 CSI 插件及自动化运维体系; 深度优化容器镜像分发,将千节点集群的 Pod 启动时
更新于 2026-05-27深圳|北京
