米哈游GPU 集群 SRE 工程师(稳定性治理)
社招全职5年以上程序&技术类地点:上海状态:招聘
工作描述
任职要求 1、本科及以上学历,计算机相关专业,5 年以上 SRE / 云原生平台 / 大规模集群运维经验,其中 GPU 或异构算力集群相关 2 年以上; 2、具备完整的 SRE 工程实践经验,且主导落地过其中至少两项:SLO 定义与度量、告警治理、oncall 体系、变更管理、故障演练与复盘改进; 3、熟悉 NVIDIA GPU 软件栈(驱动、CUDA、NCCL、DCGM),具备掉卡、显存异常、通信退化与性能问题的定位经验; 4、熟悉 RoCE / InfiniBand 组网,有 PFC/ECN 调优与网络侧故障排查经验; 5、深入理解 Kubernetes 与云原生生态,有生产级集群运维与故障排查经验;熟悉 Volcano / Kueue / Koordinator 等 AI 批调度框架,理解配额与队列管理机制; 6、熟练掌握 Prometheus / Grafana,理解指标口径设计、高基数问题与告警治理方法; 7、熟练掌握 Go 或 Python,…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
CUDA+
https://developer.nvidia.com/blog/even-easier-introduction-cuda/
This post is a super simple introduction to CUDA, the popular parallel computing platform and programming model from NVIDIA.
https://www.youtube.com/watch?v=86FAWCzIe_4
Lean how to program with Nvidia CUDA and leverage GPUs for high-performance computing and deep learning.
NCCL+
https://developer.nvidia.com/nccl
The NVIDIA Collective Communication Library (NCCL) implements multi-GPU and multi-node communication primitives optimized for NVIDIA GPUs and networking.
Kubernetes+
https://kubernetes.io/docs/tutorials/kubernetes-basics/
This tutorial provides a walkthrough of the basics of the Kubernetes cluster orchestration system.
https://kubernetes.io/zh-cn/docs/tutorials/kubernetes-basics/
本教程介绍 Kubernetes 集群编排系统的基础知识。每个模块包含关于 Kubernetes 主要特性和概念的一些背景信息,还包括一个在线教程供你学习。
https://www.youtube.com/watch?v=s_o8dwzRlu4
Hands-On Kubernetes Tutorial | Learn Kubernetes in 1 Hour - Kubernetes Course for Beginners
https://www.youtube.com/watch?v=X48VuDVv0do
Full Kubernetes Tutorial | Kubernetes Course | Hands-on course with a lot of demos
Volcano+
[英文] Tutorials
https://volcano.sh/en/docs/tutorials/
This section provides guidance to help you quickly get started with Volcano, from deploying a basic Volcano Job/Deployment, to integrating with Volcano Queues
Prometheus+
https://grafana.com/docs/grafana/latest/getting-started/get-started-grafana-prometheus/
Prometheus is an open source monitoring system for which Grafana provides out-of-the-box support.
https://prometheus.io/docs/tutorials/getting_started/
Prometheus is a system monitoring and alerting system.
Grafana+
Go+
https://www.youtube.com/watch?v=8uiZC0l4Ajw
学习Golang的完整教程!从开始到结束不到一个小时,包括如何在Go中构建API的完整演示。没有多余的内容,只有你需要知道的知识。
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
还有更多 •••
相关职位
社招5年以上云智能集团
1、精通C/C++/Go等核心开发语言,具备Python、Rust、Shell等一种或多种语言的开发经验,拥有规范的工程化编码能力。 2、深入理解Linux系统,具有Kubernetes及容器化技术的
更新于 2025-12-06杭州
社招5年以上云智能集团
1、精通C/C++/Go等核心开发语言,具备Python、Rust、Shell等一种或多种语言的开发经验,拥有规范的工程化编码能力。 2、深入理解Linux系统,具有Kubernetes及容器化技术的
更新于 2025-12-06杭州
社招5年以上云智能集团
1、精通C/C++/Go等核心开发语言,具备Python、Rust、Shell等一种或多种语言的开发经验,拥有规范的工程化编码能力。 2、深入理解Linux系统,具有Kubernetes及容器化技术的
更新于 2025-12-30杭州