月之暗面AI Infra SRE 工程师
社招全职3年以上技术类/Technical地点:深圳 | 北京状态:招聘
工作描述
职位描述 负责月之暗面大规模 GPU 训练与推理集群的稳定性保障,支撑业务 7×24 高可用运行; 负责 Kubernetes 及云原生周边系统(监控、日志、镜像分发、存储)的运维保障与平台化工具开发以及疑难问题的排查解决; 负责 GPU 节点硬件故障的自动化巡检、自愈体系与告警治理; 参与 OnCall 值班,响应集群级突发事…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
高可用+
https://redis.io/blog/high-availability-architecture/
A high available architecture is when there are a number of different components, modules, or services that work together to maintain optimal performance, irrespective of peak-time loads.
https://www.ibm.com/think/topics/high-availability
High availability (HA) is a term that refers to a system’s ability to be accessible and reliable close to 100% of the time.
Kubernetes+
https://kubernetes.io/docs/tutorials/kubernetes-basics/
This tutorial provides a walkthrough of the basics of the Kubernetes cluster orchestration system.
https://kubernetes.io/zh-cn/docs/tutorials/kubernetes-basics/
本教程介绍 Kubernetes 集群编排系统的基础知识。每个模块包含关于 Kubernetes 主要特性和概念的一些背景信息,还包括一个在线教程供你学习。
https://www.youtube.com/watch?v=s_o8dwzRlu4
Hands-On Kubernetes Tutorial | Learn Kubernetes in 1 Hour - Kubernetes Course for Beginners
https://www.youtube.com/watch?v=X48VuDVv0do
Full Kubernetes Tutorial | Kubernetes Course | Hands-on course with a lot of demos
分布式系统+
https://www.distributedsystemscourse.com/
The home page of a free online class in distributed systems.
https://www.youtube.com/watch?v=7VbL89mKK3M&list=PLOE1GTZ5ouRPbpTnrZ3Wqjamfwn_Q5Y9A
Prometheus+
https://grafana.com/docs/grafana/latest/getting-started/get-started-grafana-prometheus/
Prometheus is an open source monitoring system for which Grafana provides out-of-the-box support.
https://prometheus.io/docs/tutorials/getting_started/
Prometheus is a system monitoring and alerting system.
Grafana+
还有更多 •••
相关职位
社招3年以上广告(数据计算平
1.计算机、软件工程、网络工程等相关专业本科及以上学历,具备 3 年及以上大型互联网或 AI Infra 运维 / SRE 经验; 2.熟悉 Linux 操作系统原理,具备扎实的网络、存储、系统调优
更新于 2026-09-01深圳
社招1年以上住宿业务AI &
职位描述参与推荐系统 GPU 推理引擎的研发工作,支撑生成式推荐、排序、召回等业务场景的在线推理服务落地。参与 CUDA 算子开发与优化,包括算子融合、量化(INT8/FP8)、Tensor Core
更新于 2026-08-31上海
社招5年以上运营管理(运营管
1.2年以上相关开发经验,熟悉Go/C++/Python/Java等至少两门语言; 2.有TensorFlow、Pytorch使用或者优化经验,有机器学习平台开发经验优先; 3.了解GPU、CUD
更新于 2026-06-16深圳