
地平线【地瓜机器人】后端运维工程师(AI基础设施-GPU基建)
社招全职5年以上软件序列地点:北京状态:招聘♡ 收藏
工作描述
任职要求
1. 核心技能(必备)
- 5年以上工作经验,具备 3 年以上大算力 AI 集群相关经验,有 数千 P 级集群自建或核心运维经历(重点加分)。
- 精通 NVIDIA B 系列、H 系列等高端算力卡的部署、调优与故障排查,熟悉 GPU 并行计算原理与加速技术(NCCL/TensorRT 等)。
- 具备 国产算力卡适配实战经验,至少熟练掌握 1 种国产卡(阿里PPU、华为昇腾 等)的驱动安装、框架适配、模型迁移流程。
- 熟悉 AI 集群核心组件:分布式存储(Juicefs / Alluxio / Ceph / GlusterFS 其一)、高速网络(InfiniBand)、调度系统(Kubernetes/Volcano/Slurm)、监控工具(Prometheus/Grafana)。
- 具备 Shell/Python 脚本开发能力,能独立开发运维自动化工具(如集群巡检、故障告警、资源调度脚本)。
2. 经验与背景(优先)
- 有智算中心、…登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
NCCL+
https://developer.nvidia.com/nccl
The NVIDIA Collective Communication Library (NCCL) implements multi-GPU and multi-node communication primitives optimized for NVIDIA GPUs and networking.
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
Ceph+
https://docs.ceph.com/en/squid/start/beginners-guide/
The purpose of A Beginner’s Guide to Ceph is to make Ceph comprehensible.
https://www.youtube.com/watch?v=oEKJnHAfSiw
Kubernetes+
https://kubernetes.io/docs/tutorials/kubernetes-basics/
This tutorial provides a walkthrough of the basics of the Kubernetes cluster orchestration system.
https://kubernetes.io/zh-cn/docs/tutorials/kubernetes-basics/
本教程介绍 Kubernetes 集群编排系统的基础知识。每个模块包含关于 Kubernetes 主要特性和概念的一些背景信息,还包括一个在线教程供你学习。
https://www.youtube.com/watch?v=s_o8dwzRlu4
Hands-On Kubernetes Tutorial | Learn Kubernetes in 1 Hour - Kubernetes Course for Beginners
https://www.youtube.com/watch?v=X48VuDVv0do
Full Kubernetes Tutorial | Kubernetes Course | Hands-on course with a lot of demos
Volcano+
[英文] Tutorials
https://volcano.sh/en/docs/tutorials/
This section provides guidance to help you quickly get started with Volcano, from deploying a basic Volcano Job/Deployment, to integrating with Volcano Queues
Slurm+
https://researchcomputing.princeton.edu/education/external-online-resources/slurm
All the Research Computing clusters at Princeton rely on a workload manager called SLURM to allocate resources to jobs of different users.
https://www.youtube.com/watch?v=NH_Fb7X6Db0&list=PLZfwi0jHMBxB-Bd0u1lTT5r0C3RHUPLj-
Introduction to the Slurm Resource Manager for users and system administrators.
Prometheus+
https://grafana.com/docs/grafana/latest/getting-started/get-started-grafana-prometheus/
Prometheus is an open source monitoring system for which Grafana provides out-of-the-box support.
https://prometheus.io/docs/tutorials/getting_started/
Prometheus is a system monitoring and alerting system.
Grafana+
Bash+
[英文] The Bash Guide
https://guide.bash.academy/
A quality-driven guide through the shell's many features.
https://www.youtube.com/watch?v=tK9Oc6AEnR4
Understanding how to use bash scripting will enhance your productivity by automating tasks, streamlining processes, and making your workflow more efficient.
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
还有更多 •••
相关职位

社招5年以上软件序列
1. 学历与专业:计算机科学与技术、软件工程、云计算、大数据、网络工程等相关专业本科及以上学历。 2. 工作经验:5 年及以上分布式存储、大数据存储相关开发 / 运维经验,熟悉主流分布式存储架构与业务
更新于 2026-07-07北京

社招1-3年软件序列
任职要求 必要条件 - 计算机、人工智能、数学等相关专业,本科及以上学历,**硕士优先** - 1-3 年大模型 / NLP / RL 相关研发经验 - 扎实的 **Python** 编程能力,熟练使
更新于 2026-04-10北京

实习业务拓展序列
1、本科/硕士在读,计算机/AI/机器人/自动化等理工科背景优先; 2、对具身智能、人形机器人、大模型等方向有强兴趣,愿意持续学习前沿; 3、能读懂基础技术资料(论文/技术博客/产品白皮书),并能翻译
更新于 2026-06-18北京