智谱SRE运维工程师
社招全职5年以上工程研发地点:北京状态:招聘
工作描述
岗位职责: 负责面向全球ToC的业务稳定性建设,包括但不限于Web服务、APP后端、API网关、云基础设施、安全 负责线上业务服务7x24小时运维保障,及时定位处理业务故障,快速恢复应用服务 配合研发团队完成业务架构选型、应用部署、版本发布、监控接入、链路调优及日常维护 负责Kubernetes集群的建设与稳定性保障,包括版本升级、故障排查、性能调优、资源利用率优化 负责设计基于云平台基础设施的高可用架构,保障基础设施、中间件、业务服务的稳定性 负责云平台资源生命周期管理及成本优化 主导容器化架构调优(如Pod调度策略、网络插件选型、存储方案设计),优化资源请求/限制配置以减少资源争用。 建立容器安全防护体系,包括漏洞扫描、运行时…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
Web+
https://web.dev/learn
Explore our growing collection of courses on key web design and development subjects.
Kubernetes+
https://kubernetes.io/docs/tutorials/kubernetes-basics/
This tutorial provides a walkthrough of the basics of the Kubernetes cluster orchestration system.
https://kubernetes.io/zh-cn/docs/tutorials/kubernetes-basics/
本教程介绍 Kubernetes 集群编排系统的基础知识。每个模块包含关于 Kubernetes 主要特性和概念的一些背景信息,还包括一个在线教程供你学习。
https://www.youtube.com/watch?v=s_o8dwzRlu4
Hands-On Kubernetes Tutorial | Learn Kubernetes in 1 Hour - Kubernetes Course for Beginners
https://www.youtube.com/watch?v=X48VuDVv0do
Full Kubernetes Tutorial | Kubernetes Course | Hands-on course with a lot of demos
性能调优+
https://goperf.dev/
The Go App Optimization Guide is a series of in-depth, technical articles for developers who want to get more performance out of their Go code without relying on guesswork or cargo cult patterns.
https://web.dev/learn/performance
This course is designed for those new to web performance, a vital aspect of the user experience.
https://www.ibm.com/think/insights/application-performance-optimization
Application performance is not just a simple concern for most organizations; it’s a critical factor in their business’s success.
https://www.oreilly.com/library/view/optimizing-java/9781492039259/
Performance tuning is an experimental science, but that doesn’t mean engineers should resort to guesswork and folklore to get the job done.
高可用+
https://redis.io/blog/high-availability-architecture/
A high available architecture is when there are a number of different components, modules, or services that work together to maintain optimal performance, irrespective of peak-time loads.
https://www.ibm.com/think/topics/high-availability
High availability (HA) is a term that refers to a system’s ability to be accessible and reliable close to 100% of the time.
中间件+
https://www.youtube.com/watch?v=1oWPUpMheGk
安全防护+
https://roadmap.sh/cyber-security
Step by step guide to becoming a Cyber Security Expert
https://www.w3schools.com/cybersecurity/
This course serves as an excellent primer to the many different domains of Cyber security.
Falco+
https://falco.org/docs/getting-started/
Falco is a cloud native security tool. It provides near real-time threat detection for cloud, container, and Kubernetes workloads by leveraging runtime insights.
CI+
https://www.ibm.com/cn-zh/think/topics/continuous-integration
持续集成 (CI) 是一种软件开发实践,开发人员在整个开发周期中会定期将新的代码和代码变更集成到中央代码存储库中。它是 DevOps 和敏捷方法的关键组成部分。
https://www.youtube.com/watch?v=42UP1fxi2SY
CD+
https://www.redhat.com/zh-cn/topics/devops/what-is-ci-cd
CI/CD 是持续集成和持续交付/部署的缩写,旨在简化并加快软件开发生命周期。
https://www.youtube.com/watch?v=R8_veQiYBjI&list=PLy7NrYWoggjzSIlwxeBbcgfAdYoxCIrM2
Prometheus+
https://grafana.com/docs/grafana/latest/getting-started/get-started-grafana-prometheus/
Prometheus is an open source monitoring system for which Grafana provides out-of-the-box support.
https://prometheus.io/docs/tutorials/getting_started/
Prometheus is a system monitoring and alerting system.
数据分析+
[英文] Data Analyst Roadmap
https://roadmap.sh/data-analyst
Step by step guide to becoming an Data Analyst in 2025
还有更多 •••
相关职位

社招5年以上运维
1. 计算机相关专业,五年以上工作经验; 2. 熟悉Linux操作系统,了解Linux操作系统基本原理; 3. 熟悉Elk、Prometheus、Grafana等监控日志工具使用; 4. 熟悉虚拟化和
更新于 2026-07-28北京|苏州|上海

社招5年以上
任职要求: 1、5 年以上互联网运维/运维开发经验,本科及以上学历,计算机相关专业优先; 2、精通 Linux 操作系统与网络原理,熟练掌握 Python/Golang/Shell 至少一门脚本语言,
更新于 2026-06-30苏州|北京
社招3年以上程序&技术类
1.全日制本科,3年以上相关工作经验; 2.熟悉至少一种主流公有云/私有云,包括但不限于Aliyun/腾讯云/AWS/GCP; 3.有大型中间件系统或基础设施运维经验者优先; 4.有丰富的生产环境故障
上海