logo of apple

苹果Site Reliability Engineer — Insight Team

社招全职Software and Services地点:上海状态:招聘

任职要求


Minimum Qualifications
BS or MS in Computer Science, Software Engineering, or an equivalent technical discipline 
Good programming skills in Java, Python or Go
Hands-on experience in SRE, DevOps, or platform engineering, with demonstrated ability to work on technical initiatives
Experience applying AI/ML techniques to development & operations 
Proficiency in English for clear technical communication, documentation, and global collaboration
Availability to participate in SRE on-call rotations during China morning hours (8:00 AM CST during US Daylight Saving Time, and 9:00 AM CST during US Standard Time)


Preferred Qualifications
Cloud Technologies: Hands-on experience with cloud platforms (AWS or GCP), containerization and orchestration (Docker, Kubernetes), and multi-region infrastructure d…
登录查看完整任职要求
微信扫码,1秒登录

工作职责


You will operate and improve services within a very large-scale, highly available Big Data ecosystem supporting Exabytes level of data with sustained, rapid growth. Your work directly enables the engineering and operations teams that build every Apple product.
-	Own the reliability, performance, and scalability of services and cloud infrastructure within the Insight ecosystem
-	Build and advance AIOps capabilities across the ecosystem — developing AI-driven alerting, anomaly detection, LLM-assisted operational tooling, and automated incident triage that measurably improve observability and response
-	Instrument services for deep observability: dashboards, meaningful alerts, runbooks, and clearly defined SLOs and SLIs that give the team and stakeholders an accurate view of ecosystem health

-	Partner with incident management to drive effective incident response and conduct thorough post-incident reviews; translate findings into architectural and operational improvements that prevent recurrence
-	Collaborate with cross-functional engineering teams across Apple's global manufacturing services, communicating effectively across time zones and cultures
包括英文材料
Go+
Java+
Python+
DevOps+
AWS+
Docker+
Kubernetes+
Kafka+
MySQL+
PostgreSQL+
Redis+
还有更多 •••
相关职位

logo of jd
社招运维工程师岗

1. 负责国际业务线上服务的稳定性保障与性能优化,通过监控、告警、容量规划及故障应急处理,确保核心业务系统的高可用性与可扩展性; 2. 参与系统架构设计与技术选型,构建高可靠、高性能的运维体系,推动自动化工具链与平台建设,提升运维效率与标准化水平; 3. 深入理解业务需求,通过日志分析、链路追踪与性能剖析,定位系统瓶颈,实施针对性优化方案,支撑业务快速增长与稳定运行; 4. 负责变更流程管理与风险控制,制定并执行应急预案,主导故障复盘与改进措施落地,持续提升系统韧性与团队应急响应能力; 5. 与研发、产品及安全团队紧密协作,推动DevOps文化落地,完善CI/CD流程与安全合规实践,保障业务交付质量与效率; 6. 跟踪业界先进运维技术与最佳实践,结合业务场景进行技术探索与创新,推动运维体系持续演进,支撑业务国际化拓展与战略目标实现; 7. 日常在英国伦敦办公室工作。

更新于 2026-06-17北京
logo of tesla
社招基础架构

The Role We are Tesla Infrastructure SRE team, responsible for ensuring the stability of Tesla's globally distributed business operations in China. Our team spans two technical pillars - Traffic and Splunk - supporting critical workloads across manufacturing, energy, connected vehicles, and global content delivery. Traffic owns Tesla core traffic infrastructure - including F5 load balancing, CDN/Edge, Varnish/Envoy, L7 proxies, and cross-site high-availability architectures - ensuring reliable delivery of Tesla official website, vehicle firmware, streaming media, and APIs. Splunk owns the build, operation, and enablement of Tesla Splunk observability platform - covering data ingestion, indexing, search, alerting, dashboards, and knowledge base curation - empowering business teams to drive incident response and service improvement with data. We are building AI-native operations capabilities - integrating Large Language Models (LLMs), MCP toolchains, and Splunk / Confluence knowledge bases to enable automated alert investigation, intelligent runbook retrieval, and troubleshooting assistance, so engineers can focus on judgment and decision-making rather than repetitive context-switching and manual queries. This is a hybrid role with a primary focus on Traffic: under the guidance of senior engineers, you will spend most of your time on traffic infrastructure operations and delivery, while also growing into Splunk platform work as a secondary pillar. Responsibilities Shared Deploy, monitor, troubleshoot, and maintain production systems to ensure platform availability. Develop automation scripts and operational tooling (Shell / Python / Go) to reduce toil and improve troubleshooting efficiency. Participate in on-call (24/7) rotations, respond to alerts per runbooks, and assist with root cause analysis and issue closure. Collaborate with global Traffic and Splunk teams and China business stakeholders to drive standardization and best practices. Traffic Assist with configuration and change execution for traffic components including F5 load balancing, Varnish, and Envoy. Participate in HTTP/TCP/SSL/TLS troubleshooting and assist with traffic anomaly and availability analysis. Configure and manage TCP/L7 proxy technologies to ensure efficient traffic management and content delivery. Participate in testing and validation of cross-site traffic architecture changes and cross-border traffic analysis. Troubleshoot and resolve critical issues across multiple layers, including CDN, Loadbalancers, storage, OS, network, K8S, virtualization, and application/DB stack. Splunk Assist with deployment, configuration, and routine health checks of Splunk cluster components (UF, Indexer, Search Head, etc.). Support business teams with machine data onboarding, field extraction, index planning, and search optimization. Help create and maintain alerts, dashboards, and SPL queries; troubleshoot common search performance issues. Build up user community to grow end-user capabilities in using Splunk as a powerful tool.

上海
logo of tesla
社招软件平台

THE ROLE We're the small, expert team creating the next-generation server-side infrastructure to support the manufacturing and functionality of fleets of Tesla products, and we're looking for seasoned SREs with domain expertise in one or more of: containers, public clouds and cloud-native apps. Today, Tesla owners rely on our services to safely and securely summon their cars with a tap on their mobile phones -- a feature enabled by one of the many over-the-air updates we've delivered to the Tesla vehicle fleet. Tesla engineering relies on our data and analytics platform to make Tesla products better and safer. And, when an owner needs assistance, Tesla service and support rely our applications to understand and respond to the situation. Tomorrow, we will apply fleet learning to dispatch and deliver real-time road conditions to millions of autonomous vehicles and manage distributed energy generation & storage at grid scale. Join us and you will work alongside world-class software and data engineers on some of the newest and most challenging IoT, manufacturing and service engineering problems in the world today. The platform you help us build and automate will be used daily by millions of Tesla owners (and tens of thousands of Tesla employees) to improve and enhance the functionality of our cars, chargers, and batteries worldwide. RESPONSIBILITIES Design and write software that enables rapid prototyping by development teams, while ensuring the highest levels of reliability and availability. Work directly with our factory firmware team to provide highly available factory-facing services. Drive the migration of large-scale, distributed fleet applications towards cloud-native microservices. Influence architectural decisions with focus on security, scalability and high-performance. Automate the build and deployment of infrastructure using Docker, Kubernetes & other orchestration technologies in a hybrid-cloud environment. Setup and maintain monitoring, metrics & reporting systems for fine-grained observability and actionable alerting.

上海
logo of tesla
社招IT-基础架构与

THE ROLE Tesla's Platform Engineering is looking for a Site Reliability Engineer to join our team. As a member of the team, you will be building and maintaining Kubernetes clusters using infrastructure-as-code tools like Ansible, Terraform, ArgoCD and Helm and helping the application teams to be successful on our platform. The underlying infrastructure is a a mix of on-premises VMs, bare metal hosts and public clouds such as AWS located all around the globe, which presents unique challenges and opportunity to work with different types of infrastructure technologies. A successful candidate will be expected to possess expert knowledge in Linux fundamentals, architecture and performance tuning; as well as software development skills to match. Experience running Kubernetes in production will be a strong plus. We prefer Golang or Python for any automation or tools we have to build along the way. We are the team that runs production critical workloads for every aspect of the business at Tesla and sets the standards for other teams, a group of well-rounded generalists that not only solve the hardest problems in the industry but also push other engineering teams at large to be better. RESPONSIBILITIES • Manage our Kubernetes clusters on-prem and in the cloud to support our growing workloads. • Participating in the architecture design process and troubleshooting of live applications with the product teams. • Participating in a 24x7 on-call rotation. • Influence architectural decisions with focus on security, scalability and high-performance. • Setup and maintain monitoring, metrics & reporting systems for fine-grained observability and actionable alerting. • Authoring technical documentation for workflows/processes/best practices.

上海