苹果Site Reliability Engineer, Ai & Data Platforms
任职要求
Minimum Qualifications 4+ years of experience in SRE, DevOps, platform engineering, data infrastructure, or production operations. Strong Linux troubleshooting skills. Strong Kubernetes/EKS operations experience, including kubectl, deployments, statefulsets, pods, resource limits, service accounts, secrets, logs, events, and workload debugging. Hands-on experience supporting Apache Spark in production, preferably Spark on Kubernetes. Experience tuning Spark jobs, reading driver/executor logs, troubleshooting memory issues, shuffle failures, failed stages, and performance bottlenecks. Production experience with Apache Airflow, including DAG operations, task failures, retries, scheduling, SLA misses, and dependency troubleshooting. Practical experience with Kafka or similar messaging/streaming platforms, including consumer groups, lag, offsets, partitions, and secure client connectivity. Familiarity with Snowflake loading patterns, connectivity, roles, warehouses, stages, query/load history, and load monitoring. Familiarity with Datalake or Lakehouse architectures using object storage such as S3 or equivalent. Solid SQL skills and ability to analyze large operational or data pipeline datasets. Experienc…
工作职责
Key Responsibility • Operate and support production ETL pipelines, including extractors, loaders, batch jobs, streaming ingestion, and data consolidation workflows. • Support loader jobs running in an EKS Airflow cluster that load data into Snowflake. • Support Spark jobs running on EKS that load data into the Datalake or Lakehouse. • Own platform and pipeline reliability, including uptime, throughput, job success rate, SLA adherence, and data freshness. • Triage and resolve production incidents such as failed jobs, delayed loads, OOMKilled pods, CrashLoopBackOff pods, stalled pipelines, driver/executor failures, credential issues, and network/connectivity problems. • Perform root cause analysis for incidents and drive corrective actions through automation, configuration changes, capacity tuning, or code/platform improvements. • Tune Kubernetes workloads, including CPU/memory requests and limits, autoscaling, pod scheduling, service accounts, secrets, config maps, and workload health. • Tune Spark workloads, including executor count, cores, memory, memory overhead, shuffle partitions, dynamic allocation, adaptive query execution, spill reduction, and failed stage analysis. • Monitor and troubleshoot Airflow DAGs, task retries, scheduling issues, dependency failures, executor/resource constraints, and SLA misses. • Diagnose data latency across the full pipeline path, such as source event, Kafka or messaging layer, staging/object storage, ETL processing, Snowflake, and Datalake availability. • Support Kafka or equivalent streaming systems, including consumer lag, offsets, partitions, topic health, producer delay, and SASL/TLS client configuration. • Manage platform and job configuration through approved source-of-truth or GitOps-style processes instead of ad hoc production changes. • Manage secrets, certificates, truststores, JDBC credentials, Snowflake credentials, object-storage credentials, and key rotations following least-privilege practices. • Build and maintain dashboards, alerts, log queries, and runbooks for job health, freshness, backlog, failures, infrastructure usage, and incident response. • Write scripts and automation using Python, Bash, or similar tools to reduce toil and improve recovery speed. • Coordinate with application engineering, data engineering, platform, security, network, Snowflake, and Datalake teams. • Participate in on-call or production support rotation for China business hours and critical incidents.
1. 负责国际业务线上服务的稳定性保障与性能优化,通过监控、告警、容量规划及故障应急处理,确保核心业务系统的高可用性与可扩展性; 2. 参与系统架构设计与技术选型,构建高可靠、高性能的运维体系,推动自动化工具链与平台建设,提升运维效率与标准化水平; 3. 深入理解业务需求,通过日志分析、链路追踪与性能剖析,定位系统瓶颈,实施针对性优化方案,支撑业务快速增长与稳定运行; 4. 负责变更流程管理与风险控制,制定并执行应急预案,主导故障复盘与改进措施落地,持续提升系统韧性与团队应急响应能力; 5. 与研发、产品及安全团队紧密协作,推动DevOps文化落地,完善CI/CD流程与安全合规实践,保障业务交付质量与效率; 6. 跟踪业界先进运维技术与最佳实践,结合业务场景进行技术探索与创新,推动运维体系持续演进,支撑业务国际化拓展与战略目标实现; 7. 日常在英国伦敦办公室工作。
The Role We are Tesla Infrastructure SRE team, responsible for ensuring the stability of Tesla's globally distributed business operations in China. Our team spans two technical pillars - Traffic and Splunk - supporting critical workloads across manufacturing, energy, connected vehicles, and global content delivery. Traffic owns Tesla core traffic infrastructure - including F5 load balancing, CDN/Edge, Varnish/Envoy, L7 proxies, and cross-site high-availability architectures - ensuring reliable delivery of Tesla official website, vehicle firmware, streaming media, and APIs. Splunk owns the build, operation, and enablement of Tesla Splunk observability platform - covering data ingestion, indexing, search, alerting, dashboards, and knowledge base curation - empowering business teams to drive incident response and service improvement with data. We are building AI-native operations capabilities - integrating Large Language Models (LLMs), MCP toolchains, and Splunk / Confluence knowledge bases to enable automated alert investigation, intelligent runbook retrieval, and troubleshooting assistance, so engineers can focus on judgment and decision-making rather than repetitive context-switching and manual queries. This is a hybrid role with a primary focus on Traffic: under the guidance of senior engineers, you will spend most of your time on traffic infrastructure operations and delivery, while also growing into Splunk platform work as a secondary pillar. Responsibilities Shared Deploy, monitor, troubleshoot, and maintain production systems to ensure platform availability. Develop automation scripts and operational tooling (Shell / Python / Go) to reduce toil and improve troubleshooting efficiency. Participate in on-call (24/7) rotations, respond to alerts per runbooks, and assist with root cause analysis and issue closure. Collaborate with global Traffic and Splunk teams and China business stakeholders to drive standardization and best practices. Traffic Assist with configuration and change execution for traffic components including F5 load balancing, Varnish, and Envoy. Participate in HTTP/TCP/SSL/TLS troubleshooting and assist with traffic anomaly and availability analysis. Configure and manage TCP/L7 proxy technologies to ensure efficient traffic management and content delivery. Participate in testing and validation of cross-site traffic architecture changes and cross-border traffic analysis. Troubleshoot and resolve critical issues across multiple layers, including CDN, Loadbalancers, storage, OS, network, K8S, virtualization, and application/DB stack. Splunk Assist with deployment, configuration, and routine health checks of Splunk cluster components (UF, Indexer, Search Head, etc.). Support business teams with machine data onboarding, field extraction, index planning, and search optimization. Help create and maintain alerts, dashboards, and SPL queries; troubleshoot common search performance issues. Build up user community to grow end-user capabilities in using Splunk as a powerful tool.
You will operate and improve services within a very large-scale, highly available Big Data ecosystem supporting Exabytes level of data with sustained, rapid growth. Your work directly enables the engineering and operations teams that build every Apple product. - Own the reliability, performance, and scalability of services and cloud infrastructure within the Insight ecosystem - Build and advance AIOps capabilities across the ecosystem — developing AI-driven alerting, anomaly detection, LLM-assisted operational tooling, and automated incident triage that measurably improve observability and response - Instrument services for deep observability: dashboards, meaningful alerts, runbooks, and clearly defined SLOs and SLIs that give the team and stakeholders an accurate view of ecosystem health - Partner with incident management to drive effective incident response and conduct thorough post-incident reviews; translate findings into architectural and operational improvements that prevent recurrence - Collaborate with cross-functional engineering teams across Apple's global manufacturing services, communicating effectively across time zones and cultures
Engage with our product teams to understand requirements, design and implement resilient and scalable infrastructure solutions. Operate, monitor, and triage all aspects of our production and non-production environments. Collaborate on code, infrastructure, design reviews, and process enhancements Evaluate and integrate new technologies to improve system reliability, security, and performance. Develop and implement automation to provision, configure, deploy, and monitor Apple services. Participate in an oncall rotation providing hands-on technical expertise during service impacting events. Contribute to capacity planning, scale testing, and disaster recovery exercises Approach operational problems with a software engineering mindset.