京东Site Reliability Engineer
任职要求
1. 学历要求:本科及以上学历,计算机科学、软件工程、网络工程、信息安全等相关专业; 2. 工作经验:不限工作经验,具备互联网或相关行业运维、系统稳定性保障、高可用架构设计经验者优先; 3. 能力要求: 技术能力:熟悉Linux操作系统、网络协议、分布式系统原理;掌握至少一种脚本语言(如Shell、Python);具备容器化技术(如Docker、Kubernetes)及云平台运维实践经验;熟悉监控告警、日志分析、自动化运维工具; 业务理解:能够理解业务需求,参与系统架构设计与稳定性规划,保障业务系统高可用与性能优化; 项目管理:具备跨团队…
工作职责
1. 负责国际业务线上服务的稳定性保障与性能优化,通过监控、告警、容量规划及故障应急处理,确保核心业务系统的高可用性与可扩展性; 2. 参与系统架构设计与技术选型,构建高可靠、高性能的运维体系,推动自动化工具链与平台建设,提升运维效率与标准化水平; 3. 深入理解业务需求,通过日志分析、链路追踪与性能剖析,定位系统瓶颈,实施针对性优化方案,支撑业务快速增长与稳定运行; 4. 负责变更流程管理与风险控制,制定并执行应急预案,主导故障复盘与改进措施落地,持续提升系统韧性与团队应急响应能力; 5. 与研发、产品及安全团队紧密协作,推动DevOps文化落地,完善CI/CD流程与安全合规实践,保障业务交付质量与效率; 6. 跟踪业界先进运维技术与最佳实践,结合业务场景进行技术探索与创新,推动运维体系持续演进,支撑业务国际化拓展与战略目标实现; 7. 日常在英国伦敦办公室工作。
The Role We are Tesla Infrastructure SRE team, responsible for ensuring the stability of Tesla's globally distributed business operations in China. Our team spans two technical pillars - Traffic and Splunk - supporting critical workloads across manufacturing, energy, connected vehicles, and global content delivery. Traffic owns Tesla core traffic infrastructure - including F5 load balancing, CDN/Edge, Varnish/Envoy, L7 proxies, and cross-site high-availability architectures - ensuring reliable delivery of Tesla official website, vehicle firmware, streaming media, and APIs. Splunk owns the build, operation, and enablement of Tesla Splunk observability platform - covering data ingestion, indexing, search, alerting, dashboards, and knowledge base curation - empowering business teams to drive incident response and service improvement with data. We are building AI-native operations capabilities - integrating Large Language Models (LLMs), MCP toolchains, and Splunk / Confluence knowledge bases to enable automated alert investigation, intelligent runbook retrieval, and troubleshooting assistance, so engineers can focus on judgment and decision-making rather than repetitive context-switching and manual queries. This is a hybrid role with a primary focus on Traffic: under the guidance of senior engineers, you will spend most of your time on traffic infrastructure operations and delivery, while also growing into Splunk platform work as a secondary pillar. Responsibilities Shared Deploy, monitor, troubleshoot, and maintain production systems to ensure platform availability. Develop automation scripts and operational tooling (Shell / Python / Go) to reduce toil and improve troubleshooting efficiency. Participate in on-call (24/7) rotations, respond to alerts per runbooks, and assist with root cause analysis and issue closure. Collaborate with global Traffic and Splunk teams and China business stakeholders to drive standardization and best practices. Traffic Assist with configuration and change execution for traffic components including F5 load balancing, Varnish, and Envoy. Participate in HTTP/TCP/SSL/TLS troubleshooting and assist with traffic anomaly and availability analysis. Configure and manage TCP/L7 proxy technologies to ensure efficient traffic management and content delivery. Participate in testing and validation of cross-site traffic architecture changes and cross-border traffic analysis. Troubleshoot and resolve critical issues across multiple layers, including CDN, Loadbalancers, storage, OS, network, K8S, virtualization, and application/DB stack. Splunk Assist with deployment, configuration, and routine health checks of Splunk cluster components (UF, Indexer, Search Head, etc.). Support business teams with machine data onboarding, field extraction, index planning, and search optimization. Help create and maintain alerts, dashboards, and SPL queries; troubleshoot common search performance issues. Build up user community to grow end-user capabilities in using Splunk as a powerful tool.
You will operate and improve services within a very large-scale, highly available Big Data ecosystem supporting Exabytes level of data with sustained, rapid growth. Your work directly enables the engineering and operations teams that build every Apple product. - Own the reliability, performance, and scalability of services and cloud infrastructure within the Insight ecosystem - Build and advance AIOps capabilities across the ecosystem — developing AI-driven alerting, anomaly detection, LLM-assisted operational tooling, and automated incident triage that measurably improve observability and response - Instrument services for deep observability: dashboards, meaningful alerts, runbooks, and clearly defined SLOs and SLIs that give the team and stakeholders an accurate view of ecosystem health - Partner with incident management to drive effective incident response and conduct thorough post-incident reviews; translate findings into architectural and operational improvements that prevent recurrence - Collaborate with cross-functional engineering teams across Apple's global manufacturing services, communicating effectively across time zones and cultures
Engage with our product teams to understand requirements, design and implement resilient and scalable infrastructure solutions. Operate, monitor, and triage all aspects of our production and non-production environments. Collaborate on code, infrastructure, design reviews, and process enhancements Evaluate and integrate new technologies to improve system reliability, security, and performance. Develop and implement automation to provision, configure, deploy, and monitor Apple services. Participate in an oncall rotation providing hands-on technical expertise during service impacting events. Contribute to capacity planning, scale testing, and disaster recovery exercises Approach operational problems with a software engineering mindset.
The Role At Tesla, logs from software and hardware systems—including applications and microservices, operating systems and middleware, production-line equipment, vehicle-related systems, and network and security components—are ingested into Splunk as the standard platform for enterprise log analytics and troubleshooting. Splunk supports cross-system root-cause analysis, reduces MTTD/MTTR, helps sustain high availability of production and business systems, and provides a unified data entry point for security audit, change impact assessment, capacity planning, and compliance analysis. Platform stability, search performance, and data quality directly affect business continuity and risk control. This role is a Site Reliability Engineer (SRE) responsible for the architecture, build, and operations of the Splunk logging platform, data ingestion pipelines, and observability capabilities. The focus is on automation and platform-ready delivery, in close collaboration with application, manufacturing, and infrastructure teams, to ensure high-quality ingestion and efficient retrieval of large-scale logs that support troubleshooting and business decisions. Responsibilities 1. Splunk Platform & Reliability Design, build, and maintain Splunk infrastructure, including Multi-Site deployments and high-availability (HA) architectures, ensuring reliability, scalability, security, and high performance. Operate data-ingestion platforms (pipelines, routing, filtering/enrichment, and integration with Splunk) to improve ingestion quality, cost efficiency, and operability. Drive machine-data onboarding, field extraction, search and data-model optimization; create alerts, troubleshoot search performance issues, and define data storage and lifecycle policies. Develop automation, monitoring, and diagnostic tools; participate in on-call, respond quickly to bridge calls, minimize incident impact; mentor junior engineers. 2. Observability (Log / Metric / Trace) Design and implement enterprise observability solutions that support service health assessment, dependency analysis, and fault localization. Own architecture, ingestion, storage, query, and capacity management for Prometheus / Mimir; establish dashboard and alerting standards with Grafana. Deploy, upgrade, scale, and troubleshoot observability components and collection pipelines in Kubernetes environments. Drive correlation and unified views across Log / Metric / Trace; co-define naming, labeling, collection, and SLO/alerting standards. 3. Splunk AI OPS Design and deliver AI tools and services for troubleshooting, analysis, reporting, and knowledge Q&A (e.g., intelligent search assistance, alert interpretation, root-cause suggestions, runbooks, and ops assistants). Build platform AI operations capabilities—alert noise reduction and correlation, intelligent recommendations, and closed-loop feedback—integrating securely with Splunk search, alerts, dashboards, permissions, and audit; measure outcomes via adoption, accuracy, MTTR, and noise rate. Document reusable tools and best practices to improve how users leverage Splunk and observability capabilities. Minimum