logo of tesla

特斯拉Site Reliability Engineer, Traffic&Splunk

社招全职基础架构地点:上海状态:招聘

任职要求


Must Have
2-4 years of Linux systems operations or SRE experience.
Solid understanding of TCP/IP, HTTP, DNS, and other core networking protocols.
Deep understanding of CDN and TCP/L7 proxy technologies. Hands-on experience with load balancers such as F5 OR proxy technology.
Understanding of CDN/Edge networking, including topology, protocols, hardware, and architecture.
Foundational knowledge of containerization technologies such as Kubernetes.
Experience with configuration and management using tools such as Ansible or Puppet.
Proficiency in Shell; ability to write automation scripts in Python or Go.
Strong analytical thinking, problem-solving, and communication skills; ability to read and write technical documentation in Chinese and English.
High sense of ownership; comfortable with on-call rotations and a fast-paced production environment.
Able…
登录查看完整任职要求
微信扫码,1秒登录

工作职责


The Role
We are Tesla Infrastructure SRE team, responsible for ensuring the stability of Tesla's globally distributed business operations in China. Our team spans two technical pillars - Traffic and Splunk - supporting critical workloads across manufacturing, energy, connected vehicles, and global content delivery.

Traffic owns Tesla core traffic infrastructure - including F5 load balancing, CDN/Edge, Varnish/Envoy, L7 proxies, and cross-site high-availability architectures - ensuring reliable delivery of Tesla official website, vehicle firmware, streaming media, and APIs.
Splunk owns the build, operation, and enablement of Tesla Splunk observability platform - covering data ingestion, indexing, search, alerting, dashboards, and knowledge base curation - empowering business teams to drive incident response and service improvement with data.

We are building AI-native operations capabilities - integrating Large Language Models (LLMs), MCP toolchains, and Splunk / Confluence knowledge bases to enable automated alert investigation, intelligent runbook retrieval, and troubleshooting assistance, so engineers can focus on judgment and decision-making rather than repetitive context-switching and manual queries.
This is a hybrid role with a primary focus on Traffic: under the guidance of senior engineers, you will spend most of your time on traffic infrastructure operations and delivery, while also growing into Splunk platform work as a secondary pillar.

Responsibilities
Shared
Deploy, monitor, troubleshoot, and maintain production systems to ensure platform availability.
Develop automation scripts and operational tooling (Shell / Python / Go) to reduce toil and improve troubleshooting efficiency.
Participate in on-call (24/7) rotations, respond to alerts per runbooks, and assist with root cause analysis and issue closure.
Collaborate with global Traffic and Splunk teams and China business stakeholders to drive standardization and best practices.

Traffic
Assist with configuration and change execution for traffic components including F5 load balancing, Varnish, and Envoy.
Participate in HTTP/TCP/SSL/TLS troubleshooting and assist with traffic anomaly and availability analysis.
Configure and manage TCP/L7 proxy technologies to ensure efficient traffic management and content delivery.
Participate in testing and validation of cross-site traffic architecture changes and cross-border traffic analysis.
Troubleshoot and resolve critical issues across multiple layers, including CDN, Loadbalancers, storage, OS, network, K8S, virtualization, and application/DB stack.

Splunk
Assist with deployment, configuration, and routine health checks of Splunk cluster components (UF, Indexer, Search Head, etc.).
Support business teams with machine data onboarding, field extraction, index planning, and search optimization.
Help create and maintain alerts, dashboards, and SPL queries; troubleshoot common search performance issues.
Build up user community to grow end-user capabilities in using Splunk as a powerful tool.
包括英文材料
Linux+
TCP/IP+
HTTP+
Kubernetes+
Ansible+
Bash+
Python+
Go+
还有更多 •••
相关职位

logo of jd
社招运维工程师岗

1. 负责国际业务线上服务的稳定性保障与性能优化,通过监控、告警、容量规划及故障应急处理,确保核心业务系统的高可用性与可扩展性; 2. 参与系统架构设计与技术选型,构建高可靠、高性能的运维体系,推动自动化工具链与平台建设,提升运维效率与标准化水平; 3. 深入理解业务需求,通过日志分析、链路追踪与性能剖析,定位系统瓶颈,实施针对性优化方案,支撑业务快速增长与稳定运行; 4. 负责变更流程管理与风险控制,制定并执行应急预案,主导故障复盘与改进措施落地,持续提升系统韧性与团队应急响应能力; 5. 与研发、产品及安全团队紧密协作,推动DevOps文化落地,完善CI/CD流程与安全合规实践,保障业务交付质量与效率; 6. 跟踪业界先进运维技术与最佳实践,结合业务场景进行技术探索与创新,推动运维体系持续演进,支撑业务国际化拓展与战略目标实现; 7. 日常在英国伦敦办公室工作。

更新于 2026-06-17北京
logo of tesla
社招软件平台

THE ROLE We're the small, expert team creating the next-generation server-side infrastructure to support the manufacturing and functionality of fleets of Tesla products, and we're looking for seasoned SREs with domain expertise in one or more of: containers, public clouds and cloud-native apps. Today, Tesla owners rely on our services to safely and securely summon their cars with a tap on their mobile phones -- a feature enabled by one of the many over-the-air updates we've delivered to the Tesla vehicle fleet. Tesla engineering relies on our data and analytics platform to make Tesla products better and safer. And, when an owner needs assistance, Tesla service and support rely our applications to understand and respond to the situation. Tomorrow, we will apply fleet learning to dispatch and deliver real-time road conditions to millions of autonomous vehicles and manage distributed energy generation & storage at grid scale. Join us and you will work alongside world-class software and data engineers on some of the newest and most challenging IoT, manufacturing and service engineering problems in the world today. The platform you help us build and automate will be used daily by millions of Tesla owners (and tens of thousands of Tesla employees) to improve and enhance the functionality of our cars, chargers, and batteries worldwide. RESPONSIBILITIES Design and write software that enables rapid prototyping by development teams, while ensuring the highest levels of reliability and availability. Work directly with our factory firmware team to provide highly available factory-facing services. Drive the migration of large-scale, distributed fleet applications towards cloud-native microservices. Influence architectural decisions with focus on security, scalability and high-performance. Automate the build and deployment of infrastructure using Docker, Kubernetes & other orchestration technologies in a hybrid-cloud environment. Setup and maintain monitoring, metrics & reporting systems for fine-grained observability and actionable alerting.

上海
logo of tesla
社招IT-基础架构与

THE ROLE Tesla's Platform Engineering is looking for a Site Reliability Engineer to join our team. As a member of the team, you will be building and maintaining Kubernetes clusters using infrastructure-as-code tools like Ansible, Terraform, ArgoCD and Helm and helping the application teams to be successful on our platform. The underlying infrastructure is a a mix of on-premises VMs, bare metal hosts and public clouds such as AWS located all around the globe, which presents unique challenges and opportunity to work with different types of infrastructure technologies. A successful candidate will be expected to possess expert knowledge in Linux fundamentals, architecture and performance tuning; as well as software development skills to match. Experience running Kubernetes in production will be a strong plus. We prefer Golang or Python for any automation or tools we have to build along the way. We are the team that runs production critical workloads for every aspect of the business at Tesla and sets the standards for other teams, a group of well-rounded generalists that not only solve the hardest problems in the industry but also push other engineering teams at large to be better. RESPONSIBILITIES • Manage our Kubernetes clusters on-prem and in the cloud to support our growing workloads. • Participating in the architecture design process and troubleshooting of live applications with the product teams. • Participating in a 24x7 on-call rotation. • Influence architectural decisions with focus on security, scalability and high-performance. • Setup and maintain monitoring, metrics & reporting systems for fine-grained observability and actionable alerting. • Authoring technical documentation for workflows/processes/best practices.

上海
logo of ymtc
社招5年以上系统解决方案类

1)代工厂导入,生产测试环境设计,测试系统搭建 2)协助测试平台导入,测试OI 功能模块设计,控制逻辑定义,操作防呆和优化 3)协助新产品导入、整个生命周期的流程管控、硬件测试以及重大质量问题的跟踪解决 4)开发自动化脚本,监控测试数据,持续提升测试良率,和相关部门解决质量问题,保证出货产品的品质 5)参与制定产品测试流程,开发相应测试平台上的软/硬件系统,并导入到代工厂 6)分析并解决代工厂测试问题, 跟进问题批次直到Release。 7)代工厂管理, 推动代工厂工作改善,维护代工厂的合作关系

更新于 2025-08-01武汉|上海