特斯拉Sr. Site Reliability Engineer, Splunk
任职要求
Core Skills 7+ years of systems administration experience, with strong Linux background (RHEL / CentOS / Ubuntu, etc.). 5+ years designing, maintaining, and troubleshooting mid-to-large-scale Splunk infrastructure; deep understanding of distributed Splunk architecture. Strong SPL skills for complex queries; solid grasp of best practices for reports, alerts, and dashboards. Hands-on experience operating data-ingestion platforms (collection, processing, routing, and integration with downstream systems). Experience with large-scale distributed systems and high-availability architectures; strong analytical and problem-solving skills. Preferred Bachelor’s degree in Engineering, Mathematics, Computer Science, Information Technology, or equivalent experience. Experience extending the Splunk ecosystem (Custom Commands, Modular Inputs, App development, REST API, external service integration). AIOps experience (alert noise reduction, event correlatio…
工作职责
The Role At Tesla, logs from software and hardware systems—including applications and microservices, operating systems and middleware, production-line equipment, vehicle-related systems, and network and security components—are ingested into Splunk as the standard platform for enterprise log analytics and troubleshooting. Splunk supports cross-system root-cause analysis, reduces MTTD/MTTR, helps sustain high availability of production and business systems, and provides a unified data entry point for security audit, change impact assessment, capacity planning, and compliance analysis. Platform stability, search performance, and data quality directly affect business continuity and risk control. This role is a Site Reliability Engineer (SRE) responsible for the architecture, build, and operations of the Splunk logging platform, data ingestion pipelines, and observability capabilities. The focus is on automation and platform-ready delivery, in close collaboration with application, manufacturing, and infrastructure teams, to ensure high-quality ingestion and efficient retrieval of large-scale logs that support troubleshooting and business decisions. Responsibilities 1. Splunk Platform & Reliability Design, build, and maintain Splunk infrastructure, including Multi-Site deployments and high-availability (HA) architectures, ensuring reliability, scalability, security, and high performance. Operate data-ingestion platforms (pipelines, routing, filtering/enrichment, and integration with Splunk) to improve ingestion quality, cost efficiency, and operability. Drive machine-data onboarding, field extraction, search and data-model optimization; create alerts, troubleshoot search performance issues, and define data storage and lifecycle policies. Develop automation, monitoring, and diagnostic tools; participate in on-call, respond quickly to bridge calls, minimize incident impact; mentor junior engineers. 2. Observability (Log / Metric / Trace) Design and implement enterprise observability solutions that support service health assessment, dependency analysis, and fault localization. Own architecture, ingestion, storage, query, and capacity management for Prometheus / Mimir; establish dashboard and alerting standards with Grafana. Deploy, upgrade, scale, and troubleshoot observability components and collection pipelines in Kubernetes environments. Drive correlation and unified views across Log / Metric / Trace; co-define naming, labeling, collection, and SLO/alerting standards. 3. Splunk AI OPS Design and deliver AI tools and services for troubleshooting, analysis, reporting, and knowledge Q&A (e.g., intelligent search assistance, alert interpretation, root-cause suggestions, runbooks, and ops assistants). Build platform AI operations capabilities—alert noise reduction and correlation, intelligent recommendations, and closed-loop feedback—integrating securely with Splunk search, alerts, dashboards, permissions, and audit; measure outcomes via adoption, accuracy, MTTR, and noise rate. Document reusable tools and best practices to improve how users leverage Splunk and observability capabilities. Minimum
THE ROLE Tesla's Platform Engineering is looking for a Site Reliability Engineer to join our team. As a member of the team, you will be building and maintaining Kubernetes clusters using infrastructure-as-code tools like Ansible, Terraform, ArgoCD and Helm and helping the application teams to be successful on our platform. The underlying infrastructure is a a mix of on-premises VMs, bare metal hosts and public clouds such as AWS located all around the globe, which presents unique challenges and opportunity to work with different types of infrastructure technologies. A successful candidate will be expected to possess expert knowledge in Linux fundamentals, architecture and performance tuning; as well as software development skills to match. Experience running Kubernetes in production will be a strong plus. We prefer Golang or Python for any automation or tools we have to build along the way. We are the team that runs production critical workloads for every aspect of the business at Tesla and sets the standards for other teams, a group of well-rounded generalists that not only solve the hardest problems in the industry but also push other engineering teams at large to be better. RESPONSIBILITIES • Manage our Kubernetes clusters on-prem and in the cloud to support our growing workloads. • Participating in the architecture design process and troubleshooting of live applications with the product teams. • Participating in a 24x7 on-call rotation. • Influence architectural decisions with focus on security, scalability and high-performance. • Setup and maintain monitoring, metrics & reporting systems for fine-grained observability and actionable alerting. • Authoring technical documentation for workflows/processes/best practices.
STM (Site Technical Manager) is the single threaded owner for manufacturing technical issues for assigned projects with our Contract Manufacturers, focus on: - Ensure technical readiness for product ramp and serve as manufacturing engineering owner for product PVT and mass production. - Driving manufacture test flow optimization, process and yield improvements to exceed our volume production goals. - Define the manufacture process qualification criteria, manage the qualification activities, and complete the documentation. - Leading our efforts for root cause and corrective action for all manufacturing process related issues, dive deep to analyze the problems found both in and out of the factory. - Participation in design & planning through DFM, Line balancing , as well as Process Yield, Capacity and Cost Modeling. - Work with CM’s engineering teams to identify and escalate manufacturing challenges by enforcing DFM and DFT principles. - Review and approval of Fixtures Designs & Qualification (FATP, Device build) - Review and execute Manufacturing Test Coverage Documents to ensure new products launch. - Work with sustaining engineering team on the opportunities to improve our product, owns the technical readiness of PRQ in CMs. - Periodically audit the CM for manufacture process quality management system.
- Drive two-digital revenue and market share growth within assigned strategic/named accounts in the Logistics sector - Develop and execute comprehensive account plans and manage all key customer relationships, including strong C-level engagement - Accelerate customer adoption by identifying new opportunities, expanding existing usage, and leading customers through cloud migration journeys - Maintain a robust customer pipeline with accurate forecasting and reporting - Manage complex contract negotiations and work with partners to extend reach and drive adoption
The Role We are Tesla Energy Team. We are building highly distributed energy network to support company vision of accelerate world’s transition to a sustainable energy. The Energy Team is seeking hardworking and passionate software engineers as various levels. Responsibilities • Design and develop high quality, scalable and stable back-end services. • Develop back-end web services. • Follow Tesla’s high standards for security-best practices in all development. • Partner closely with security team for code analysis and design reviews. • Perform unit testing. • Process bug reports and release fixes. • Participate in code reviews. • Participate in agile processes. • Always think innovatively to solve customer problems.