苹果Data Engineer, IS&T Ai & Data Platforms
任职要求
Minimum Qualifications The offered position requires at least a bachelor's degree or equivalent in Computer Science, Information Systems, or a related field. The beneficiary has at least 5 years of subsequent data engineering experience across the required skills. Architecting scalable data processing systems for real-time, near-real-time, and batch data pipelines. Develop data engineering solutions in Python and advanced SQL. Develop self service data engineering applications for business users. Design co…
工作职责
The offered position requires at least 5 years subsequent data engineering experience including the following: • Writing advanced SQL for large scale relational data warehouses and data lakes. • Designing data pipelines for structured and semi-structured data. • Programing data engineering solutions in Python, Java, or Scala. • Working with distributed frameworks including Apache Kafka and Apache Spark and Flink • Designing reporting solutions with Tableau, Streamlit or Superset.
Design, build, and manage cloud-based data warehouses and data access layers that support teams across Apple. Develop highly scalable and secure ETL pipelines to ingest, process, and serve data from a multitude of source systems. Perform data discovery and in-depth analysis across large, multi-source datasets, quickly learning new data to uncover insights and validate quality. Build proofs of concept and run ad-hoc and recurring analyses to answer high-priority business questions. Analyze and optimize existing data systems and queries, including Spark and Trino, for reliability, performance, and cost at scale. Use modern AI tools to enhance and automate data workflows such as orchestration, scheduling, and monitoring. Collaborate closely with US-based and regional teams and business partners, presenting findings and recommendations in English.
Turn business and ML requirements into scalable platform and decisioning systems Build resilient, low-latency services for real-time inference and streaming data Define gold-standard CI/CD, observability, and SLA practices for Data and ML Ship automated pipelines powering features, training, and online serving Drive performance and cost optimizations across Data and ML workloads Own data and model health with monitoring, validation, and drift detection
You will design and implement large-scale, secure, and highly available systems, while collaborating across teams to drive the future of secure scalable inference platform. The mindset required and to be developed is how to process thousands of transactions per second, how to achieve the consistency without sacrificing the performance. Work with cross functional teams to drive requirements, size scope and effort, mentor junior engineers, lead the project to completion and provide support for any Production issues.
Design and develop scalable backend services using Java and distributed system principles. Build and maintain high-throughput APIs, batch processing pipelines, and Mass Action services for fraud operations and NPI events. Develop new platform capabilities for ACM Next Gen, Sky Vault, and AML Services. Lead the engineering work required for infrastructure readiness and migration on Alibaba Cloud, including: ◦ Building ACM infrastructure on Alibaba Cloud. ◦ Supporting Kubernetes, OpenSearch, Solr, Kafka, and other platform components. ◦ Migrating AWS-based services to Alibaba Cloud services including OSS, ACK, and ApsaraMQ. ◦ Replacing AWS SDKs with Alibaba Cloud SDKs. ◦ Implementing China-specific data routing and ingestion logic. ◦ Updating deployment automation, CI/CD pipelines, and load testing for the Alibaba Cloud environment. Design reliable, scalable, and observable distributed systems with strong operational excellence. Partner with Product, Architecture, and partner engineering teams to define technical solutions and platform direction. Use AI tools and modern engineering practices to improve development productivity, code quality, automation, and platform innovation. Participate in production support, debugging, performance tuning, and operational improvements.