西门子Principal Engineer, AI and Machine Learning Systems
任职要求
The role Siemens builds the systems the physical world runs on: factories, power grids, buildings, trains, hospitals. Industrial and physical AI is a major opportunity in applied AI, and one of the harder ones to get right. There is a generation of AI-powered products to build. We are forming engineering pods in China to build them. As Principal Engineer, you own the technical vision and system architecture for the AI-powered platforms the pod ships. You take evolving product and research requirements and turn them into systems that run reliably, scale economically, and stay maintainable as the work grows. This is a senior individual contributor role for a deeply technical engineer who thrives in ambiguity. You are hands on. Your primary impact is through architectural leadership, technical judgment, and raising the engineering bar across the team. Key responsibilities • Define the technical vision and system architecture for AI-powered platforms and products in the pod • Convert evolving product and research requirements into scalable, reliable ML systems • Partner with the Senior Principal Product Manager to align technical decisions with product strategy • Partner with the Senior Principal Applied Scientist and Principal Scientists to take models from experimentation into production grade systems • Own architectural decisions across model training, inference, data pipelines, and system integration • Identify critical risks early in performance, scalability, cost, and reliability, and…
工作职责
N/A
THE ROLE: AMD is looking for a strategic software engineering lead to drive next-generation AI inference systems, intelligent model routing, and cloud-native deployment technologies for AMD Instinct GPUs. In this role, you will work at the intersection of LLM serving, semantic routing, Kubernetes, Envoy, AI gateways, and open-source infrastructure. You will be a member of a core team of talented industry specialists focused on enabling high-performance, production-ready AI software on the latest AMD hardware and ROCm software stack. This role is especially focused on advancing intelligent routing and system-level optimization for LLM inference, including vLLM, vLLM Semantic Router, multi-model serving, policy-driven routing, semantic caching, observability, privacy-aware routing, and workload-aware optimization across AMD GPU platforms. THE PERSON: The ideal candidate is a hands-on technical leader with deep expertise in cloud-native infrastructure, open-source development, and AI inference systems. The candidate should be passionate about building scalable systems from 0 to 1, driving open-source communities, and solving complex performance, reliability, and deployment challenges. The candidate should be comfortable working across engineering, architecture, product, partner, and open-source communities. Strong communication, technical writing, public speaking, and community leadership skills are essential. The successful candidate will be able to translate emerging AI infrastructure trends into practical software solutions that strengthen AMD’s position in the open-source AI ecosystem. KEY RESPONSIBILITIES: Lead the design and development of intelligent routing technologies for LLM serving on AMD Instinct GPUs, including semantic routing, workload-aware routing, policy-based routing, and multi-model inference orchestration.Drive AMD enablement and optimization for vLLM Semantic Router and related open-source AI gateway technologies, ensuring strong support for ROCm and AMD GPU platforms.Collaborate with AMD architecture, ROCm, kernel, compiler, and AI framework teams to identify and optimize bottlenecks in LLM inference workloads.Develop production-quality software components for AI inference systems, including routers, gateways, control-plane services, observability tools, policy engines, and deployment automation.Build and optimize integrations across vLLM, Kubernetes, Envoy, Gateway API, service mesh, and AI gateway ecosystems.Apply a data-driven approach to performance analysis, including benchmarking, profiling, latency analysis, throughput optimization, and cost-efficiency evaluation.Contribute to open-source communities and represent AMD in key AI infrastructure projects, including vLLM, Kubernetes, Envoy, Gateway API, and related CNCF ecosystems.Develop technical relationships with external partners, customers, researchers, and community maintainers to accelerate AMD adoption in AI inference workloads.Create technical documentation, blogs, demos, reference architectures, and conference presentations to showcase AMD’s AI software capabilities.Participate in new AMD GPU platform bring-up activities by validating AI inference software stacks, debugging system-level issues, and developing early proof points for emerging workloads.Research and prototype new approaches for system intelligence in LLM serving, including semantic caching, prompt classification, privacy-aware routing, safety-aware routing, tool routing, agent routing, and workload-router-pool architectures.PREFERRED EXPERIENCE: Strong software engineering background with experience building distributed systems, cloud-native infrastructure, AI infrastructure, or high-performance serving systems.Deep experience with Kubernetes, Envoy, Gateway API, service mesh, ingress/gateway controllers, or cloud-native networking.Experience contributing to or maintaining major open-source projects, preferably in CNCF, Kubernetes, Envoy, Istio, vLLM, or AI infrastructure communities.Experience with LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or related inference-serving systems.Experience with semantic routing, AI gateways, model routing, policy-based routing, semantic caching, prompt classification, or multi-model inference orchestration.Strong programming skills in Go, Rust, Python, and/or C/C++. Go and Rust experience are especially valuable for cloud-native control planes, gateways, and high-performance routing systems.Experience with Linux systems, containerized deployments, distributed debugging, observability, and production reliability.Familiarity with GPU-accelerated AI workloads, ROCm, CUDA, ONNX Runtime, PyTorch, or inference performance optimization is a strong plus.Experience with performance profiling, benchmarking, latency optimization, memory optimization, and high-concurrency serving systems.Ability to write high-quality, maintainable code with strong attention to architecture, reliability, testing, and operational simplicity.Strong technical communication skills, including technical writing, public speaking, community engagement, and cross-functional collaboration.Demonstrated ability to lead complex technical projects from concept to production and influence without direct authority across organizations and open-source communities.Motivating technical leader with excellent interpersonal skills and the ability to work effectively in global, distributed teams.ACADEMIC CREDENTIALS: Bachelor’s or Master’s degree in Computer Science, Software Engineering, Computer Engineering, Electrical Engineering, or equivalent experience.Advanced degree or research experience in AI systems, distributed systems, cloud-native infrastructure, machine learning systems, or high-performance computing is a plus.#LI-JW2
Key Responsibilities: Contribution: • Strong hands-on capability of migration, integration and cloud project implementation in specific domain. push the consumption growth objective. • Identifying customer needs in specific domain, and driving adoption and expansion. Lead product implementation. optimization, troubleshooting. • Model accountability in technical solution optimization, and ownership of both successes and failures that impact the technical solutions. Challenge: • Handle customer's issues using multiple approaches to identify optimal technical solution in specific domain. • Utilize Oracle and third-party data to gain a strategic, forward-thinking perspective on technical solutions. • Stay informed about evolving technical and business trends that may impact current or future projects. Expertise: • Apply deep knowledge of Oracle Cloud technologies, industry verticals, and market trends to design tailored solutions. • Continuously deepen your product and industry expertise, educating others and leveraging this knowledge to deliver impactful solutions.
Contribution: • Strong hands-on capability of migration, integration and cloud project implementation in specific domain. push the consumption growth objective. • Identifying customer needs in specific domain, and driving adoption and expansion. Lead product implementation. optimization, troubleshooting. • Model accountability in technical solution optimization, and ownership of both successes and failures that impact the technical solutions. Challenge: • Handle customer's issues using multiple approaches to identify optimal technical solution in specific domain. • Utilize Oracle and third-party data to gain a strategic, forward-thinking perspective on technical solutions. • Stay informed about evolving technical and business trends that may impact current or future projects. Expertise: • Apply deep knowledge of Oracle Cloud technologies, industry verticals, and market trends to design tailored solutions. • Continuously deepen your product and industry expertise, educating others and leveraging this knowledge to deliver impactful solutions.
• Act as the craft owner for LLM engineering at Supercell — setting direction, sharing best practices, and raising the bar for what “great” looks like in LLM-powered systems. • Drive adoption of a wide range of LLM and agentic applications (e.g. in-game bots, player support, social insights, internal productivity tools). • Champion AI development tools as a particularly high-leverage use case — helping our teams accelerate workflows, code smarter, and build better systems faster. • Partner closely with game and functional teams to evangelize LLM capabilities and turn vague opportunities into concrete solutions. • Serve as a bridge between technical and non-technical stakeholders, clearly communicating the strengths, limitations, and business value of different approaches. • Monitor emerging developments in the LLM and agents space; assess technologies for safety, reliability, and performance; and build our internal “stack” of recommended approaches. • Lead the design and implementation of robust evaluation, monitoring, and safety frameworks for production LLM-powered systems. • Rapidly prototype and deliver production-grade systems — but also coach, mentor, and enable other engineers to build confidently themselves.