logo of microsoft

微软Principal Software Engineering Manager- GPU Inference Optimization

社招全职Software Engineering地点:北京状态:招聘

任职要求


• Bachelor's degree in computer science or related technical field AND 5+ years technical engineering experience with coding in languages including, but not limited to, C/C++, CUDA, ROCm or equivalent experience

• Practical Experience writing new GPU kernels, going beyond experience of GPU workloads with existing library kernels

• Quick learning, good communication (fluent in English) and solid problem-solving skills

• Cross-team collaboration skills and the desire to collaborate in a team of researchers and developers

• Experience in low-level performance analysis and optimization, including proficiency using GPU profiling tools such as NVIDIA Visual Profiler, and NVIDIA Nsight Compute is a plus

• Familiar with LLM inference optimization, experience in developing popular inference framework such as TensorRT-LLM, SGLang, …
登录查看完整任职要求
微信扫码,1秒登录

工作职责


• Lead the software development in C/C++, Python, and in GPU languages such as CUDA, ROCm, or Triton• Analyze metrics and identify opportunities based on offline and online testing, develop and deliver robust and scalable solutions.• Work with cutting-edge hardware stacks and a fast-moving software stack to deliver best-of-class inference and optimal cost.• Engage with key partners to understand and implement inference and training optimization for state-of-the-art LLMs and other models.
包括英文材料
C+
C+++
CUDA+
NVIDIA Visual Profiler+
Nsight+
还有更多 •••
相关职位

logo of amd
社招 Enginee

THE ROLE: The mission of the Principal Technical Lead is to orchestrate and elevate the quality, consistency, and competitiveness of AMD's GPU software ecosystem on Linux. This leader will bridge strategic objectives with technical execution across the ROCm stack and Linux driver portfolios (both packaged and inbox), ensuring a seamless, powerful, and reliable experience for developers, researchers, and enterprises choosing AMD for their accelerated computing needs.   KEY RESPONSIBILITIES: Strategic Technical Leadership & SOW Definition Act as the central technical nexus between Product Management, Software Architecture, and engineering teams (kernel, ROCm, QA, support). Translate high-level product goals and market requirements into detailed, actionable, and prioritized Technical Statements of Work (SOWs) for RSL AI validation team ensure validation plans are coherent, dependencies are managed, and resources are aligned to deliver on strategic commitments for both Radeon and Ryzen AI solutions. Quality, Test & Process Optimization: Own the definition and evolution of the product quality bar for AMD's Linux GPU software. · Champion and drive the implementation of a robust, scalable, and automated CI/CD and test infrastructure across Native Linux, WSL, and various hardware platforms. Establish key performance indicators (KPIs) for software quality, release velocity, and regression rates. Use data to drive continuous improvement in development and testing efficiency Unified User Experience & Competitive Analysis: Define and monitor a holistic user experience (UX) scorecard encompassing installation, performance predictability, documentation, and debugging. Institute a formal, ongoing competitive analysis framework to benchmark the AMD software stack (ROCm + Drivers) against key competitors across performance, feature parity, stability, and usability. Serve as the ultimate internal advocate for the end-user, ensuring customer and community feedback is systematically integrated into the development lifecycle. Linux Ecosystem & Driver Consistency: Provide technical guidance and oversight to ensure flawless synchronization between the AMD packaged driver and the upstream Linux kernel (inbox) driver. Strengthen AMD's partnership with the Linux kernel community and major distributions (e.g., Canonical, Red Hat, SUSE). Drives a consistent and high-quality user experience regardless of the driver delivery channel (OS vendor vs. AMD.com).

更新于 2025-09-24上海
logo of oracle
社招PRODEV-S

Responsibilities Collaborate with GPU sales team and SCE AIML TPM team to provide technical support for customers both at pre-sales and after-sales stage. Take ownership of problems and work to identify solutions. Design, deploy, and manage infrastructure components such as cloud resources, distributed computing systems, and data storage solutions to support AI/ML workflows. Collaborate with customers’ scientists and software/infrastructure engineers to understand infrastructure requirements for training, testing, and deploying machine learning models. Implement automation solutions for provisioning, configuring, and monitoring AI/ML infrastructure to streamline operations and enhance productivity. Optimize infrastructure performance by tuning parameters, optimizing resource utilization, and implementing caching and data pre-processing techniques. Troubleshoot infrastructure performance, scalability, and reliability issues and implement solutions to mitigate risks and minimize downtime. Stay updated on emerging technologies and best practices in AI/ML infrastructure and evaluate their potential impact on our systems and workflows. Document infrastructure designs, configurations, and procedures to facilitate knowledge sharing and ensure maintainability. Qualifications: Experience in scripting and automation using tools like Ansible, Terraform, and/or Kubernetes. Experience with containerization technologies (e.g., Docker, Kubernetes) and orchestration tools for managing distributed systems. Solid understanding of networking concepts, security principles, and best practices. Excellent problem-solving skills, with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment. Strong communication and collaboration skills, with the ability to work effectively in cross-functional teams and convey technical concepts to non-technical stakeholders. Strong documentation skills with experience documenting infrastructure designs, configurations, procedures, and troubleshooting steps to facilitate knowledge sharing, ensure maintainability, and enhance team collaboration. Strong Linux skills with hands-on experience in Oracle Linux/RHEL/CentOS, Ubuntu, and Debian distributions, including system administration, package management, shell scripting, and performance optimization.

更新于 2025-12-02
logo of microsoft
社招Research

• Initiate and advance research to advance state-of-the-art in AI for Software Engineering • Collaborate across disciplines with product teams across Microsoft and Github • Stay up to date with the research literature and product advances in AI for software engineering • Collaborate with world renowned experts in programming tools and developer tools to integrate AI across software development stack for Copilot • Build and manage large-scale AI experiments and models.

更新于 2025-10-22上海
logo of microsoft
社招Software

• Lead the technical direction and vision for the architecture, design, and the implementation of our infrastructure on Azure that scales to provision, manage, and monitor health of millions of cloud-based virtual devices.  • Mentor and help grow a team of talented, diverse software engineers.  • Work across organizations, collaborating with internal partner teams such as Azure Compute, Core OS, Microsoft Security and Identity team, and others.  • Raise the technical bar, maintain a data and results driven culture, and nurture a high-performance team to build world-class experiences for W365 ITPros, partners, and operations teams.  • Get to extend your knowledge of cloud computing, hypervisors, desktop virtualization, streaming technologies, and other technical areas including cloud-based management suites.  • Be part of a team designing fundamental capabilities involving device management, computing, storage, networking, and streaming protocols (such as Remote Desktop Protocol) for our core products to enhance the value to our customer base.  • Be a part of an agile team working with experienced engineers and product managers that behave more like a technology startup.

更新于 2025-10-28苏州