英伟达Senior Software Engineer, Automation Infrastructure
任职要求
NVIDIA's Performance Lab (PerfLab) builds the systems and automation used to evaluate the performance and quality of accelerated computing and AI workloads. We turn complex benchmark experiments into reliable, scalable, and reproducible workflows that help engineering teams make better decisions faster.We are looking for an experienced and highly self-motivated System Software Engineer to help build the next generation of PerfLab's benchmark infrastructure. You will independently own meaningful platform components and take projects from problem discovery and technical design through production deployment and adoption. The ideal candidate enjoys finding important engineering problems, understanding their root causes, and using technology to create simple, reusable solutions.You will collaborate with NVIDIA teams around the world and work on evolving areas such as large language models, agentic AI, accelerated computing, and other emerging AI workloads. What You'll Be Doing: • Design, build, and maintain reusable software, services, and workflows that automate benchmark definition, execution, result collection, validation, and reporting across local, cluster, and cloud-native environments. • Carry out performance testing and analysis as needed. Develop a deep understanding of existing performance workflows and infrastructure, and translate real-world needs into scalable, user-friendly solutions. • Improve the reliability, scalability, observability, and reproducibility of large benchmark campaigns. • Diagnose complex issues across applications, Linux systems, containers, distributed jobs, compute resources, networking, and storage. • Build strong partnerships with performance engineers, QA teams, product teams, and other customers; agree on goals, organize execution, and drive new ideas and projects from concept through adoption. • Contribute to technical designs, code reviews, documentation, and internal or open-source infrastructure projects. • Apply AI-assisted automation where it can meaningfully improve benchmark creation, failure triage, data analysis, or engineering productivity. What We Need to See: • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience in practice. • 5+ years of relevant software engine…
工作职责
N/A
• Lead code reviews to ensure adherence to engineering standards, test coverage, and secure coding practices. Provide feedback and mentorship to peers, and apply tools and patterns that enhance reliability, diagnosability, and maintainability. • Develop scalable and secure design proposals, collaborating across teams to resolve dependencies and validate design hypotheses. Ensure solutions meet performance, compliance, and cost expectations. • Create and maintain test plans that validate functionality and security. Leverage automation and AI tools to improve test reliability and coverage, and ensure testability is embedded in design. • Apply secure design principles and engineering best practices to build resilient systems. Drive automation in deployment, ensure compliance with global regulations, and integrate security monitoring and incident response mechanisms. • Translate product requirements into actionable plans, estimate effort, and guide execution. Ensure safe deployment practices, flighting strategies, and rollback plans are in place to support efficient and secure releases. • Integrate telemetry and observability into systems to monitor performance and security. Act as a Designated Responsible Individual (DRI), lead incident response efforts, and continuously improve live site operations and support documentation. • Collaborate with stakeholders to understand user needs and incorporate feedback into product design. Ensure privacy and security requirements are met, and establish feedback loops to measure impact and value.
• Lead on IC3 fundamental initiatives including but not limited to SDP, resiliency, and architectural movements. • Provide technical guidance and expertise to the development team, driving the overall software design and implementation of complex, large-scale software solutions. • Research and evaluate emerging technologies, tools, and methodologies, making recommendations for adoption and integration into existing systems and processes. • Communicate effectively with both technical and non-technical stakeholders, presenting complex technical concepts in a clear and concise manner. • Drive to make tough technical decisions in face of great ambiguity to unblock team with clear course of action. • Model Microsoft core values and culture of Growth Mindset, Diversity & Inclusion and One Microsoft.
• Learn to review and break down work items into tasks with stakeholder collaboration, provide estimations, and escalate delays, while also supporting feature deployments to customers, considering user and service impacts, and adhering to best deployment practices for safety. • Collaborate with key stakeholders to define feature requirements, integrate feedback to enhance design, and establish feedback loops for continuous improvement based on customer metrics. • Learn and apply coding standards and best practices through code reviews, developing maintainable and extensible code with guidance. Utilize debugging tools to proactively and reactively address issues in product features, ensuring code quality and reliability. • Support the identification of dependencies and design documentation for product features, learn about system interactions and back-end dependencies, and contribute to architectural processes under guidance. Produce code to test hypotheses for technical solutions and assist with technical validation efforts. Collaborate on quality assurance plans, augment test cases, and integrate automation into testing, while understanding the implications of security and compliance in system architecture. • Contribute to data analysis and feedback integration for product engineering decisions, acting as a Designated Responsible Individual (DRI) for monitoring and restoring system functionality within Service Level Agreement (SLA) timeframe. Participate in live service operations, and support telemetry data integration for system behavior insights, with a focus on performance, reliability, and safety. • Develop and apply best practices for reliable code building, understand global and local regulations, customer scaling requirements, and support communication with key partners across Microsoft for user experience enhancement and partner needs. • Ensure compliance with security, privacy, safety, and accessibility standards, leverage developer tools for code creation and debugging, contribute to automation in production and deployment, and proactively seek knowledge to improve product availability, reliability, efficiency, and performance at scale.
- Keep up to date with and utilize the latest developments in LLM system optimization.- Discover/solve impactful technical problems, advance state-of-the-art LLM technologies, and translate ideas into production.- Optimize LLM inference workloads through innovative kernel, algorithm, scheduling, and parallelization technologies.- Continuously maintain internal LLM inference infrastructure.