英伟达Infrastructure Software Engineer, Deep Learning Libraries
任职要求
• BS or equivalent experience or higher degree in Computer Science or Computer Engineering • 2+ years of relevant experience • Strong programming skills in Python (or similar) and familiarity with C/C++ development • Experience setting up, maintaining, and automating continuous integration systems (e.g. Jenkins) • Fluency in SCM (e.g. Git, Perforce) and build systems (e.g. Make, CMake, Bazel) • A pragmatic approach to solving problems and collaboration • Passion for "it just works" automation and enabling team members Ways to stand out from the crowd: • Experience designing and developing automation in Jenk…
工作职责
• Designing and developing software for testing and analysis of our codebases • Building scalable automation for build, test, integration, and release processes for publicly distributed deep learning libraries • Developing throughout the software stack, from the user experience down to the cluster and database layers • Configuring, maintaining, and building upon deployments of industry-standard tools (e.g. Kubernetes, Slurm, Jenkins, Docker, CMake, Github, Gitlab, Jira, etc) • Advancing state of the art in those industry-standard tools
• Designing and developing software for testing and analysis of our codebases • Building scalable automation for build, test, integration, and release processes for publicly distributed deep learning libraries • Developing throughout the software stack, from the user experience down to the cluster and database layers • Configuring, maintaining, and building upon deployments of industry-standard tools (e.g. Kubernetes, Jenkins, Docker, CMake, Github, Gitlab, Jira, etc) • Advancing state of the art in those industry-standard tools
• Create and implement the training infrastructure spanning pre-training, SFT, and RL post-training for Physical AI world foundation models. The work involves the framework and a comprehensive control plane across clusters to coordinate workloads efficiently. • Develop and improve the pre-training and SFT pipelines — large-scale data loading, distributed training, and checkpointing — to achieve high throughput and scalability. • Develop and improve the inference and evaluation stack, including the inference engine, inference/generation pipelines (which also support RL rollout), and evaluation pipelines. Use methods like continuous batching and KV-cache management to achieve high throughput and low latency. • Build and improve the effective interaction and data flow among the RL system's roles (policy, rollout, reward, simulation) while investigating system-level optimization opportunities. • Integrate and orchestrate simulation and robotics environments as RL environments — driving the simulation↔rollout↔training loop at scale. • Build and refine the distributed training backend — sharding/parallelism, mixed precision, activation checkpointing, and memory/throughput optimization across many GPUs. • Improve the efficiency, scalability, and resiliency of training and RL workloads — focusing on fault tolerance, fast/elastic restart, and throughput optimization under preemption and hardware failure. • Define meaningful, actionable reliability and efficiency metrics to track and improve system reliability. • Root cause, triage, and resolve failures from the application level down to the framework, GPU, and network/hardware level.
• Architect, develop, and maintain Python-based tools and services to efficiently run a performance-focused multi-tenant Linux cluster including embedded, desktop, and server systems • Work with industry standard tools (Kubernetes, Slurm, Ansible, Gitlab, Artifactory, Jira) • Actively support users doing development, functional testing, and performance testing on current and pre-production GPU cluster systems • Work with various teams at NVIDIA across different timezones to incorporate and influence the latest tools for operating GPU clusters • Collaborate with users and system administrators to seek out ways to improve UX and operational efficiency • Become an expert on the entire AI infrastructure stack