logo of nvidia

随便看看「英伟达」有没有自己喜欢的职位~

社招

• Analyze state of the art DL networks (LLM etc.), identify and prototype performance opportunities to influence SW and Architecture team for NVIDIA's current and next gen inference products • Develop analytical models for the state of the art deep learning networks and algorithm to innovate processor and system architectures design for performance and efficiency. • Specify hardware/software configurations and metrics to analyze performance, power, and accuracy in existing and future uni-processor and multiprocessor configurations. • Collaborate across the company to guide the direction of next-gen deep learning HW/SW by working with architecture, software, and product teams.

更新于 2026-07-17上海|北京
社招

We are looking for Engineers to join NVIDIA's network R&D team. NVIDIA is a fast-growing company with positive energy that emanates from our team members' internal aim to develop, market, sell, and support cutting-edge products and services. We are a strong believer in developing our employees and giving them the tools to succeed. The work environment is versatile, educational, dynamic and challenging as our employees are currently working on innovative, next-generation networking devices at the forefront of technology in terms of performance efficiency. The daily work involves all aspects of firmware development: Design, Micro- Architecture, Software interfaces and Verification. If you are looking for a rewarding career, hardworking colleagues, and a great environment in which to challenge yourself, grow, and lead, NVIDIA is the right place for you. What you’ll be doing: • Engage in next-generation SmartNIC firmware develop • Validate the functionalities of firmware and silicon design before tape-out, bring up new silicon when tape-out is done • Implement new features for congestion control, virtualization, QoS, etc. • Support local customers to resolve issues in deployment • Supply creative ideas to improve products and efficiency

更新于 2026-07-17北京|上海
社招

• Design and implement functional/performance tests for CUDA products, like driver and library. • Automate CUDA tests, design test plans and integrate into automation testing infrastructure. • Triage test results, root cause test failures or performance drops, and drive through bugs to fix. • Develop scripts/tools and optimize workflow to improve efficiency and productivity.

更新于 2026-07-17上海
社招

N/A

更新于 2026-07-16上海
社招

• Use your Design/Verification experience to define, model, and prove key design behaviors using formal methods. • Work closely with design, verification, and architecture teams to find bugs early and improve design quality. • Build expertise in advanced formal tools, properties, abstractions, and debug methods with mentorship from an experienced FV team. • Take part in the AI revolution, working on cutting-edge silicon architectures.

更新于 2026-07-16上海|北京
社招

• Understand the Switch architecture and data flows • Refine the full Chip working flow to improve the entire team's efficiency & smooth the project exeuction • Work closely with multiple teams within organizations such as Architecture, Micro- Architecture, and FW to meet project milestones.

更新于 2026-07-15上海
社招

• Working directly with key application developers to understand the current and future problems they are solving. You will build and optimize core parallel algorithms and data structures to deliver the most effective solutions using GPUs, through both library development and direct contribution to applications. This includes training and inference optimization for large language models (LLM), contributing to frameworks and open-source projects in the large language models ecosystem, such as Megatron and TRTLLM, SGLang, vLLM... • Collaborating closely with the architecture, research, libraries, tools, and system software teams at NVIDIA to influence the build of next-generation architectures, software platforms, and programming models. This includes investigating impact on application performance and developer efficiency, and turning real-world developer feedback into actionable platform improvements. • Engaging in deep optimization of high-performance operators, involving but not limited to GPU kernel optimization, instruction-level tuning, and compiler optimization. These optimizations will directly support customers or be coordinated within computation libraries and open-source projects across the community, like cuDNN, cuBLAS, and CUTLASS and Open- source libs like DeepGEMM, FlashMLA, FlashAttention, Flashinfer... • Improving communication for broad distributed large language models workloads. You will spearhead advancements in distributed training and inference by refining communication libraries(NCCL,NCCL GIN , NVSHMEM) and engaging in open-source communication libraries(like DeepEP, NCCL EP). This demands in-depth study of interconnect topologies(NVLINK) and network protocols(InfiniBand/RoCE) to design efficient data transfer strategies and methods for compute-communication overlap.

更新于 2026-07-15上海|北京|深圳
社招

• Work in a combined design and verification team specializing in Switch Fullchip works • Understand the Switch architecture and build on testplan accordingly • Maintain and optimize Fullchip verification enviornment to meet feature requirements efficiently • Develop Fullchip test suites, maintain regressions, debug failures and sign-off coverages

更新于 2026-07-15上海