英伟达CPU Performance Developer Technology Engineer
任职要求
• BS, MS, or PhD in Computer Science, Computer Engineering, or a related field. • 5+ years of relevant experience in performance engineering or CPU optimization. • Strong programming proficiency in C/C++ and/or Python, with a deep understanding of algorithms and software architecture. • Solid grasp of CPU microarchitecture, performance analysis tools, and optimization methodologies. • Proven track record of CPU benchmarking and bottleneck-driven performance tuning. • Excellent communication and organizational skills, wit…
工作职责
• Collaborate with developers, researchers, and framework maintainers across industries to identify and resolve performance challenges in diverse workloads such as AI, data analytics, simulation, and numerical computing. • Profile, analyze, and optimize CPU performance from application-level algorithms down to low-level microarchitecture. • Contribute to open-source frameworks, key software stacks, reference implementations, and performance libraries to unlock full CPU potential. • Work closely with NVIDIA’s architecture, research, libraries, tools, and system software teams to improve our overall platform performance. • Provide insights that shape next-generation CPU designs, compiler toolchains, and development workflows for better developer productivity and throughput.
• Collaborate with developers, researchers, and framework maintainers across industries to identify and resolve performance challenges in diverse workloads such as AI, data analytics, simulation, and numerical computing. • Profile, analyze, and optimize CPU performance from application-level algorithms down to low-level microarchitecture. • Contribute to open-source frameworks, key software stacks, reference implementations, and performance libraries to unlock full CPU potential. • Work closely with NVIDIA’s architecture, research, libraries, tools, and system software teams to improve our overall platform performance. • Provide insights that shape next-generation CPU designs, compiler toolchains, and development workflows for better developer productivity and throughput.
• Collaborate with developers, researchers, and framework maintainers across industries to identify and resolve performance challenges in diverse workloads such as AI, data analytics, simulation, and numerical computing. • Profile, analyze, and optimize CPU performance from application-level algorithms down to low-level microarchitecture. • Contribute to open-source frameworks, key software stacks, reference implementations, and performance libraries to unlock full CPU potential. • Work closely with NVIDIA’s architecture, research, libraries, tools, and system software teams to improve our overall platform performance. • Provide insights that shape next-generation CPU designs, compiler toolchains, and development workflows for better developer productivity and throughput.

1.参与自动驾驶CPU(高性能AP核)的需求和规格的定义与分析; 2.完成CPU子系统的交付,包括RTL集成、时钟/复位设计、电源域划分、低功耗流程、静态时序分析与物理协同,支持验证团队测试并协助后端完成物理实现; 3.为CPU计算子系统打造产品竞争力,包括SOC场景需求分析、微架构及方案制定、性能和成本分析、时序面积功耗优化等工作; 4.支持系统级验证与硅后调试,完成量产问题跟踪、良率提升等相关工作;