vivoGPU系统工程师-27届暑期实习
任职要求
1、硕士及以上学历,计算机体系结构、计算机图形、微电子专业等相关专业;
2、在计算机图形学、体系结构或微电子方向理论基础扎实;具有较好的工程和编码能力;同时具备研究与创新能力;
3、优选加分项:
(1)熟悉游戏…工作职责
本岗位面向硕士及以上毕业生,旨在培养具备计算机图形学、GPU性能分析、GPU微架构设计能力与产业落地视野的高潜人才,职责涵盖图形渲染算法研究、GPU软硬件协同设计、GPU架构预研与前瞻技术探索,具体如下: 1. 针对移动端游戏的高性能渲染与能效需求,基于移动端GPU的设计特点对游戏渲染过程进行系统性分析,结合GPU软、硬件架构与游戏负载建模,识别游戏软、硬件瓶颈并制定针对性优化策略,为移动端GPU的设计迭代及游戏优化提供量化评估模型与数据支撑; 2. 面向下一代图形渲染技术演进(如实时路径追踪、神经辐射场渲染等),构建场景驱动的GPU能效评估体系,通过量化分析当前GPU软、硬件在典型负载下的PPA表现,跟踪业界最新的图形渲染技术和架构演进方向,推导出面向未来游戏场景的渲染技术和相应的GPU架构演进路径。为与供应商合作提供场景和数据支撑。
工作职责 1、万卡级 GPU 调度系统建设: 参与大规模 GPU 集群调度系统建设,围绕 Quota、优先级、抢占、弹性伸缩、碎片整理、拓扑感知调度等能力提升资源效率。 2、训推统一调度: 面向大模型训练、后训练、推理服务等不同负载,设计训推统一调度、潮汐混部、在线离线协同和资源弹性策略。 3、资源利用率治理: 建设 GPU 资源利用率分析体系,基于真实负载数据识别低效资源、资源碎片、潮汐空闲和调度瓶颈。 4、LLMOps 平台融合: 参与构建面向大模型训练、微调、推理、部署全流程的 LLMOps 能力,与云原生平台深度融合,支撑大模型生产链路稳定高效落地。 5、集群稳定性建设: 与云原生、IDC、网络、存储和业务团队协作,提升大规模 AI 集群的故障恢复能力、资源周转效率和任务稳定性。 6、前沿技术探索: 持续关注 Kubernetes、Volcano、Kueue、Ray、GPU 虚拟化、弹性调度等相关技术,探索下一代 AI 资源调度系统。
1、负责AI硬件的云端GPU和XPU的推理引擎日常研发; 2、跟进支撑智能业务,整合团队资源,协调外部团队,推动业务上线; 3、探索计算优化和芯片异构前沿技术,比如:GPU,XPU等,跟进前沿加速技术:flash attention,paged attention,MTP。 职位要求
1、万卡级 GPU 调度系统建设: 参与大规模 GPU 集群调度系统建设,围绕 Quota、优先级、抢占、弹性伸缩、碎片整理、拓扑感知调度等能力提升资源效率。 2、训推统一调度: 面向大模型训练、后训练、推理服务等不同负载,设计训推统一调度、潮汐混部、在线离线协同和资源弹性策略。 3、资源利用率治理: 建设 GPU 资源利用率分析体系,基于真实负载数据识别低效资源、资源碎片、潮汐空闲和调度瓶颈。 4、LLMOps 平台融合: 参与构建面向大模型训练、微调、推理、部署全流程的 LLMOps 能力,与云原生平台深度融合,支撑大模型生产链路稳定高效落地。 5、集群稳定性建设: 与云原生、IDC、网络、存储和业务团队协作,提升大规模 AI 集群的故障恢复能力、资源周转效率和任务稳定性。 6、前沿技术探索: 持续关注 Kubernetes、Volcano、Kueue、Ray、GPU 虚拟化、弹性调度等相关技术,探索下一代 AI 资源调度系统。
THE ROLE: As a core member of the team, you will play a pivotal role in optimizing and developing deep learning frameworks for AMD GPUs. Your strong experience will be critical in enhancing GPU kernels, deep learning models, and training/inference performance across multi-GPU and multi-node systems. You will engage with both internal GPU library teams and open-source maintainers to ensure seamless integration of optimizations, utilizing cutting-edge compiler technologies and advanced engineering principles to drive continuous improvement. THE PERSON: Highly skilled engineer with strong technical and analytical expertise in C++ development within Linux environments. The ideal candidate will thrive in both collaborative team settings and independent work, with the ability to define goals, manage development efforts, and deliver high-quality solutions. Strong problem-solving skills, a proactive approach, and a keen understanding of software engineering best practices are essential. KEY RESPONSIBILITIES: Optimize Deep Learning Frameworks: Enhance and optimize frameworks like TensorFlow and PyTorch for AMD GPUs in open-source repositories. Develop GPU Kernels: Create and optimize GPU kernels to maximize performance for specific AI operations. Develop & Optimize Models: Design and optimize deep learning models specifically for AMD GPU performance. Collaborate with GPU Library Teams: Work closely with internal teams to analyze and improve training and inference performance on AMD GPUs. Collaborate with Open-Source Maintainers: Engage with framework maintainers to ensure code changes are aligned with requirements and integrated upstream. Work in Distributed Computing Environments: Optimize deep learning performance on both scale-up (multi-GPU) and scale-out (multi-node) systems. Utilize Cutting-Edge Compiler Tech: Leverage advanced compiler technologies to improve deep learning performance. Optimize Deep Learning Pipeline: Enhance the full pipeline, including integrating graph compilers. Software Engineering Best Practices: Apply sound engineering principles to ensure robust, maintainable solutions. Mentor and Guide: Provide mentorship to junior team members, fostering growth and collaboration through code reviews, knowledge sharing, and technical guidance. PREFERRED EXPERIENCE: GPU Kernel Development & Optimization: Strong experience in designing and optimizing GPU kernels for deep learning on AMD GPUs using HIP, CUDA, and assembly (ASM). Strong knowledge of AMD architectures (GCN, RDNA) and low-level programming to maximize performance for AI operations, leveraging tools like Compute Kernel (CK), CUTLASS, and Triton for multi-GPU and multi-platform performance. Deep Learning Integration: Strong experience in integrating optimized GPU performance into machine learning frameworks (e.g., TensorFlow, PyTorch) to accelerate model training and inference, with a focus on scaling and throughput. Software Engineering: Expert skills in Python and C++, with experience in debugging, performance tuning, and test design to ensure high-quality, maintainable software solutions. High-Performance Computing: Strong experience in running large-scale workloads on heterogeneous compute clusters, optimizing for efficiency and scalability. Compiler Optimization: Sound understanding of compiler theory and tools like LLVM and ROCm for kernel and system performance optimization. ACADEMIC CREDENTIALS: Master’s and/or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field. 5+ years of professional experience in technical software development, with a focus on GPU optimization, performance engineering, and framework development. #LI-FL1