蚂蚁金服算子开发和优化
任职要求
1. 熟悉常见的CPU和GPU的架构和微架构; 2. 熟悉核函数性能分析和优化…
工作职责
1. 分析AI算子性能瓶颈,优化算子性能; 2. 面向模型结构和训推设置设计和实现高性能算子; 3. 算子测试和验证并完成算子与训推框架的集成。
An exciting internship opportunity to make an immediate contribution to AMD's next generation of technology innovations awaits you! We have a multifaceted, high-energy work environment filled with a diverse group of employees, and we provide outstanding opportunities for developing your career. During your internship, our programs provide the opportunity to collaborate with AMD leaders, receive one-on-one mentorship, attend amazing networking events, and much more. Being part of AMD means receiving hands-on experience that will give you a competitive edge. Together We Advance your career! JOB DETAILS: Location: Shanghai, China Onsite/Hybrid: This role require the student to work at least 3 days/week, either in a hybrid (minimum 3 Days in Office) or onsite work structure throughout the duration of the co-op/intern term. Duration: Jan - June 2026 WHAT YOU WILL BE DOING: We are seeking a highly motivated Machine Learning (ML)/Artificial Intelligence (AI) intern/co-op to join our team and contribute to the development of next-generation product differentiation features alongside expert ML/AI engineers. In this role, you will: Gain hands-on experience with cutting-edge technologies in ML, AI, and High-Performance Computing. Learn to analyze and optimize GPU Kernel to maximize performance for specific AI operations. Contribute to projects such as: Researching, developing, and deploying machine learning and computer vision solutions for AMD's current and future products. Work closely with internal teams to analyze and improve training and inference performance on AMD GPUs. Design and optimize deep learning models specifically for AMD GPU performance. Assisting AI software teams with roadmap planning, collateral development, and customer engagements. Engage with framework maintainers to ensure code changes are aligned with requirements and integrated upstream. Apply sound engineering principles to ensure robust, maintainable solutions.
1.针对大模型训练、强化学习、推理场景,负责GPU kernel开发与调优 2.深入分析GPU/主流AI芯片的硬件架构特性,设计并实现高性能算子、算法和相关组件 3.保持关注行业前沿技术,且有能力和热情开展创新研究

1.基于地平线BPU芯片架构设计开发高性能计算库,针对LLM核心算子(Attention/FFN/MOE等)进行硬件适配优化 2.与芯片设计团队协作,提出计算库需求并验证硬件性能瓶颈,推动芯片架构迭代 3.开发混合精度(FP8/INT4)算子库,适配地平线征程系列芯片的异构计算特性