腾讯AI编译优化工程师(北京/上海/深圳)
任职要求
1.熟悉Linux开发环境,掌握Python/C++等语言, 有良好的编程基础、系统设计优化能力; 2.熟悉GPU/AI DSA体系结构和原理,有Nvidia/AMD/Intel或 AI 芯片等编译相关的开发优化经验; 3.精通MLIR、Triton、tilelan…
工作职责
1.基于自研芯片,研发高性能图和算子编译器,解决芯片落地过程中编译相关性能、易用性等问题; 2.探索和研发自动生成算子解决方案,提升生成算子效率; 3.不断迭代编译器功能、性能和易用性,和业务一起构建AI生态。
1、参与面向大模型与物理AI场景的AI编译器研发,优化计算图表示、算子融合与内存调度策略; 2、针对自研AI芯片进行算子定制与性能调优,实现端到端推理与训练加速; 3、设计并实现自动代码生成工具链,支持多后端(GPU/NPU/CPU)的高效算子发射; 4、调研SOTA大语言模型的压缩和加速算法,并针对小鹏的模型结构做优化和实现; 5、与算法、芯片团队深度协作,推动编译优化在百万级量产环境中的稳定落地。
【部门介绍】引擎架构部提供小红书搜广推,CV和NLP业务的深度学习模型高性能推理服务。主导SOTA推理引擎的架构设计与核心模块开发,支撑搜广推业务在长序列建模、生成式推荐、Agent等前沿场景在GPU,XPU等异构计算部件上规模落地。 1. 参与推理引擎的架构设计与核心模块的开发,参与AI编译器前后端的设计与实现,优化IR Compile模式下DSL特征处理引擎和AI推理引擎的性能。 2. 分析I/O性能瓶颈、优化编译耗时和codegen性能,改进编译优化算法,不断优化编译器,解决编译部署问题。 3. 优化IR Compile模式下搜广推、长序列、多模态、MoE等深度学习模型的推理效率。 4. 针对GPU/NPU等异构计算芯片,探索基于IR编译优化的片内多部件并行流水线等前沿技术,构建业界影响力。
负责AI编译器研发,通过理解云上AI负载特征,挖掘CPU/GPU硬件与编译框架的协同优化潜力;主导核心功能实现与性能优化,解决AI训练、推理场景中的关键工程挑战,提升计算效率;同时紧密协同跨团队资源,推动AI编译器产品的业务落地与技术价值转化。 职责描述: 1. 理解云上AI负载场景与特点,挖掘CPU/GPU硬件架构与编译器,编程框架协同优化的潜力, 负责核心功能的实现与性能优化。 2. 能够与各团队紧密合作,理解和实现客户需求,参与AI编译器产品的设计、功能开发及业务落地工作。
An exciting internship opportunity to make an immediate contribution to AMD's next generation of technology innovations awaits you! We have a multifaceted, high-energy work environment filled with a diverse group of employees, and we provide outstanding opportunities for developing your career. During your internship, our programs provide the opportunity to collaborate with AMD leaders, receive one-on-one mentorship, attend amazing networking events, and much more. Being part of AMD means receiving hands-on experience that will give you a competitive edge. Together We Advance your career! JOB DETAILS: Location: Beijing,China Onsite/Hybrid: at least 3 days a week, either in a hybrid or onsite or remote work structure throughout the duration of the co-op/intern term. Duration: at least 6 months WHAT YOU WILL BE DOING: We are seeking highly motivated AI Compiler Software Engineering intern/co-op to join our team. In this role – We will involve you in extending Triton’s compiler infrastructure to support new AI workloads and hardware targets. We will assign you tasks to implement and optimize GPU kernels using Triton’s Python-based DSL. We will train you to analyze kernel performance using profiling tools and help you identify bottlenecks and optimization opportunities. We will understand how modern compilers translate high-level abstractions into efficient machine code.