logo of amd

AMD模型优化工程师(推理&训练)Model Optimization Engineer (Inference & Training)

社招全职 Engineering地点:北京状态:招聘

任职要求


* Strong software engineering in Python and C/C++. * Practical experience with PyTorch/JAX and building/extending deep learning frameworks. * Hands‑on CUDA and/or ROCm development; experience writing or optimizing GPU kernels. * Experience with Triton (kernel development/optimization) is highly desired. * Proven experience with model optimization techniques, especially low‑bitwidth quantization and other compression methods. * Familiarity with GenAI inference engines and optimizations (e.g., vLLM, SGLang, xDiT, continuous batching, speculative decoding). * Skilled at profiling and performance debugging across stack layers (operator → model → framework → hardware). PREFERRED QUALIFICATIONS * Publications or contributions in model optimization / ML systems are a strong plus. * Experience with distributed traini…
登录查看完整任职要求
微信扫码,1秒登录

工作职责


THE ROLE We are looking for a hands‑on Engineer to design, implement, and optimize AI model training and inference solutions for AMD platforms. The role focuses on end‑to‑end performance and accuracy improvements at the framework, model, and operator levels, with strong emphasis on low‑bitwidth quantization, model compression, and real‑world deployment. You will work closely with AMD hardware and software teams, support customers, and contribute to open‑source projects and inference/training frameworks. KEY RESPONSIBILITIES * Design, implement, and optimize inference and training pipelines for AMD GPUs/accelerators at the framework, model, and operator levels. * Lead research and development of model optimization algorithms: low‑bitwidth quantization, pruning/sparsity, compression, efficient attention mechanisms, and lightweight architectures. * Implement and tune CUDA/ROCm/Triton kernels for critical operators; profile and eliminate performance bottlenecks. * Integrate and optimize models for PyTorch/JAX and common distributed training/inference stacks (Torchtitan, Megatron, DeepSpeed, HF Transformers, etc.). * Reduce latency and increase throughput for large‑model inference (e.g., batching strategies, caching, speculative decoding). * Contribute to and/or maintain open‑source inference/training tools, ensuring production readiness and community adoption. * Provide technical support and guidance to customers and internal teams to achieve target accuracy and performance on AMD platforms. TECHNICAL
包括英文材料
Python+
C+
C+++
PyTorch+
JAX+
开发框架+
还有更多 •••
相关职位

logo of weibo
社招新浪&微博

1、负责 CTR 模型在 GPU 上的训练推理性能优化,支撑高并发、低延迟场景下的线上业务稳定落地; 2、结合大模型技术,推动 CTR 模型升级的工程优化,提升推理效率与资源利用率。

更新于 2026-04-03北京
logo of alibaba
社招1年以上

- 设计模型分布式并行策略,对推理性能进行分析与优化,在给定硬件配置下找到最优部署方案 - 针对模型实现低精度量化方案,完成精度对齐,确保可在生产环境部署 - 协同框架、算子、运行时等相关团队,制定模型系统性部署方案,为最终模型交付质量负责 - 构建大模型Model Zoo,打造模型推理参考样例,优化模型评测标准和流程

更新于 2026-07-16成都|北京|杭州
logo of didi
社招技术

参与车载端AI模型的优化、部署与调度,保障车端系统高效、稳定运行; 深入参与车端系统稳定性建设,对系统级风险与问题具备独立排查与根因分析能力,高效解决复杂疑难问题; 支持大模型在仿真、标注等环境中的服务化部署与验证,助力业务模型快速迭代; 跟踪业界前沿技术动态,持续探索并引入先进的性能优化方法与工具链;

更新于 2026-07-07北京
logo of kuaishou
社招3-5年J0011

1、和领域内最顶尖的算法工程师合作,一起研发业内领先的大模型推理优化方案,优化目标包括但不限于视频生成大模型、多模态大模型; 2、调研大模型推理优化方向最新论文,方向包括但不限于高性能算子开发、大模型量化、分布式大模型并行推理、投机推理等。

更新于 2026-06-22北京|上海