蚂蚁金服【Plan A】AI工程师-大模型评测-百灵-27届
校招全职2027届蚂蚁星- Plan A人才计划地点:北京 | 上海 | 杭州状态:招聘
工作描述
任职要求 1. 大规模评测体系构建 * 设计支撑万亿参数、百万亿 token 规模训练的评测框架与流水线; * 结合分布式训练体系(TP / PP / SP / EP / CP),排查评测结果中的系统性噪声与偏差; * 建立模型能力的定量评估框架,支撑架构 / 超参 / 数据 scaling laws 的验证。 2. 评测与训练系统协同 * 深入 Megatron / DeepSpeed / FSDP 等框架底层,理解训练过程对评测结果的潜在影响; * 结合 CUDA / NCCL 通信优化原理,定位评测流程中的性能瓶颈并推动优化; * 与系统团队协作,完成评测系统与训练系统的架构级 co-design,确保评测结果的可信度。 3. 能力边界的定量刻画 * 从算法目标反推评测指标设计,构建能反映深度思考、长…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
Megatron+
https://www.youtube.com/watch?v=hc0u4avAkuM
DeepSpeed+
https://www.youtube.com/watch?v=pDGI668pNg0
FSDP+
https://docs.pytorch.org/tutorials/intermediate/FSDP_tutorial.html
In DistributedDataParallel (DDP) training, each rank owns a model replica and processes a batch of data, finally it uses all-reduce to sync gradients across ranks.
https://www.youtube.com/watch?v=PjEwLgyzuzQ
FSDP provides a comprehensive framework for large model training in PyTorch.
CUDA+
https://developer.nvidia.com/blog/even-easier-introduction-cuda/
This post is a super simple introduction to CUDA, the popular parallel computing platform and programming model from NVIDIA.
https://www.youtube.com/watch?v=86FAWCzIe_4
Lean how to program with Nvidia CUDA and leverage GPUs for high-performance computing and deep learning.
NCCL+
https://developer.nvidia.com/nccl
The NVIDIA Collective Communication Library (NCCL) implements multi-GPU and multi-node communication primitives optimized for NVIDIA GPUs and networking.
还有更多 •••
相关职位
实习蚂蚁星- Pla
1. 计算机、人工智能、数学、统计学或相关专业博士及以上学历,面向应届毕业生; 2. 熟悉 Transformer、语言模型训练和常见后训练方法,能够使用 PyTorch 完成模型实验与分析; 3.
更新于 2026-05-12北京|杭州
校招2027届蚂蚁星
1. 计算机、人工智能、数学、统计学或相关专业博士及以上学历,面向应届毕业生; 2. 熟悉 Transformer、语言模型训练和常见后训练方法,能够使用 PyTorch 完成模型实验与分析; 3.
更新于 2026-05-14北京|杭州
实习蚂蚁星- Pla
1. 深入理解分布式训练体系(TP / PP / SP / EP / CP 等); 2. 熟悉主流大模型训练框架底层实现(Megatron / DeepSpeed / FSDP 等); 3. 熟悉 C
更新于 2026-05-12北京|上海|杭州