月之暗面Infra 系统工程师 - 推理系统
社招全职技术类/Technical地点:北京状态:招聘
工作描述
工作职责:
优化超大规模线上推理集群运行效率和稳定性
跟踪最新模型进展,设计技术架构将模型应用到产品
设计和优化工作流,通过自动化手段加速模型到产品的迭代效率
任职要求:
拥有良好的代码开发…登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
系统设计+
https://roadmap.sh/system-design
Everything you need to know about designing large scale systems.
https://www.youtube.com/watch?v=F2FmTdLtb_4
This complete system design tutorial covers scalability, reliability, data handling, and high-level architecture with clear explanations, real-world examples, and practical strategies.
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
Go+
https://www.youtube.com/watch?v=8uiZC0l4Ajw
学习Golang的完整教程!从开始到结束不到一个小时,包括如何在Go中构建API的完整演示。没有多余的内容,只有你需要知道的知识。
Rust+
https://www.youtube.com/watch?v=BpPEoZW5IiY
In this comprehensive Rust course for beginners, you will learn about the core concepts of the language and underlying mechanisms in theory.
https://www.youtube.com/watch?v=lzKeecy4OmQ
Full Rust 101 Crash Course for beginners.
https://www.youtube.com/watch?v=rQ_J9WH6CGk
数据结构+
https://www.youtube.com/watch?v=8hly31xKli0
In this course you will learn about algorithms and data structures, two of the fundamental topics in computer science.
https://www.youtube.com/watch?v=B31LgI4Y4DQ
Learn about data structures in this comprehensive course. We will be implementing these data structures in C or C++.
https://www.youtube.com/watch?v=CBYHwZcbD-s
Data Structures and Algorithms full course tutorial java
还有更多 •••
相关职位
校招
岗位职责: 1. 参与内部训练平台、LLM 推理集群服务的建设和维护。 2. 通过自动化手段,保障内部超大规模训练任务的高效运行。 3. 参与平台的稳定性建设,包括异常监控、故障排查与修复等,确保平台
更新于 2025-03-25北京
社招技术类/Tech
工作职责: - 基于 Kubernetes 构建稳定、可扩展的大模型训练平台,实现 GPU 算力资源的自动化管理与高效调度。 - 通过自动化手段,保障内部超大规模训练任务的高效运行。 - 优化训练作业
更新于 2026-04-01北京
社招3年以上程序&技术类
1)本科及以上学历,计算机科学、软件工程、人工智能、分布式系统或相关专业 2)3 年以上基础设施、后端平台、分布式系统或 AI Infra 相关研发经验,有大规模训练平台、在线推理服务或高性能系统建设
上海|北京