
MiniMaxMaaS 架构师
社招全职5年以上研发地点:上海 | 北京状态:招聘
工作描述
作为 MaaS 架构师,你将全面负责大模型线上服务的全链路架构设计与质量保障,构建高性能、高可用、可弹性伸缩的模型服务平台,确保模型在生产环境中的 SLA、延迟、吞吐量达到业界领先水平。工作包括不限于: -负责 MaaS 平台架构设计,明确模型从产出到上线的全链路环节,对模型服务的 SLA、延迟、吞吐量等核心指标负责 -主导大模型推理网关的设计与建设,包括多模型路由、流量调度、优先级队列、多租户隔离与 Token 级计量能力 -设计 GPU 资源弹性伸缩策略,结合模型特征与负载信号实现智能调度与资源高效利用 -推动 KV Cache 感知调度的方案设计与落地,包括 Prefix Caching、Paged Attentio…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
系统设计+
https://roadmap.sh/system-design
Everything you need to know about designing large scale systems.
https://www.youtube.com/watch?v=F2FmTdLtb_4
This complete system design tutorial covers scalability, reliability, data handling, and high-level architecture with clear explanations, real-world examples, and practical strategies.
高可用+
https://redis.io/blog/high-availability-architecture/
A high available architecture is when there are a number of different components, modules, or services that work together to maintain optimal performance, irrespective of peak-time loads.
https://www.ibm.com/think/topics/high-availability
High availability (HA) is a term that refers to a system’s ability to be accessible and reliable close to 100% of the time.
流量调度+
https://ferrishall.dev/istio-service-mesh-deepish-dive-architecture-traffic-control-security-and-observability
We’ll be diving a bit deeper, specifically into Istio's architecture, traffic management, security, and observability features.
[英文] Traffic Management
https://istio.io/latest/docs/concepts/traffic-management/
Istio’s traffic routing rules let you easily control the flow of traffic and API calls between services.
https://www.cloudnativedeepdive.com/managing-traffic-in-kubernetes-for-gateways-service-mesh/
There is always traffic moving throughout your cluster, whether it's north/south (ingress/egress) or east/west (service-to-service).
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
性能调优+
https://goperf.dev/
The Go App Optimization Guide is a series of in-depth, technical articles for developers who want to get more performance out of their Go code without relying on guesswork or cargo cult patterns.
https://web.dev/learn/performance
This course is designed for those new to web performance, a vital aspect of the user experience.
https://www.ibm.com/think/insights/application-performance-optimization
Application performance is not just a simple concern for most organizations; it’s a critical factor in their business’s success.
https://www.oreilly.com/library/view/optimizing-java/9781492039259/
Performance tuning is an experimental science, but that doesn’t mean engineers should resort to guesswork and folklore to get the job done.
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
SGLang+
[英文] Install SGLang
https://docs.sglang.ai/get_started/install.html
SGLang is a fast serving framework for large language models and vision language models.
https://github.com/sgl-project/sgl-learning-materials
还有更多 •••
相关职位
社招3年以上云智能集团
1、基础开发能力:具备Python/Java/Go/Node.js 等至少一种语言的 3 年以上开发经验,有后端分布式系统设计、大数据平台或AI应用构建经验。 2、业务架构能力:深刻理解AI应用与Ma
更新于 2026-07-21成都|北京|武汉
社招5年以上云智能集团
1、基础开发能力:具备Python/Java/Go/Node.js 等至少一种语言的 3 年以上开发经验,有后端分布式系统设计、大数据平台或AI应用构建经验。 2、业务架构能力:深刻理解AI应用与Ma
更新于 2026-07-22成都|北京|武汉

社招3年以上
1、基础开发能力:具备Python/Java/Go/Node.js 等至少一种语言的 3 年以上开发经验,有后端分布式系统设计、大数据平台或AI应用构建经验。 2、业务架构能力:深刻理解AI应用与Ma
更新于 2026-04-07北京|杭州|上海