
MiniMax大模型技术招聘负责人
社招全职5年以上地点:上海 | 北京状态:招聘
工作描述
关于这个角色 从 Scaling Law 到深度推理(Reasoning),从 Pre-training 到 Post-training & RL(强化学习),大语言模型(LLM)正从简单的文本续写,迈向具备慢思考、自我修正与复杂逻辑推演的“智能大脑”时代! 我们不需要一个死板的“简历筛选器”,我们寻找的是一位“懂算法演进、有极客追求、人际敏感度极高”的 Talent Architect(人才架构师)。你将作为研发团队的“Talent Co-pilot”,帮我们把全球最顶尖的算法大脑、系统专家与研究天才,聚集到通用大模型与文本推理的主战场上。 职责描述 1. 全链路 Talent Hunt(组建 AGI 核心“智囊团”) 统筹并实施通用大模型与文本推理方向人才的全流程招聘。 重点攻坚方向包括但不限于:Pre-training(大规模预训练)、Post-training(SFT/RLHF/DPO)、Reasoning & Search(复杂逻辑推理/Chain-of-Thought/过程奖励模型 PRM)、MoE 架构、Long-Context(长文本)、Agent 与代码/数学推理、高效推理加速与基建等。 2. 深度 BP 与“技术 Prompt”对齐 与 LLM 预训练、Post-Training 及推理算法…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
AI agent+
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
Prompt+
https://cloud.google.com/vertex-ai/generative-ai/docs/learn/prompts/introduction-prompt-design
A prompt is a natural language request submitted to a language model to receive a response back.
https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/prompt-engineering
These techniques aren't recommended for reasoning models like gpt-5 and o-series models.
https://www.youtube.com/watch?v=LWiMwhDZ9as
Learn and master the fundamentals of Prompt Engineering and LLMs with this 5-HOUR Prompt Engineering Crash Course!
NeurIPS+
https://neurips.cc/
还有更多 •••
相关职位
社招A182329A
1、计算机、人工智能、软件工程、数字媒体等相关专业本科及以上学历;具备AI应用、模型服务、技术支持、解决方案或企业项目交付经验者优先; 2、熟练掌握Python或Golang,具备代码调试、日志分析和
更新于 2026-07-17深圳
社招A199128
核心能力要求 数据判断:买训练语料时,能判断数据质量(清洗程度、OCR 准确度、去重后的真实增量)、领域稀缺性与可交付性;会设计抽样验证,不被精选样本误导,给出"值不值得买、值多少钱"的明确判断。 A
更新于 2026-06-26北京
社招3年以上核心本地商业-基
1.计算机、人工智能或信息安全相关专业背景,硕士及以上学历,在安全领域(主机安全、网络安全、流量安全、风控等)有实际业务经验者优先。 2.熟悉模型训练、优化与部署流程,具备安全模型训练的实际项目经验。
更新于 2026-04-30北京