小米大模型安全能力研究员-MiMo
社招全职A41838地点:北京状态:招聘
工作描述
任职要求 岗位要求 - 计算机、网络安全等相关方向硕博学历以上,具备 CTF 参赛经历或漏洞挖掘实战经验; - 熟悉大模型训练全链路(预训练、SFT、RLHF) - 具备扎实的系统编程能力(C/C++/Rust 其一) - 了解软件漏洞原理,如缓冲区溢出、UAF、逻辑漏洞等 加分项 - 熟悉逆向工程、fuzzing(AFL/LibFuzzer)、符号执行 - 有 LLM for co…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
还有更多 •••
相关职位
实习阿里巴巴2027
1、扎实的机器学习/深度学习基础,熟悉大模型评测方法(LLM-as-a-Judge、红队测试、对抗评测等); 2、深入理解 Claude、OpenAI 等顶尖实验室的模型评测技术体系,具备改进与创新能
更新于 2026-03-23北京|杭州
实习阿里巴巴2027
1.计算机、数学、人工智能、网络安全等相关专业硬士及以上学历,3年以上算法研发或安全研发经验,有智能体安全、系统安全、大模型安全或偏好对齐领域背景者优先; 2.精通Python/C++/Java/Go
更新于 2026-05-11北京|杭州
实习A179699
1、2027届本科及以上学历在读,计算机、统计、数据科学或相关专业; 2、可提供至少3个月实习周期;具备机器学习基础知识,了解NLP或CV方向模型训练流程,理解准确率、召回率等模型评估指标; 3、熟练
更新于 2026-05-28北京