米哈游【提前批-大模型】Code Agentic算法研究员
校招全职程序&技术类地点:北京状态:招聘
工作描述
任职要求 1、计算机科学、人工智能、机器学习、软件工程或相关领域硕士 / 博士学历。 2、熟悉 Transformer 架构和大模型训练流程,对 SFT、RLHF、DPO、PPO、GRPO 等 Post-training 方法有深入理解和实践经验。 3、具备扎实的代码能力,熟悉软件工程开发流程,理解代码库结构、测试体系、版本控制、CI、调试和工程协作流程。 4、熟练使用 Python 和 PyTorch,熟悉 DeepSpeed、Megatron-LM、FSDP、vLLM 等训练或推理框架中的一种或多种。 5、有代码生成、代码修复、Agent、工具调用、程序验证、自动化测试、SWE-bench 类任务等相关经验者优先。 加分项 1、在 NeurIPS、ICML、I…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
机器学习+
https://www.youtube.com/watch?v=0oyDqO8PjIg
Learn about machine learning and AI with this comprehensive 11-hour course from @LunarTech_ai.
https://www.youtube.com/watch?v=i_LwzRVP7bg
Learn Machine Learning in a way that is accessible to absolute beginners.
https://www.youtube.com/watch?v=NWONeJKn6kc
Learn the theory and practical application of machine learning concepts in this comprehensive course for beginners.
https://www.youtube.com/watch?v=PcbuKRNtCUc
Learn about all the most important concepts and terms related to machine learning and AI.
学历+
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
CI+
https://www.ibm.com/cn-zh/think/topics/continuous-integration
持续集成 (CI) 是一种软件开发实践,开发人员在整个开发周期中会定期将新的代码和代码变更集成到中央代码存储库中。它是 DevOps 和敏捷方法的关键组成部分。
https://www.youtube.com/watch?v=42UP1fxi2SY
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
还有更多 •••
相关职位
校招程序&技术类
1、计算机科学、人工智能、机器学习、软件工程或相关领域在读硕士 / 博士。 2、熟悉 LLM 评测方法,对代码生成、代码修复、Agent 评测或自动化测试有实践经验。 3、熟悉 Python,具备较强
北京
校招程序&技术类
1.具备大语言模型数据研发相关工作经验; 2.熟悉大模型完整训练链路,包含预训练、中段训练、后训练阶段; 3.具备agent模型训练与评测实操经验; 4.有 agentic, reasoning, c
上海|北京
校招程序&技术类
1.熟练掌握 Apache Spark、Ray 等大规模分布式数据处理框架; 2.熟悉大模型完整训练链路,包含预训练、中段训练、后训练阶段; 3.对数据质量具备严苛把控意识,能够针对各类代码、文本语料
上海|北京