月之暗面多模态强化学习实习生
校招全职北京市地点:北京状态:招聘
工作描述
参与开发多模态大模型对齐及强化学习算法研发、数据构造,提升大模型在应用场景的效果 与 Kimi 产品、Kimi 开放平台等团队合作,实现大模型实际应用中需要的功能点 评估并跟踪大模型在实际应用中的表现,通过优化…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
还有更多 •••
相关职位
实习阿里巴巴研究型实
1. 计算机、人工智能或数学相关专业博士,有扎实的计算机知识和LLM功底。 2. 掌握Qwen/DeepSeek-R1等LLM训练方式,常见PPO/GRPO/Self-Play等强化学习算法原理,有R
更新于 2026-03-20杭州
社招2年以上AI类
工作职责: 1. 负责强化学习训练框架的架构设计、研发与性能优化,根据业务需求持续演进训练策略与系统能力,提升大规模模型训练效率。 2. 深度分析与定位训练系统中的性能瓶颈(包括计算、通信、存储等),
更新于 2026-08-13上海