阿里巴巴研究型实习生-1688-多模态强化学习算法工程师
实习兼职阿里巴巴研究型实习生地点:杭州状态:招聘
工作描述
任职要求 1. 计算机、人工智能或数学相关专业博士,有扎实的计算机知识和LLM功底。 2. 掌握Qwen/DeepSeek-R1等LLM训练方式,常见PPO/GRPO/Self-Play等强化学习算法原理,有RL实操经验。 3. 熟悉DeepRes…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
还有更多 •••
相关职位
校招算法类
1、硕士及以上学历,毕业时间:2025年9月-2026年8月,计算机、人工智能、软件工程等相关专业; 2、扎实的自然语言处理与深度学习理论基础,熟悉Transformer等主流大模型架构及训练机制,理
更新于 2026-03-10北京
实习算法类
1、硕士及以上学历,毕业时间:2026年9月-2027年8月,计算机、人工智能、软件工程等相关专业; 2、扎实的自然语言处理与深度学习理论基础,熟悉Transformer等主流大模型架构及训练机制,理
更新于 2026-06-02北京
实习阿里巴巴2027
1.本科及以上学历,计算机科学、人工智能、电子与通信等相关专业;面向2026年11月及以后的海内外高校在校生。 2.精通Diffusion模型及相关技术,掌握T2V基础模型及相关技术原理,有图像/视频
更新于 2026-05-28杭州