阿里巴巴数据技术及产品部-大模型评测算法-RSI(Recursive Self-Improvement,递归自我改进 / AI R&D 自动化)方向
社招全职5年以上技术类-算法地点:北京 | 杭州状态:招聘
工作描述
Qualifications 1. 硕士及以上,机器学习 / AI 相关专业;3 年以上大模型相关研发经验,或具备高质量论文 / 开源成果(不唯年限,重在深度)。 2. 深刻理解 AI R&D 自动化 / agentic coding / 自动化科研的评测与训练范式,熟悉长程 agentic 任务、RLHF / RLAIF、reward modeling、self-improvement 等技术。 3. 熟悉主流 AI R&D / coding agent benchmark(RE-Bench、MLE-bench、PaperBench、SWE-bench、The AI Scientist 等)及其 rubric / LLM-judge 评分方法与局限。 4. 具备设计严谨评测与训练实验的能力,能识别并防范 reward hacking、数据泄…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
机器学习+
https://www.youtube.com/watch?v=0oyDqO8PjIg
Learn about machine learning and AI with this comprehensive 11-hour course from @LunarTech_ai.
https://www.youtube.com/watch?v=i_LwzRVP7bg
Learn Machine Learning in a way that is accessible to absolute beginners.
https://www.youtube.com/watch?v=NWONeJKn6kc
Learn the theory and practical application of machine learning concepts in this comprehensive course for beginners.
https://www.youtube.com/watch?v=PcbuKRNtCUc
Learn about all the most important concepts and terms related to machine learning and AI.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
R+
[英文] R Tutorial
https://www.w3schools.com/r/
R is often used for statistical computing and graphical presentation to analyze and visualize data.
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
还有更多 •••
相关职位
社招2年以上技术类-算法
1. 计算机科学、人工智能、机器学习或相关领域硕士及以上学历,博士优先。 2. 对大模型评测方法论有体系化认知,能判断不同方法在不同模态/场景下的适用边界。 3. 对数据合成与数据质量有研究深度,理解
更新于 2026-08-04北京|杭州
社招5年以上技术类-算法
1. 硕士及以上,计算机 / AI 相关专业;3 年以上大模型相关研发经验,或具备高质量论文 / 开源成果(不唯年限,重在深度)。 2. 熟悉 Agent 训练、工具调用(tool use / fun
更新于 2026-09-15北京|杭州
社招5年以上技术类-算法
1. 硕士及以上,计算机视觉 / 机器学习 / 机器人等相关专业;3 年以上大模型相关研发经验,或具备高质量论文 / 开源成果(不唯年限,重在深度)。 2. 熟悉世界模型 / 视频生成 / 视频预测
更新于 2026-09-15北京|杭州