阿里巴巴1688AI大模型应用算法专家-阿里星
实习兼职阿里巴巴2027届实习生地点:杭州状态:招聘
工作描述
任职要求 1.计算机、人工智能、数学/统计等相关专业硕士及以上;对大模型与 Agent 有强烈好奇心与持续动手动力。 2.熟悉大模型落地常用技术栈:SFT 与对齐(DPO/RLHF 等)、RAG/知识增强、工具/函数调用、多轮对话与长上下文处理。 3.具备 Agent/多智能体研发经验:任务规划与编排、记忆/状态管理、工具/技能体系(可包含 MCP 等),能用评测驱动快速迭代。 4.工程与结果导向:能围绕业务指标搭建离线评测与线上 A/B,把原型快速打到可用、可扩、可稳定的线上系统;善于定位疑难问题并持续优化。 5.加分项:顶会论…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
AI agent+
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
还有更多 •••
相关职位
实习阿里巴巴2027
硬性要求: ● 计算机科学、人工智能、运筹学、自动化等相关方向博士 ● 扎实的强化学习理论功底,精通主流RL算法(Policy Gradient/PPO/GRPO及其变体),有实际实现与调优经验 ●
更新于 2026-05-28杭州
实习阿里巴巴2027
1. 计算机/电子/通信/网络/人工智能等相关专业,毕业时间在 2026 年 11 月以后的硕士及以上学历在校生; 2. 熟悉网络传输协议(TCP/QUIC/RTP/WebRTC 等),深入理解拥塞控
更新于 2026-05-28杭州
实习阿里巴巴2027
1、毕业时间在 2026年11月及以后的在校硕博同学,计算机视觉、计算机图形学、机器学习等相关专业 2、具备计算机图形学和计算机视觉理论基础; 3、具备极佳的工程实现能力,熟练掌握C++/Java/P
更新于 2026-05-28杭州