阿里巴巴AI数据工程师
实习兼职阿里巴巴2027届实习生地点:北京 | 杭州状态:招聘
工作描述
任职要求 1.基础条件 ● 计算机、软件工程、数学、统计、人工智能、大数据、机器人等相关专业硕士/博士优先(非此类专业,有相关经验亦可)。 ● 有顶会论文/高影响项目/开源贡献者加分。 2.专业能力 ● 大数据处理技术:深入理解大规模分布式数据处理系统原理,熟悉Spark/Flink/Ray等开源技术栈;深入理解流批处理原理(计算模型、调度和资源管理、容错与一致性等);可独立完成面向全模态数据(结构化/文本/图像/音频/视频)的批流一体数据处理开发与优化。 ● 大模型技术的理解与掌握:深入理解大模型核心原理,包括Transformer架构、上下文学习(ICL)、指令微调(Instruction Tuning)、检索增强生成(RAG)及推理机制(如思维链CoT)等关键技术;熟悉大模型在预训练、监督微调(SFT)和强化学习对齐(RLHF/RLAIF)等阶段的数据需求与优化逻辑。能够基于领域场景设计高质量数据处理与合成算法,通过系统化的数据迭代、评估反馈与模型微调闭环,持续驱动大模型在特定领域的能力提升与性能优化 。 ● AI编程意识与工程思维:能持续快速学习AI研发新范式,熟练运用主流AI工具,独立完成从需求分析、架构设计到高质量代码实现的系统级开发任务,并确保代码的可…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大数据+
https://www.youtube.com/watch?v=bAyrObl7TYE
https://www.youtube.com/watch?v=H4bf_uuMC-g
With all this talk of Big Data, we got Rebecca Tickle to explain just what makes data into Big Data.
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
Ray+
https://github.com/ray-project/ray
Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
https://www.youtube.com/watch?v=FhXfEXUUQp0
In this video, I'll teach you everything you need to know about Apache Ray!
https://www.youtube.com/watch?v=fMiAyj2kgac
Using powerful machine learning algorithms is easy using Ray.io and Python.
https://www.youtube.com/watch?v=q_aTbb7XeL4
Parallel and Distributed computing sounds scary until you try this fantastic Python library.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
RAG+
https://www.youtube.com/watch?v=sVcwVQRHIc8
Learn how to implement RAG (Retrieval Augmented Generation) from scratch, straight from a LangChain software engineer.
还有更多 •••
相关职位
社招2年以上金融服务平台
1、AI能力要求: 1)熟悉大模型基本原理,了解 SFT/DPO/RAG 等主流训练与应用范式;有模型微调、数据标注体系搭建或 AI 数据工程实践经验者优先 2)了解 Agent 开发框架(如 Lan
更新于 2026-04-23上海
社招
1. 24-25届毕业生,本科及以上学历,计算机、数据科学、人工智能、统计等相关专业 2. 熟悉大数据相关技术,有机器学习/算法基础或相关项目经验优先; 3. 善于用AI工具提升工作效率与解决复杂问题
更新于 2026-01-05深圳|长沙
实习MEG
-本科及以上学历,计算机/人工智能/数据科学相关专业 -熟悉数据结构和算法设计,熟悉C++/python程序开发 -能够运用prompt工程辅助数据生成;熟悉多模态模型的训练,能提出改机训练效果的方案
更新于 2025-10-28北京