阿里巴巴阿里国际站-高级数据研发工程师-智能数据部
社招全职4年以上技术类-数据地点:杭州状态:招聘
工作描述
任职要求 1、较为丰富的数据仓库及数据平台架构经验,熟悉数据建模、分层设计、ETL/实时链路开发与数据资产沉淀。期望通过对业务的深入理解发挥数据价值,并反哺数据体系建设;有跨境电商、B2B、国际零售、营销、供应链、商家/商品运营背景尤佳;具备较为系统的海量数据性能处理与优化经验。 2、有从事分布式数据存储与计算平台应用开发经验,熟悉 Hadoop 生态相关技术并有实际开发经验,具备 Spark/Flink 等离线/实时计算框架的开发经验尤佳。熟悉 AI 业务场景下的数据链路(如特征工程、样本构建、模型训练数据供给、数据回流、实时特征服务等)者优先。 3、具有良好的商业敏感度,能够高效地将业务问题转化为数据、算法与 AI 问题,为业务带来切实的数据价值。能综合运用现有数据、算法、产品等能力形成数据化解决方案,并通过 Dat…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
数据仓库+
https://www.youtube.com/watch?v=9GVqKuTVANE
From Zero to Data Warehouse Hero: A Full SQL Project Walkthrough and Real Industry Experience!
https://www.youtube.com/watch?v=k4tK2ttdSDg
ETL+
https://www.ibm.com/think/topics/etl
ETL—meaning extract, transform, load—is a data integration process that combines, cleans and organizes data from multiple sources into a single, consistent data set for storage in a data warehouse, data lake or other target system.
https://www.youtube.com/watch?v=OW5OgsLpDCQ
It explains what ETL is and what it can do for you to improve your data analysis and productivity.
Hadoop+
https://www.runoob.com/w3cnote/hadoop-tutorial.html
Hadoop 为庞大的计算机集群提供可靠的、可伸缩的应用层计算和存储支持,它允许使用简单的编程模型跨计算机群集分布式处理大型数据集,并且支持在单台计算机到几千台计算机之间进行扩展。
[英文] Hadoop Tutorial
https://www.tutorialspoint.com/hadoop/index.htm
Hadoop is an open-source framework that allows to store and process big data in a distributed environment across clusters of computers using simple programming models.
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
特征工程+
https://www.ibm.com/think/topics/feature-engineering
Feature engineering preprocesses raw data into a machine-readable format. It optimizes ML model performance by transforming and selecting relevant features.
https://www.kaggle.com/learn/feature-engineering
Better features make better models. Discover how to get the most out of your data.
还有更多 •••
相关职位
社招3年以上
1. 计算机、数学、统计、人工智能等相关专业,具备扎实的工程基础、数据分析能力和系统设计能力。 2. 有 AI、数据平台、后端平台或 LLM 应用相关研发经验,熟练掌握 Python 和 SQL,熟悉
更新于 2026-09-07杭州
社招A07123
1、精通Unix/Linux操作系统下Java或Scala开发,有良好的编码习惯,有扎实的计算机理论基础; 2、熟练掌握大数据处理技术栈,有丰富的Hadoop/Spark的实际项目使用经验,使用过Fl
更新于 2026-02-04北京
社招3-5年J0012
1、有丰富的数据仓库及数据平台架构经验,通过对AI领域的深入理解,进行数据仓库、数据体系和数据价值的建设和优化; 2、有从事分布式数据存储与计算平台应用开发经验,熟悉Hive,Kafka,Spark,
更新于 2026-06-18北京