拼多多搜广推数据开发工程师
社招全职技术类地点:上海状态:招聘
工作描述
任职要求 1、精通数据仓库模型设计,具备丰富的ETL开发及海量数据加工处理经验; 2、具备丰富的Flink实时计算开发与调优经验,精通高并发、高可用、可扩展分布式系统设计原则与实践; 3、具备分布式数据存储与计算平台(Hadoop生态)应用开发经验,熟练掌握HDFS, MapRedu…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
数据仓库+
https://www.youtube.com/watch?v=9GVqKuTVANE
From Zero to Data Warehouse Hero: A Full SQL Project Walkthrough and Real Industry Experience!
https://www.youtube.com/watch?v=k4tK2ttdSDg
ETL+
https://www.ibm.com/think/topics/etl
ETL—meaning extract, transform, load—is a data integration process that combines, cleans and organizes data from multiple sources into a single, consistent data set for storage in a data warehouse, data lake or other target system.
https://www.youtube.com/watch?v=OW5OgsLpDCQ
It explains what ETL is and what it can do for you to improve your data analysis and productivity.
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
高并发+
https://www.baeldung.com/concurrency-principles-patterns
In this tutorial, we’ll discuss some of the design principles and patterns that have been established over time to build highly concurrent applications.
https://www.baeldung.com/java-concurrency
Handling concurrency in an application can be a tricky process with many potential pitfalls. A solid grasp of the fundamentals will go a long way to help minimize these issues.
https://www.oreilly.com/library/view/concurrency-in-go/9781491941294/
You’ll understand how Go chooses to model concurrency, what issues arise from this model, and how you can compose primitives within this model to solve problems.
https://www.oreilly.com/library/view/modern-concurrency-in/9781098165406/
With this book, you'll explore the transformative world of Java 21's key feature: virtual threads.
https://www.youtube.com/watch?v=qyM8Pi1KiiM
https://www.youtube.com/watch?v=wEsPL50Uiyo
高可用+
https://redis.io/blog/high-availability-architecture/
A high available architecture is when there are a number of different components, modules, or services that work together to maintain optimal performance, irrespective of peak-time loads.
https://www.ibm.com/think/topics/high-availability
High availability (HA) is a term that refers to a system’s ability to be accessible and reliable close to 100% of the time.
还有更多 •••
相关职位
社招3-5年引擎
1、计算机相关专业,本科以上学历,两年以上工作经验; 2、扎实的计算机专业基础知识,精通数据结构/算法设计; 3、熟练使用Hadoop、Spark、Hive、Flink等大数据处理框架。有海量数据
更新于 2026-08-27北京|上海|深圳
社招3-5年引擎
1、计算机相关专业,本科以上学历,两年以上工作经验; 2、扎实的计算机专业基础知识,精通数据结构/算法设计; 3、熟练使用Hadoop、Spark、Hive、Flink等大数据处理框架。有海量数据
更新于 2026-08-27北京|上海|深圳
实习引擎
1. 本科及以上学历,计算机、软件工程、数据科学等相关专业; 2. 熟悉至少一种编程语言,如 Python、Java 或 C++; 3. 对大数据处理工具有一定的了解,如 Flink、Hadoop、S
更新于 2026-06-30上海|北京