字节跳动数据湖存储专家-Hudi
社招全职Y8213地点:杭州状态:招聘
任职要求
1、具备良好的 Java / Scala 编程基础和良好的计算机技术基础,同时具备良好的沟通能力和团队协作能力; 2、熟悉开源数据湖存储方案 Hudi,Iceberg,Delta Lake 的原理及源码,有内核开发经验或社区贡献者优先,开源社区 committer / PMC 优先; 3、精通任一个 Parquet,ORC,Arrow 列存格式,或者 Avro,Protobuf 行存格式者优先; 4、熟悉 Spark、Flink、Presto、Hive 等主流大数据计算引擎者优先。
工作职责
数据引擎-数据湖团队,旨在打造业界领先的 EB 级超大规模数据湖,支持字节跳动众多业务线,如抖音、今日头条、电商。同时基于内部最佳实践,在火山引擎上打造一款云原生实时湖仓一体的 toB 产品——湖仓一体分析服务LAS(LakeHouse Analytics Service)。 职位描述: 1、打造业界领先的基于 HUDI 构建的 EB 级湖仓一体解决方案,支撑字节跳动众多核心业务线(如抖音,今日头条,电商)和 ToB 业务; 2、负责围绕数据湖构建一站式全托管优化服务,数据湖内核的极致优化,以及流批一体的数据湖加速层的设计与研发; 3、负责数据湖存储的生态研发,与 Spark、Flink、Presto、Hive 等计算引擎深度结合; 4、与开源社区紧密合作,持续构建开源影响力,有机会成长为 HUDI committer / PMC。
包括英文材料
Java+
https://www.youtube.com/watch?v=eIrMbAQSU34
Master Java – a must-have language for software development, Android apps, and more! ☕️ This beginner-friendly course takes you from basics to real coding skills.
Scala+
内核+
https://www.youtube.com/watch?v=C43VxGZ_ugU
I rummage around the Linux kernel source and try to understand what makes computers do what they do.
https://www.youtube.com/watch?v=HNIg3TXfdX8&list=PLrGN1Qi7t67V-9uXzj4VSQCffntfvn42v
Learn how to develop your very own kernel from scratch in this programming series!
https://www.youtube.com/watch?v=JDfo2Lc7iLU
Denshi goes over a simple explanation of what computer kernels are and how they work, alonside what makes the Linux kernel any special.
Parquet+
https://www.youtube.com/watch?v=KLFadWdomyI
Learn all about Apache Parquet, a column-based file format that's popular in the Hadoop/Spark ecosystem.
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
Presto+
[英文] What is Presto?
https://prestodb.io/what-is-presto/
https://www.tutorialspoint.com/apache_presto/index.htm
Hive+
[英文] Hive Tutorial
https://www.tutorialspoint.com/hive/index.htm
Hive is a data warehouse infrastructure tool to process structured data in Hadoop. It resides on top of Hadoop to summarize Big Data, and makes querying and analyzing easy.
https://www.youtube.com/watch?v=D4HqQ8-Ja9Y
大数据+
https://www.youtube.com/watch?v=bAyrObl7TYE
https://www.youtube.com/watch?v=H4bf_uuMC-g
With all this talk of Big Data, we got Rebecca Tickle to explain just what makes data into Big Data.
相关职位
社招R9882
1、打造业界领先的基于HUDI构建的EB级湖仓一体解决方案,支撑字节跳动众多核心业务线(如抖音,今日头条,电商)和ToB业务; 2、负责围绕数据湖构建一站式全托管优化服务,数据湖内核的极致优化,以及流批一体的数据湖加速层的设计与研发; 3、负责数据湖存储的生态研发,与Spark、Flink、Presto、Hive等计算引擎深度结合; 4、与开源社区紧密合作,持续构建开源影响力,有机会成长为HUDI Committer/PMC。
更新于 2023-03-06
社招2年以上A38455
1、负责多模态数据湖内核与存储引擎的研发工作,在Data+AI场景提供行业数据湖解决方案; 2、负责与上层数据处理产品深度联动,建设多模数据湖生态; 3、结合字节跳动、国内头部大模型客户场景,支持多模态数据管理需求; 4、与开源社区深度合作,提升开源影响力。
更新于 2025-05-19
社招A149021
数据引擎-存储引擎团队,负责开源数据湖 Hudi 的内核研发。团队内部有多名 Apache Committer,在国内外有较强的技术影响力,和国内顶尖的大数据计算、存储领域的专家一起合作,一起打造业界领先的 EB 级超大规模数据湖,并通过火山引擎的湖仓一体平台 LAS 对外输出。 职位描述: 1、打造业界领先的 EB 级湖仓一体解决方案,支撑字节跳动众多业务线(如抖音,今日头条,电商),并通过火山引擎 LAS 产品对外输出; 2、负责数据湖存储产品的架构设计、核心开发和应用落地; 3、负责数据湖产品的长期竞争力规划与推进落地。
更新于 2023-11-27