58同城AI Infra工程师——数据方向
校招全职技术类地点:北京状态:招聘
工作描述
职位描述:负责 Hive、Spark、Flink、YARN、HDFS 等开源大数据组件的二次封装、任务调度优化、资源调度治理、容错机制迭代。沉淀标准化离线/实时计算组件,统一数据开发范…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
Hive+
[英文] Hive Tutorial
https://www.tutorialspoint.com/hive/index.htm
Hive is a data warehouse infrastructure tool to process structured data in Hadoop. It resides on top of Hadoop to summarize Big Data, and makes querying and analyzing easy.
https://www.youtube.com/watch?v=D4HqQ8-Ja9Y
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
Yarn+
[英文] Introduction
https://yarnpkg.com/getting-started
Yarn is an established open-source package manager used to manage dependencies in JavaScript projects.
HDFS+
https://hadoop.apache.org/docs/r1.2.1/hdfs_design.html
The Hadoop Distributed File System (HDFS) is a distributed file system designed to run on commodity hardware.
https://www.ibm.com/cn-zh/think/topics/hdfs
Hadoop 分布式文件系统 (HDFS) 是一种管理大型数据集的文件系统,可在商用硬件上运行。
大数据+
https://www.youtube.com/watch?v=bAyrObl7TYE
https://www.youtube.com/watch?v=H4bf_uuMC-g
With all this talk of Big Data, we got Rebecca Tickle to explain just what makes data into Big Data.
还有更多 •••
相关职位
校招研发类
1、计算机科学、软件工程、人工智能、计算机工程、机器学习、数据工程等相关专业,具备扎实的计算机系统基础; 2、具备较强的代码编写和算法实现能力,熟悉模型架构、数据工程、操作系统、分布式计算与存储系统、
更新于 2026-07-23北京|杭州|上海
实习研发类
1、计算机科学、软件工程、人工智能、计算机工程、机器学习、数据工程等相关专业,具备扎实的计算机系统基础。 2、具备较强的代码编写和算法实现能力,熟悉模型架构、数据工程、操作系统、分布式计算与存储系统、
更新于 2026-03-18北京|上海|东莞

社招3年以上技术类-数据
1、本科及以上学历,计算机科学、统计学、数学、电子工程或相关专业;硕士及以上优先。 2、3 年以上数据开发、数据分析或AI Infra相关工作经验; 3、有 AI Infra或云计算平台相关数据开发或
更新于 2026-04-02杭州