小红书大数据生产平台Agent开发工程师(Java方向)
社招全职1-3年后端开发地点:北京 | 上海 | 杭州状态:招聘
工作描述
任职要求 1. 本科及以上学历,计算机相关专业,有生产平台开发经验者优先。 2. 熟练掌握Java编程语言,深入理解Hadoop生态组件(HDFS/YARN/Hive/Spark/Flink等),有实际项目落地经验。 3. 有高可用系统的设计经验和能力,具备高并发、海量数据的处理能力。 4. 熟…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
Java+
https://www.youtube.com/watch?v=eIrMbAQSU34
Master Java – a must-have language for software development, Android apps, and more! ☕️ This beginner-friendly course takes you from basics to real coding skills.
Hadoop+
https://www.runoob.com/w3cnote/hadoop-tutorial.html
Hadoop 为庞大的计算机集群提供可靠的、可伸缩的应用层计算和存储支持,它允许使用简单的编程模型跨计算机群集分布式处理大型数据集,并且支持在单台计算机到几千台计算机之间进行扩展。
[英文] Hadoop Tutorial
https://www.tutorialspoint.com/hadoop/index.htm
Hadoop is an open-source framework that allows to store and process big data in a distributed environment across clusters of computers using simple programming models.
HDFS+
https://hadoop.apache.org/docs/r1.2.1/hdfs_design.html
The Hadoop Distributed File System (HDFS) is a distributed file system designed to run on commodity hardware.
https://www.ibm.com/cn-zh/think/topics/hdfs
Hadoop 分布式文件系统 (HDFS) 是一种管理大型数据集的文件系统,可在商用硬件上运行。
Yarn+
[英文] Introduction
https://yarnpkg.com/getting-started
Yarn is an established open-source package manager used to manage dependencies in JavaScript projects.
Hive+
[英文] Hive Tutorial
https://www.tutorialspoint.com/hive/index.htm
Hive is a data warehouse infrastructure tool to process structured data in Hadoop. It resides on top of Hadoop to summarize Big Data, and makes querying and analyzing easy.
https://www.youtube.com/watch?v=D4HqQ8-Ja9Y
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
高可用+
https://redis.io/blog/high-availability-architecture/
A high available architecture is when there are a number of different components, modules, or services that work together to maintain optimal performance, irrespective of peak-time loads.
https://www.ibm.com/think/topics/high-availability
High availability (HA) is a term that refers to a system’s ability to be accessible and reliable close to 100% of the time.
还有更多 •••
相关职位
实习核心本地商业-基
1. 27届及以后毕业的本硕博在读同学,计算机类/数学类专业,有扎实的学科理论基础,有相关专业竞赛经历 2. 英语读写流利,cet六级及以上 3. 具有良好的逻辑思维和分析能力,能够准确理解数据任务的
更新于 2026-07-21北京|成都
实习A181972A
1、本科及以上学历在读,计算机、数据科学、统计等相关专业优先; 2、具备较强的数据分析能力,熟练使用Excel/SQL/Python中至少一种工具,能够独立完成数据清洗、分析和结论输出; 3、有Age
更新于 2026-03-23北京
实习A161913A
1、本科及以上学历在读,专业不限,理科、社会科学、建筑学、法学、经济学或具备文理交叉背景的同学优先; 2、有大模型数据相关实习经历,了解大模型训练和评估的基本链路,对数据生产、标注、评估等环节有实践经
更新于 2026-05-28北京