快手高级数据研发工程师(直播)-【数据平台】
社招全职1-3年J0012地点:北京状态:招聘
任职要求
1、有Hudi、Hive、Kafka、Spark、Flink、HBase等两种以上两年以上使用经验;(熟悉Spark MLlib, Flink ML或其他大数据生态中的机器学习库者优先;了解如何利用这些框架支持AI工作流); 2、熟悉数据仓库理论方法及ETL相关技术,对于数据的架构和设计有一定的思考,具备良好的数学思维和建模思维;具备将业务问题转化为数据问题和潜在AI应用场景的思维;理解特征工程的基本概念及其在数据管道中的重要性; 3、熟悉分布式计算框架,掌握分布式计算的设计与优化能力,了解流式计算;了解分布式机器学习/深度学习框架(如Tens…
登录查看完整任职要求
微信扫码,1秒登录
工作职责
1、负责直播数据仓库的建设,构建直播领域垂直应用的数据集市;探索并应用智能化的数据管理方法(如元数据自动化管理、数据质量智能监控)提升效率; 2、定义并开发业务核心指标数据,负责垂直业务数据建模;结合业务目标,探索利用AI/ML技术(如预测性指标设计、用户/内容理解特征工程)提升数据洞察深度; 3、根据业务需求,提供大数据计算应用服务,并持续优化改进;支持基于AI模型训练、推理、评估的数据计算需求,并探索利用AI技术(如自动化调优、智能资源调度)优化计算性能和成本; 4、参与直播数据平台的开发工作,支持业务需求;关注并评估AI技术(如AutoML、向量数据库、LLM应用)在数据平台中的应用潜力,协助平台智能化演进。
包括英文材料
Hudi+
[英文] Spark Quick Start
https://hudi.apache.org/docs/quick-start-guide
we will walk through code snippets that allows you to insert, update, delete and query a Hudi table.
https://www.oreilly.com/library/view/apache-hudi-the/9781098173821/
Overcome challenges in building transactional guarantees on rapidly changing data by using Apache Hudi.
https://www.youtube.com/watch?v=pyK18sDYnS0
In this video, I'll introduce you to one of the most popular Data Lake solutions out there, Apache Hudi!
Hive+
[英文] Hive Tutorial
https://www.tutorialspoint.com/hive/index.htm
Hive is a data warehouse infrastructure tool to process structured data in Hadoop. It resides on top of Hadoop to summarize Big Data, and makes querying and analyzing easy.
https://www.youtube.com/watch?v=D4HqQ8-Ja9Y
Kafka+
https://developer.confluent.io/what-is-apache-kafka/
https://www.youtube.com/watch?v=CU44hKLMg7k
https://www.youtube.com/watch?v=j4bqyAMMb7o&list=PLa7VYi0yPIH0KbnJQcMv5N9iW8HkZHztH
In this Apache Kafka fundamentals course, we introduce you to the basic Apache Kafka elements and APIs, as well as the broader Kafka ecosystem.
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
HBase+
[英文] HBase Tutorial
https://www.tutorialspoint.com/hbase/index.htm
HBase is a data model that is similar to Google's big table designed to provide quick random access to huge amounts of structured data. This tutorial provides an introduction to HBase, the procedures to set up HBase on Hadoop File Systems, and ways to interact with HBase shell.
大数据+
https://www.youtube.com/watch?v=bAyrObl7TYE
https://www.youtube.com/watch?v=H4bf_uuMC-g
With all this talk of Big Data, we got Rebecca Tickle to explain just what makes data into Big Data.
机器学习+
https://www.youtube.com/watch?v=0oyDqO8PjIg
Learn about machine learning and AI with this comprehensive 11-hour course from @LunarTech_ai.
https://www.youtube.com/watch?v=i_LwzRVP7bg
Learn Machine Learning in a way that is accessible to absolute beginners.
https://www.youtube.com/watch?v=NWONeJKn6kc
Learn the theory and practical application of machine learning concepts in this comprehensive course for beginners.
https://www.youtube.com/watch?v=PcbuKRNtCUc
Learn about all the most important concepts and terms related to machine learning and AI.
还有更多 •••
相关职位
社招A07123
1、负责企业级数据仓库设计、建模、规范以及架构和研发工作; 2、深入理解业务需求,规划数据仓库建设的整体方向和技术路线,构建快速、准确、灵活、实用的数据仓库,挖掘发挥数据的价值; 3、保证数据仓库的及时、准确、稳定产出,构建好用的数据仓库; 4、打造行业一流的数据仓库团队; 5、与团队一起调研和实践热门数据仓库组件和技术(数据湖、TiDB、FlinkSql、实时数仓、Vibe Coding等);
更新于 2026-02-04北京
社招3-5年J0012
1、建设可灵AI的基础数据能力,提供丰富、稳定的内容生成AI类产品的公共基础数据; 2、建设核心数据资产,与业务场景深度结合,为内容生成服务提供端到端的数据服务和数据解决方案; 3、建设数据治理和管理体系,结合业务+元数据+技术,保障公司各个业务服务的数据质量和产出稳定; 4、能够在业务成长期探索挖掘更多数据研发岗带来业务增量价值的机会; 5、有利用AI进行效率提升的能力; 6、有AI内容理解能力,并进行数据化落地的候选人优先考虑。
更新于 2026-06-18北京
社招1-3年J0012
1、建设全站的基础数据能力,提供丰富、稳定的短视频社区公共基础数据,探索更多数据能力的增量价值; 2、支持运营方向各类数据专题体系的建设,通过数据+算法+产品,赋能业务,提供全链路、可分析、可复用的数据能力,提供更直观、更具分析指导性的产品化能力; 3、建设公司层面的核心数据资产,与业务场景深度结合,为社区服务提供数据服务化、数据业务化的数据&产品解决方案; 4、建设全站数据治理和管理体系,结合业务+元数据+技术,保障公司各个业务服务的数据质量和产出稳定。
更新于 2026-03-17北京