蚂蚁金服钱塘征信-数据研发工程师 (数仓/大数据/ETL开发)-征信方向
社招全职3年以上技术类-数据地点:杭州状态:招聘
工作描述
任职要求 1、有较为丰富的数仓设计&开发经验,熟悉ETL分层建设方法、数据、维度建模以及领域驱动设计; 2、熟悉HBase/Hadoop/Spark/Hive/Flink等大数据工具等,具备丰富的海量数据加工处理和优化经验; 3、熟悉Java/Python/Scala等至少一门语言,有优秀…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
ETL+
https://www.ibm.com/think/topics/etl
ETL—meaning extract, transform, load—is a data integration process that combines, cleans and organizes data from multiple sources into a single, consistent data set for storage in a data warehouse, data lake or other target system.
https://www.youtube.com/watch?v=OW5OgsLpDCQ
It explains what ETL is and what it can do for you to improve your data analysis and productivity.
DDD+
https://ddd-crew.github.io/ddd-starter-modelling-process/
This process gives you a step-by-step guide for learning and practically applying each aspect of Domain-Driven Design (DDD) - from orienting around an organisation’s business model to coding a domain model.
[英文] Domain Driven Design
https://medium.com/@matteopampana/list/domain-driven-design-c1efaabe287e
Everyone talks about DDD, but how many understand and correctly apply Domain-Driven Design? I want to be one of them.
https://redis.io/glossary/domain-driven-design-ddd/
Domain-Driven Design (DDD) is a software development philosophy that emphasizes the importance of understanding and modeling the business domain.
HBase+
[英文] HBase Tutorial
https://www.tutorialspoint.com/hbase/index.htm
HBase is a data model that is similar to Google's big table designed to provide quick random access to huge amounts of structured data. This tutorial provides an introduction to HBase, the procedures to set up HBase on Hadoop File Systems, and ways to interact with HBase shell.
Hadoop+
https://www.runoob.com/w3cnote/hadoop-tutorial.html
Hadoop 为庞大的计算机集群提供可靠的、可伸缩的应用层计算和存储支持,它允许使用简单的编程模型跨计算机群集分布式处理大型数据集,并且支持在单台计算机到几千台计算机之间进行扩展。
[英文] Hadoop Tutorial
https://www.tutorialspoint.com/hadoop/index.htm
Hadoop is an open-source framework that allows to store and process big data in a distributed environment across clusters of computers using simple programming models.
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Hive+
[英文] Hive Tutorial
https://www.tutorialspoint.com/hive/index.htm
Hive is a data warehouse infrastructure tool to process structured data in Hadoop. It resides on top of Hadoop to summarize Big Data, and makes querying and analyzing easy.
https://www.youtube.com/watch?v=D4HqQ8-Ja9Y
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
还有更多 •••
相关职位
社招5年以上产品-产品解决方
1. 了解数据行业,有3年以上数据公司、风控咨询、信贷机构或互联网金融行业背景; 2. 有征信、信用管理经验优先,熟悉数据评分设计、风险咨询、风控策略、信用风险模型优先; 3. 金融、统计、经济、数学
更新于 2026-07-27杭州
社招5年以上解决方案-产品解
1. 了解数据行业,有3年以上数据公司、风控咨询、信贷机构或互联网金融行业背景; 2. 有征信、信用管理经验优先,熟悉数据评分设计、风险咨询、风控策略、信用风险模型优先; 3. 金融、统计、经济、数学
更新于 2025-12-15杭州
社招3年以上技术类-开发
1. 计算机、软件工程等相关专业,本科及以上学历。 2. 具备较强的产品意识和业务理解能力,能够将相对模糊的业务诉求转化为明确的Agent目标、产品方案和技术实现路径,具备实际的Agent落地经验。
更新于 2026-08-07杭州