蚂蚁金服钱塘征信-数据研发工程师 (数仓/大数据/ETL开发)-征信方向
社招全职3年以上技术类-数据地点:杭州状态:招聘
任职要求
1、有较为丰富的数仓设计&开发经验,熟悉ETL分层建设方法、数据、维度建模以及领域驱动设计; 2、熟悉HBase/Hadoop/Spark/Hive/Flink等大数据工具等,具备丰富的海量数据加工处理和优化经验; 3、熟悉Ja…
登录查看完整任职要求
微信扫码,1秒登录
工作职责
1、围绕征信业务目标,通过数据分析洞察、构建分析工具,为业务目标达成提高效率; 2、探索大模型等技术在征信数据场景的应用,推动数据分析和挖掘的效能革新; 3、参与公司数据体系建设、实时和离线数仓设计、数据模型体系的构建和开发; 4、负责数据治理,建立数据规范,优化数据链路,保证数据时效和数据质量; 5、通过数据仓库的建设和治理,实现数据产品化,能够针对业务场景探索提供大数据解决方案等。
包括英文材料
ETL+
https://www.ibm.com/think/topics/etl
ETL—meaning extract, transform, load—is a data integration process that combines, cleans and organizes data from multiple sources into a single, consistent data set for storage in a data warehouse, data lake or other target system.
https://www.youtube.com/watch?v=OW5OgsLpDCQ
It explains what ETL is and what it can do for you to improve your data analysis and productivity.
DDD+
https://ddd-crew.github.io/ddd-starter-modelling-process/
This process gives you a step-by-step guide for learning and practically applying each aspect of Domain-Driven Design (DDD) - from orienting around an organisation’s business model to coding a domain model.
[英文] Domain Driven Design
https://medium.com/@matteopampana/list/domain-driven-design-c1efaabe287e
Everyone talks about DDD, but how many understand and correctly apply Domain-Driven Design? I want to be one of them.
https://redis.io/glossary/domain-driven-design-ddd/
Domain-Driven Design (DDD) is a software development philosophy that emphasizes the importance of understanding and modeling the business domain.
HBase+
[英文] HBase Tutorial
https://www.tutorialspoint.com/hbase/index.htm
HBase is a data model that is similar to Google's big table designed to provide quick random access to huge amounts of structured data. This tutorial provides an introduction to HBase, the procedures to set up HBase on Hadoop File Systems, and ways to interact with HBase shell.
Hadoop+
https://www.runoob.com/w3cnote/hadoop-tutorial.html
Hadoop 为庞大的计算机集群提供可靠的、可伸缩的应用层计算和存储支持,它允许使用简单的编程模型跨计算机群集分布式处理大型数据集,并且支持在单台计算机到几千台计算机之间进行扩展。
[英文] Hadoop Tutorial
https://www.tutorialspoint.com/hadoop/index.htm
Hadoop is an open-source framework that allows to store and process big data in a distributed environment across clusters of computers using simple programming models.
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Hive+
[英文] Hive Tutorial
https://www.tutorialspoint.com/hive/index.htm
Hive is a data warehouse infrastructure tool to process structured data in Hadoop. It resides on top of Hadoop to summarize Big Data, and makes querying and analyzing easy.
https://www.youtube.com/watch?v=D4HqQ8-Ja9Y
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
还有更多 •••
相关职位
社招5年以上解决方案-产品解
1.负责征信产品的设计,识别市场机会,挖掘及分析机构客户需求,设计及制定包括评分、画像、报告等一系列征信产品的数据标准、产品需求及解决方案。 2.负责征信公司内外部数据价值的挖掘、评估,及应用方向规划,明确数据挖掘方向,建立数据资产价值体系。 3.负责征信产品的管理及运营,配合商务团队制定产品的推广计划,同时负责产品增值服务的设计,持续优化产品提升客户体验。
更新于 2025-12-15杭州
社招3年以上技术类-开发
负责公司业务和平台核心系统的设计、研发以及维护工作。 1. 负责复杂商业产品、大规模高并发系统的设计和开发; 2. 承担系统迭代优化任务,提升业务扩展性、系统稳定性、响应性能等; 3. 主导相关系统分析与设计工作,承担核心功能或组件的代码编写; 4. 参与团队的稳定性建设和安全生产,保障日常大促的系统稳定和线上问题的预警监控以及应急解决方案。
更新于 2025-11-03杭州
社招5年以上解决方案-产品解
1.为金融机构提供数据产品使用的增值服务,深入研究三方数据的使用场景,提供贷前、贷中、贷后全生命周期的数据产品评测服务,同时推荐合适的三方数据产品和使用建议; 2.深入研究数据产品的性能和联系,为金融机构提供成本、性能等多目标优化下最佳的数据产品调用组合推荐; 3.协同内部各部门同学,与数据、算法、产品、技术同学紧密合作,内部通过产品化提升工作效率,外部作为一个整体,为金融机构提供高效、高质的服务; 4.基于对数据产品的行业认知,进行内部培训和外部交流,提升公司的内部人员认知和公司的行业影响力。
更新于 2026-07-07杭州