拼多多搜广推数据开发工程师
社招全职技术类地点:上海状态:招聘
任职要求
1、精通数据仓库模型设计,具备丰富的ETL开发及海量数据加工处理经验; 2、具备丰富的Flink实时计算开发与调优经验,精通高并发、高可用、可扩展分布式系统设计原则与实践; 3、具备分布式数据存储与计算平台(Hadoop生态)应用开发经…
登录查看完整任职要求
微信扫码,1秒登录
工作职责
1、负责电商搜索、广告、推荐等核心业务数据体系及相关工具平台的规划、建设与持续优化,高效支持算法、数据分析、工程等团队的数据需求; 2、深入理解业务逻辑,抽象业务需求并设计可扩展、高性能的数据技术架构,快速响应业务变化,构建高效、可靠的数据互通与共享机制; 3、负责数据处理链路(离线/实时)的日常运维、监控与保障,确保数据稳定、高效产出,及时解决数据问题。
包括英文材料
数据仓库+
https://www.youtube.com/watch?v=9GVqKuTVANE
From Zero to Data Warehouse Hero: A Full SQL Project Walkthrough and Real Industry Experience!
https://www.youtube.com/watch?v=k4tK2ttdSDg
ETL+
https://www.ibm.com/think/topics/etl
ETL—meaning extract, transform, load—is a data integration process that combines, cleans and organizes data from multiple sources into a single, consistent data set for storage in a data warehouse, data lake or other target system.
https://www.youtube.com/watch?v=OW5OgsLpDCQ
It explains what ETL is and what it can do for you to improve your data analysis and productivity.
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
高并发+
https://www.baeldung.com/concurrency-principles-patterns
In this tutorial, we’ll discuss some of the design principles and patterns that have been established over time to build highly concurrent applications.
https://www.baeldung.com/java-concurrency
Handling concurrency in an application can be a tricky process with many potential pitfalls. A solid grasp of the fundamentals will go a long way to help minimize these issues.
https://www.oreilly.com/library/view/concurrency-in-go/9781491941294/
You’ll understand how Go chooses to model concurrency, what issues arise from this model, and how you can compose primitives within this model to solve problems.
https://www.oreilly.com/library/view/modern-concurrency-in/9781098165406/
With this book, you'll explore the transformative world of Java 21's key feature: virtual threads.
https://www.youtube.com/watch?v=qyM8Pi1KiiM
https://www.youtube.com/watch?v=wEsPL50Uiyo
高可用+
https://redis.io/blog/high-availability-architecture/
A high available architecture is when there are a number of different components, modules, or services that work together to maintain optimal performance, irrespective of peak-time loads.
https://www.ibm.com/think/topics/high-availability
High availability (HA) is a term that refers to a system’s ability to be accessible and reliable close to 100% of the time.
还有更多 •••
相关职位
实习引擎
1. 参与搜索、广告、推荐系统的数据处理工作,协助团队完成数据清洗、预处理和特征工程; 2. 协助开发和优化大数据处理流程,提升数据处理效率和质量; 3. 学习和研究最新的大数据处理技术和工具,为团队带来新的思路和方法。
更新于 2026-06-30上海|北京
实习机器学习平台
【业务介绍】 作为公司统一的模型训练引擎团队,支撑公司内所有搜推广类业务的训练工程侧工作,包括模型训练、参数服务器、特征样本流水线等,通过引擎能力的持续建设结合多元异构算力为业务提供高效、灵活、稳定的搜广推模型服务。 你将专注于大规模AI训练系统最核心的性能优化赛道,直面千亿参数模型训练中的效率瓶颈,解决工业级AI系统在性能与规模上面临的真实挑战。 【岗位职责】 1、深入参与GPU异构计算栈的研发与调优,从算子、内存、通信多维度挖掘硬件极限性能;通过CUDA编程、内核融合、混合精度训练、通信与计算重叠等高级优化技术,不断提升训练引擎效率。 2、推动自动化扩展、智能资源调度、跨架构设备兼容(NV GPU、GPGPU、XPU等)、AI系统可观测性等先进技术在公司模型训练平台落地; 3、跟踪并推动AI系统领域的最新技术趋势(如生成式推荐、AI编译优化、RDMA/NCCL通信计算并发等),持续保持平台业界领先优势。
更新于 2026-02-12上海
社招技术类
1、负责搜广推业务系统和索引系统架构设计、核心模块开发与维护,支撑数十亿级商品库的高效检索; 2、优化在线服务的性能与稳定性,满足高并发场景下的低延迟、高吞吐等需求; 3、优化引擎的索引构建、检索性能、排序效率,持续提升系统吞吐能力和响应速度; 4、负责召回、粗排、精排、重排等全链路性能优化; 5、与算法团队协作,将排序模型、策略设计、产品需求等技术方案工程化落地; 6、深刻理解业务,抽象和设计合理的技术架构,以适应不断变化的需求。
更新于 2026-06-05上海