快手数据采集技术专家(Tech Lead)-【大模型专项】
社招全职5-10年J0012地点:北京状态:招聘
工作描述
任职要求 1、本科及以上学历,计算机相关专业;8年以上互联网后端/数据工程经验,其中3年以上采集或爬虫系统相关工作经验; 2、精通 Java 或 Python,具备大规模分布式系统设计与落地能力,有高吞吐、高可用架构实战经验; 3、深入理解主流采集技术栈(Scrapy、Puppeteer、Playwright、Frida 等),熟悉数据调度框架(如 Airflow、自研调度系统); 4、熟悉主流互联网风控策略与反爬机制,具备完整的对抗实战经验与方案沉淀能力; 5、具备良好的技术判断力和项目推…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
Java+
https://www.youtube.com/watch?v=eIrMbAQSU34
Master Java – a must-have language for software development, Android apps, and more! ☕️ This beginner-friendly course takes you from basics to real coding skills.
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
分布式系统+
https://www.distributedsystemscourse.com/
The home page of a free online class in distributed systems.
https://www.youtube.com/watch?v=7VbL89mKK3M&list=PLOE1GTZ5ouRPbpTnrZ3Wqjamfwn_Q5Y9A
高可用+
https://redis.io/blog/high-availability-architecture/
A high available architecture is when there are a number of different components, modules, or services that work together to maintain optimal performance, irrespective of peak-time loads.
https://www.ibm.com/think/topics/high-availability
High availability (HA) is a term that refers to a system’s ability to be accessible and reliable close to 100% of the time.
Puppeteer+
https://oxylabs.io/blog/puppeteer-tutorial
There are a few methods to accessing and parsing web pages, but in this tutorial we will be covering how to do it with Google Puppeteer.
[英文] Getting started
https://pptr.dev/guides/getting-started
You launch/connect a browser, create some pages, and then manipulate them with Puppeteer's API.
https://www.youtube.com/watch?v=nIJV-LbV_vM
This tutorial walks you through every thing you need to know about Puppeteer and headless browsers, so you can automate website testing, web scraping, fetching and downloading content, and more.
https://www.youtube.com/watch?v=Sag-Hz9jJNg
Learn puppeteer in less than one hour.
还有更多 •••
相关职位
社招2年以上技术类-数据
1. 5 年以上后端、爬虫平台、搜索引擎、数据平台、数据治理、安全合规或 AI 数据工程经验,有复杂系统设计和跨团队项目推进经验。 2. 熟悉大规模网页采集系统,理解 URL 调度、网页解析、内容抽取
更新于 2026-07-28杭州
社招3-5年J0012
1、本科及以上学历,计算机相关专业,有强烈的好奇心和技术敏锐度,对AI大模型和采集相关技术有浓厚的兴趣; 2、熟悉Java、Python等语言,具备扎实的编码能力;熟悉主流采集技术及框架工具,如Fri
更新于 2026-06-15北京
社招3-5年J0012
1、本科及以上学历,计算机相关专业,有强烈的好奇心和技术敏锐度,对AI大模型和采集相关技术有浓厚的兴趣; 2、熟悉Java、Python等语言,具备扎实的编码能力;熟悉主流采集技术及框架工具,如Fri
更新于 2026-06-18北京