深度求索预训练数据工程师
社招全职全栈开发/算法地点:北京 | 杭州状态:招聘
工作描述
【团队使命】 预训练数据开采自人类世界全部已有知识所沉积的富矿,是驱动模型强大能力的核心燃料。 预训练数据团队围绕大模型预训练,构建了从数据采集、清洗处理到底层基础设施的完整数据体系,持续为模型训练输送高质量、规模化、多模态的数据燃料,不断拓展模型的世界知识边界。 如何寻找蒙尘的富矿,并淘洗炼制成可用的智能燃料,是一门极难的硬功夫。我们以亿万数据筑就模型智能的坚实底座,更要用普惠的方式,把这份属于全人类的财富,重新还给全人类。 招聘方向【实习/全职】:数据采集 pipeline 方向、语言数据处理方向、多模态数据方向、数据基建方向。 数据采集 pipeline 方向 【岗位职责】 1.全网数据选取策略设计开发 (1)根据数据清洗反馈信号设计全网数据选取、调度系统,包括链接质量预估、属性聚合、站点抓取配额、数据优先级等,保证数据多样性,最大化数据覆盖,为模型训练提供丰富的世界知识。 (2)面向 RAG 等时效场景,负责时效性种子挖掘、调度与有效性验证,保障关键信息的新鲜度。 2.全网数据采集环路 (1)负责全网数据采集环路的设计与开发,包括但不限于下载、任务调度、流式数据处理、DNS、连通性等核心组件,构建高吞吐、可容错的分布式采集系统。 3.链接规模控制 (1)通过识别低质站群、重复镜像站点、清洗 URL 冗余参数、规范化链接等手段,控制链接规模、去除重复与噪声,提升链接库的有效性与数据纯净度。 【岗位要求】 1.计算机或数学相关专业,本科及以上学历。 2.熟练使用至少一种编程语言,包括但不限于 Python/C++/Rust。 【加分项】 1.数学或信息学竞赛中取得优秀成绩。 2.有大规模数据采集、爬虫或数据 pipeline 系统研发经验。 语言数据处理方向 【岗位职责】 1.参与大模型预训练语料清洗框架研发,建设稳定、高效、可扩展的大规…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
学历+
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
C+++
https://www.learncpp.com/
LearnCpp.com is a free website devoted to teaching you how to program in modern C++.
https://www.youtube.com/watch?v=ZzaPdXTrSb8
Rust+
https://www.youtube.com/watch?v=BpPEoZW5IiY
In this comprehensive Rust course for beginners, you will learn about the core concepts of the language and underlying mechanisms in theory.
https://www.youtube.com/watch?v=lzKeecy4OmQ
Full Rust 101 Crash Course for beginners.
https://www.youtube.com/watch?v=rQ_J9WH6CGk
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
RAG+
https://www.youtube.com/watch?v=sVcwVQRHIc8
Learn how to implement RAG (Retrieval Augmented Generation) from scratch, straight from a LangChain software engineer.
Pandas+
[英文] 10 minutes to pandas
https://pandas.pydata.org/docs/user_guide/10min.html
This is a short introduction to pandas, geared mainly for new users.
[英文] Cookbook - pandas
https://pandas.pydata.org/docs/user_guide/cookbook.html#cookbook
This is a repository for short and sweet examples and links for useful pandas recipes.
https://www.kaggle.com/learn/pandas
Solve short hands-on challenges to perfect your data manipulation skills.
https://www.youtube.com/watch?v=2uvysYbKdjM
I'm super excited for this one. We're doing another complete Python Pandas tutorial walkthrough.
https://www.youtube.com/watch?v=Mdq1WWSdUtw
Filtering, Joins, Indexing, Data Cleaning, Visualizations
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
数据结构+
https://www.youtube.com/watch?v=8hly31xKli0
In this course you will learn about algorithms and data structures, two of the fundamental topics in computer science.
https://www.youtube.com/watch?v=B31LgI4Y4DQ
Learn about data structures in this comprehensive course. We will be implementing these data structures in C or C++.
https://www.youtube.com/watch?v=CBYHwZcbD-s
Data Structures and Algorithms full course tutorial java
还有更多 •••
相关职位
社招技术类/Tech
岗位职责 设计和优化大规模 Web 爬虫的 URL 发现与调度系统,持续扩大数据覆盖面 建设多因子抓取优先级体系,在有限资源下最大化高质量页面的获取效率 主导反爬对抗和动态渲染方案,提升核心站点的抓取
更新于 2026-06-09北京
社招研发类
1)扎实的编程能力(Python必须,Java / Golang / Javascript 至少一种),良好的数据结构和算法基础; 2)熟悉软件工程基本流程(构建 / 调试 / 测试 / CI); 3
更新于 2026-09-08上海
社招
【岗位职责】 1. 参与大模型预训练数据建设,围绕图文合训、多模态理解及视觉能力训练,探索高质量数据构建和规模化扩展方案。 2. 负责图文及视觉数据的分析、清洗、去重、筛选和结构化处理,提升训练数据的
更新于 2026-09-03北京
