千问千问事业部-大模型语料数据研发专家-杭州/北京
社招全职3年以上技术类-开发地点:北京 | 杭州状态:招聘
工作描述
任职要求 1. 具备大规模数据处理经验,熟悉网页、文档、文本、图片及音视频等多模态语料数据处理技术,对语料清洗、去重、过滤、结构化、对齐及质量评估等关键流程有深入理解; 2. 熟悉分布式数据计算与存储技术,如 Ray、Spark、Flink、Paimon 等,具备大规模数据处理系统设计与性能优化经验,能够与AI Infra及基础数据平台团队协同推进能力落地; 3. 熟悉数据分类与质量评价体系,具备语料数据治理与数据画像建设经验,能够通过数据分析与评估指导语料数据建设与优化; 4. 熟悉主流Agent开发框架,如L…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
Ray+
https://github.com/ray-project/ray
Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
https://www.youtube.com/watch?v=FhXfEXUUQp0
In this video, I'll teach you everything you need to know about Apache Ray!
https://www.youtube.com/watch?v=fMiAyj2kgac
Using powerful machine learning algorithms is easy using Ray.io and Python.
https://www.youtube.com/watch?v=q_aTbb7XeL4
Parallel and Distributed computing sounds scary until you try this fantastic Python library.
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
数据治理+
https://www.ibm.com/think/topics/data-governance
Data governance is the data management discipline that focuses on the quality, security and availability of an organization’s data.
https://www.youtube.com/watch?v=uPsUjKLHLAg
Building data fabric eliminates the technological complexities of data governance so users can connect to the right data at the right time, regardless of where it resides.
还有更多 •••
相关职位
社招3年以上技术类-开发
1. 主导设计和落地过大规模AI语料数据处理平台或数据生产系统,具备从0到1或大规模演进语料数据体系的经验,能够从工程架构层面构建支撑语料数据全生命周期的系统能力; 2. 具备大规模数据处理经验,熟悉
更新于 2026-07-01北京|杭州
社招3年以上技术类-算法
1. 硕士及以上学历,计算机科学、人工智能或相关专业背景。 2. 熟练掌握机器学习、自然语言处理、大语言模型等相关领域的基本理论和算法,具备扎实的数学基础。 3. 熟练掌握Python编程语言,熟悉主
更新于 2026-09-03北京|杭州
校招2027届蚂蚁星
1. 热爱人工智能领域,对探索新事物充满热情; 2. 硕士及以上学历,计算机科学、人工智能或相关专业背景; 3. 熟练掌握机器学习、自然语言处理、大语言模型等相关领域的基本理论和算法,具
更新于 2026-05-14北京|杭州|成都