千问千问事业部-高级数据研发专家-大模型语料
社招全职3年以上技术类-开发地点:北京 | 杭州状态:招聘
工作描述
任职要求 1. 主导设计和落地过大规模AI语料数据处理平台或数据生产系统,具备从0到1或大规模演进语料数据体系的经验,能够从工程架构层面构建支撑语料数据全生命周期的系统能力; 2. 具备大规模数据处理经验,熟悉网页、文档、文本、图片及音视频等多模态语料数据处理技术,对语料清洗、去重、过滤、结构化、对齐及质量评估等关键流程有深入理解; 3. 熟悉分布式数据计算与存储技术,如 Ray、Spark、Flink、Paimon 等,具备大规模数据处理系统设计与性能优化经验,能够与AI Infra及基础数据平台团队协同推进能力落地; 4. 熟悉数据分类与质量评价体系,具备语料数据治理与数据画像建设经验,能够通过数据分析与评估指导语料数据建设与优化; 5. 深刻理解并熟练掌握 AI Agent 的核心技术范式,熟悉主流Agent开发框架,如LangChain、Open…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
大数据+
https://www.youtube.com/watch?v=bAyrObl7TYE
https://www.youtube.com/watch?v=H4bf_uuMC-g
With all this talk of Big Data, we got Rebecca Tickle to explain just what makes data into Big Data.
性能调优+
https://goperf.dev/
The Go App Optimization Guide is a series of in-depth, technical articles for developers who want to get more performance out of their Go code without relying on guesswork or cargo cult patterns.
https://web.dev/learn/performance
This course is designed for those new to web performance, a vital aspect of the user experience.
https://www.ibm.com/think/insights/application-performance-optimization
Application performance is not just a simple concern for most organizations; it’s a critical factor in their business’s success.
https://www.oreilly.com/library/view/optimizing-java/9781492039259/
Performance tuning is an experimental science, but that doesn’t mean engineers should resort to guesswork and folklore to get the job done.
Ray+
https://github.com/ray-project/ray
Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
https://www.youtube.com/watch?v=FhXfEXUUQp0
In this video, I'll teach you everything you need to know about Apache Ray!
https://www.youtube.com/watch?v=fMiAyj2kgac
Using powerful machine learning algorithms is easy using Ray.io and Python.
https://www.youtube.com/watch?v=q_aTbb7XeL4
Parallel and Distributed computing sounds scary until you try this fantastic Python library.
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Flink+
https://nightlies.apache.org/flink/flink-docs-release-2.0/docs/learn-flink/overview/
This training presents an introduction to Apache Flink that includes just enough to get you started writing scalable streaming ETL, analytics, and event-driven applications, while leaving out a lot of (ultimately important) details.
https://www.youtube.com/watch?v=WajYe9iA2Uk&list=PLa7VYi0yPIH2GTo3vRtX8w9tgNTTyYSux
Today’s businesses are increasingly software-defined, and their business processes are being automated. Whether it’s orders and shipments, or downloads and clicks, business events can always be streamed. Flink can be used to manipulate, process, and react to these streaming events as they occur.
数据治理+
https://www.ibm.com/think/topics/data-governance
Data governance is the data management discipline that focuses on the quality, security and availability of an organization’s data.
https://www.youtube.com/watch?v=uPsUjKLHLAg
Building data fabric eliminates the technological complexities of data governance so users can connect to the right data at the right time, regardless of where it resides.
还有更多 •••
相关职位

社招3年以上技术类-开发
1. 主导过LLM、VLM、ASR或TTS大模型预训练及微调语料数据建设工作,有丰富的数据交付经验; 2. 精通大规模分布式数据处理技术(如spark/flink/ray等),拥有从0到1搭建全模态数
更新于 2026-04-06北京|杭州
社招3年以上技术类-开发
1. 主导过LLM、VLM、ASR或TTS大模型预训练及微调语料数据建设工作,有丰富的数据交付经验; 2. 精通大规模分布式数据处理技术(如spark/flink/ray等),拥有从0到1搭建全模态数
更新于 2026-02-06北京|杭州
社招5年以上技术类-数据
1. 学历与背景 - 计算机科学、软件工程、数据科学或相关专业硕士及以上学历(优秀者可放宽); - 熟悉分布式系统原理、数据库原理及大数据生态技术栈。 2. 技术能力 - 精通至
更新于 2025-12-10上海