
智能互联通义实验室-大模型数据处理与优化算法工程师-Qwen
社招全职3年以上技术类-算法地点:北京 | 杭州状态:招聘
工作描述
任职要求 1. 计算机科学、人工智能、数学、物理或相关领域博士/顶尖硕士毕业生。 2. 熟练掌握Python,熟悉SQL及数据库操作;熟悉分布式计算框架(如Spark、Hadoop、Ray);熟悉常见分类模型及深度学习训练 微调 与推理框架(如transformer bert gpt, pytorch , vllm sglang)。 3. 具备大规模数据处理经验,能够高效完成数据清洗与转换任务。 4. 学习能力强,动手能力突出,能快速上手新工具和技术。 5. 具备跨域视野与协作意识,能够与其他团队紧密合作,实现数据平台共建。 加分项: 1. 有大模型相关数据收集处理清洗经验,有处理千亿级以上数据的经验。 2. 有阿里云服务使用经验,如MaxCompute、Function Compute、OSS等。 3. 掌握HTML2Text、PDF2Text、OCR…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
SQL+
https://liaoxuefeng.com/books/sql/introduction/index.html
什么是SQL?简单地说,SQL就是访问和处理关系数据库的计算机标准语言。
https://sqlbolt.com/
Learn SQL with simple, interactive exercises.
https://www.youtube.com/watch?v=p3qvj9hO_Bo
In this video we will cover everything you need to know about SQL in only 60 minutes.
Spark+
[英文] Learning Spark Book
https://pages.databricks.com/rs/094-YMS-629/images/LearningSpark2.0.pdf
This new edition has been updated to reflect Apache Spark’s evolution through Spark 2.x and Spark 3.0, including its expanded ecosystem of built-in and external data sources, machine learning, and streaming technologies with which Spark is tightly integrated.
Hadoop+
https://www.runoob.com/w3cnote/hadoop-tutorial.html
Hadoop 为庞大的计算机集群提供可靠的、可伸缩的应用层计算和存储支持,它允许使用简单的编程模型跨计算机群集分布式处理大型数据集,并且支持在单台计算机到几千台计算机之间进行扩展。
[英文] Hadoop Tutorial
https://www.tutorialspoint.com/hadoop/index.htm
Hadoop is an open-source framework that allows to store and process big data in a distributed environment across clusters of computers using simple programming models.
Ray+
https://github.com/ray-project/ray
Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
https://www.youtube.com/watch?v=FhXfEXUUQp0
In this video, I'll teach you everything you need to know about Apache Ray!
https://www.youtube.com/watch?v=fMiAyj2kgac
Using powerful machine learning algorithms is easy using Ray.io and Python.
https://www.youtube.com/watch?v=q_aTbb7XeL4
Parallel and Distributed computing sounds scary until you try this fantastic Python library.
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
还有更多 •••
相关职位
社招1年以上技术类-算法
1. 计算机、机器学习等方向相关专业,博士及硕士优先。 2. 具有 post-training 或强化学习相关方向经验。 3. 精通 Python 以及 Pytorch 等深度学习框架,具有较强的代码
更新于 2026-04-02北京|杭州|上海
社招1年以上技术类-算法
1. 计算机科学、人工智能、机器学习等领域的博士/硕士毕业生。 2. 精通Python、C/C++等至少一门编程语言。 3. 有良好的学术调研能力,工程能力,逻辑和数据分析能力,热衷于Agentic
更新于 2026-04-02北京|杭州

社招1年以上技术类-算法
1. 计算机科学、人工智能等相关专业硕士及以上学历,1 年以上大模型研发相关工作经验。 2. 了解 LLM 常见的能力维度(推理、知识、指令遵循、安全、多轮对话等)及对应的评测策略。 3. 具备大模型
更新于 2026-04-07北京|杭州