蚂蚁金服【蚂蚁星】多模态理解与生成评测算法研究-27届
校招全职2027届蚂蚁星- Plan A人才计划地点:北京 | 杭州状态:招聘
工作描述
Qualifications 1. 计算机科学、人工智能、机器学习、语音、多模态、计算机视觉、图形学等相关专业硕士及以上学历,博士或具备同等研究能力者优先; 2. 熟练使用 Python,熟悉 PyTorch 或其他深度学习框架,能够独立完成算法实验、评测原型和数据分析; 3. 熟悉以下一个或多个方向:Multimodal LLM / Omni-modal Model、Vision-Language Model、Speech LM / Audio LM、Streaming ASR / Streaming TTS、Audio-Visual Understanding、Digital Human / Virtual Avatar、Talking Head / Full-body Avatar、Motion Generation / 3D Facial Animation、Neural Rendering / 3D Gaussian Splatting、Video Diffusion / Diffusion Transformer; 4. 具备较强的实验设计和问题抽象能力,能够把“自然不自然”“像不像真人”“交互顺不顺”这类主观体验转化为可验证的评测指标; 5. 对多模态前沿技术有持续兴趣,能够快速阅读论文、复现方法,并结合业务场景形成可落地方案; 6. 具备良好的工程能力和沟通能力,能够与算法、产品、工程团队协作推动评测体系落地。 加分项: 1. 在 ICLR、NeurIPS、ICML、ACL、EMNLP、CVPR、ICCV、ECCV、SIGGRAPH、AAAI 等会议发表过高质量论文; 2. 有多模态大模型、语音大模型、数字人、视频生成、3D 生成、虚拟环境或世界模型相关项目经验; 3. 参与过 Benchmark、评测框架、数据集、开源项目或算法竞赛; 4. 有 VLM-as-Judge、LLM-as-Judge、…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
机器学习+
https://www.youtube.com/watch?v=0oyDqO8PjIg
Learn about machine learning and AI with this comprehensive 11-hour course from @LunarTech_ai.
https://www.youtube.com/watch?v=i_LwzRVP7bg
Learn Machine Learning in a way that is accessible to absolute beginners.
https://www.youtube.com/watch?v=NWONeJKn6kc
Learn the theory and practical application of machine learning concepts in this comprehensive course for beginners.
https://www.youtube.com/watch?v=PcbuKRNtCUc
Learn about all the most important concepts and terms related to machine learning and AI.
OpenCV+
https://learnopencv.com/getting-started-with-opencv/
At LearnOpenCV we are on a mission to educate the global workforce in computer vision and AI.
https://opencv.org/university/free-opencv-course/
This free OpenCV course will teach you how to manipulate images and videos, and detect objects and faces, among other exciting topics in just about 3 hours.
学历+
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
数据分析+
[英文] Data Analyst Roadmap
https://roadmap.sh/data-analyst
Step by step guide to becoming an Data Analyst in 2025
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
语音识别+
https://developer.nvidia.com/blog/essential-guide-to-automatic-speech-recognition-technology/
Over the past decade, AI-powered speech recognition systems have slowly become part of our everyday lives, from voice search to virtual assistants in contact centers, cars, hospitals, and restaurants.
语音合成+
https://www.ibm.com/think/topics/text-to-speech
Text to speech (TTS) is a type of technology that converts text on a digital interface into natural-sounding audio.
Framer Motion+
https://motion.dev/docs/quick-start
Motion is an animation library that's easy to start and fun to master.
https://www.youtube.com/watch?v=znbCa4Rr054
Framer Motion is not only the simplest way to get up and running with animations in React JS, but also one of the most powerful.
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
ICLR+
https://iclr.cc/
NeurIPS+
https://neurips.cc/
还有更多 •••
相关职位
实习蚂蚁星- Pla
Qualifications 1. 计算机科学、人工智能、机器学习、语音、多模态、计算机视觉、图形学等相关专业硕士及以上学历,博士或具备同等研究能力者优先; 2. 熟练使用 Python,熟悉 PyT
更新于 2026-08-26北京|杭州
实习蚂蚁星- Pla
1. 计算机科学、人工智能、数学等相关专业硕士及以上学历,博士优先; 2. 深入掌握Transformer/BERT/GPT等架构,有1个以上千亿参数大模型实战经验(训练/推理/优化全流程); 3.
更新于 2025-07-25北京|上海|杭州
实习蚂蚁星- Pla
1. 本科及以上学历,计算机相关专业,多模态算法相关工作经验; 2. 熟练掌握计算机视觉领域的基础理论和方法,熟悉PyTorch等主流深度学习框架,能够独立实现前沿模型; 3. 有良好的自我学习能力及
更新于 2025-07-25北京|上海|杭州