爱奇艺视频生成算法工程师
社招全职算法地点:北京状态:招聘
任职要求
-学历背景: 计算机、人工智能、数学等相关专业硕士及以上学历,具有扎实的深度学习,大模型理论基础; -深入理解Diffusion Model、DiT (Diffusion Transformer)、VAE等核心架构; -熟悉主流视频生成开源模型(如Stable Video Diffusion, Wan,AnimateDiff, CogVideo等),有实际的SFT或LoRA微调经验; -熟练掌握PyTorch,熟悉Diffusers、DeepSpeed、Accelerate等训练/推理框架; -熟悉LLaMA-Factory等…
登录查看完整任职要求
微信扫码,1秒登录
工作职责
将最前沿的生成式AI技术应用于长视频内容的生产与创新。我们正在构建行业领先的文生视频与图生视频大模型,旨在通过AI重塑影视剧本可视化、智能剪辑与特效制作流程。你将有机会利用海量的高质量影视数据,训练懂镜头、懂叙事的视频生成模型。 【岗位职责】 -核心模型研发: 负责文生视频、图生视频大模型的SFT(有监督微调)与LoRA训练策略设计。基于爱奇艺独有的影视级数据,优化模型在特定风格、角色一致性及动作流畅度上的表现; -对齐与强化学习: 探索并应用RLHF/RLAIF技术于视频生成领域,利用PPO、DPO或GRPO等算法,提升模型对指令的遵循能力及视频生成的美学质量; -可控生成探索: 研发基于ControlNet、Adapter等技术的可控视频生成算法,解决视频生成中的一致性、运镜控制及物理规律遵循等痛点,满足专业内容生产(PGC)需求; -工程与落地: 负责算法模型在业务场景(如智能预告片、脚本转视频)的落地与效果迭代。
包括英文材料
学历+
深度学习+
https://d2l.ai/
Interactive deep learning book with code, math, and discussions.
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
LLaMA-Factory+
https://llamafactory.readthedocs.io/en/latest/
LLaMA Factory is an easy-to-use and efficient platform for training and fine-tuning large language models.
AI agent+
https://www.ibm.com/think/ai-agents
Your one-stop resource for gaining in-depth knowledge and hands-on applications of AI agents.
强化学习+
https://cloud.google.com/discover/what-is-reinforcement-learning?hl=en
Reinforcement learning (RL) is a type of machine learning where an "agent" learns optimal behavior through interaction with its environment.
https://huggingface.co/learn/deep-rl-course/unit0/introduction
This course will teach you about Deep Reinforcement Learning from beginner to expert. It’s completely free and open-source!
https://www.kaggle.com/learn/intro-to-game-ai-and-reinforcement-learning
Build your own video game bots, using classic and cutting-edge algorithms.
还有更多 •••
相关职位
社招3-5年策略算法
岗位介绍 我们正在探索流式视频生成大模型方向,你将参与构建面向真实世界的实时视频处理系统,覆盖TV2V、TIV2V、TI2V等前沿课题,推动模型能力从离线处理走向实时交互式处理。 1、深度参与流式视频编辑大模型的研究与落地,涵盖超分辨率、指令风格化、参考图风格化、特效生成、交互式图片等核心任务,打造业界领先的视频编辑能力; 2、探索Diffusion、自回归等生成范式在流式视频编辑领域的前沿技术,包括但不限于高效推理、时序一致性建模、多条件可控生成等方向,推动技术创新与工程落地; 3、研究Consistency Model、DMD、GAN蒸馏等加速方法在视频编辑场景的应用,解决工业级实时性与画质的平衡难题,支持大规模线上服务; 4、参与高质量视频编辑数据的挖掘、清洗与合成,结合Post-training、强化学习等方法持续优化模型效果与泛化能力; 5、跟踪视频生成与编辑领域最新进展,将相关研究总结为顶会论文、专利或技术博客,推动团队技术影响力建设。
更新于 2026-06-09北京|上海
