
同花顺多模态算法工程师
校招全职AI 算法类地点:杭州状态:招聘
工作描述
任职要求 岗位要求: 1、有多模态模型研发经验:VL、AL、AV、Video、Omni任一方向 2、熟练使用多模态开源模型,如 Qwen-omni、LLaVA、Whisper、Clap、MERT、SeamlessM4T 等 3、有大规模模型训练经验:SFT、DPO、RLHF、GRPO、MoE、长上下文训练 4、掌握音频/视频建模,例如ASR、TTS、音频编码、视频理解/生成 5、有模型推理优化经验:TensorRT、vLLM、FlashAttention、KV Cache、量化、稀疏化 6、有Agent系统、RAG、多模态增强检索、工具调用链构建经验 7、在…
登录查看完整工作描述
微信扫码,1秒登录
包括英文材料
SFT+
https://cameronrwolfe.substack.com/p/understanding-and-using-supervised
Understanding how SFT works from the idea to a working implementation...
RLHF+
[英文] What is RLHF?
https://aws.amazon.com/what-is/reinforcement-learning-from-human-feedback/
Reinforcement learning from human feedback (RLHF) is a machine learning (ML) technique that uses human feedback to optimize ML models to self-learn more efficiently.
https://www.ibm.com/think/topics/rlhf
Reinforcement learning from human feedback (RLHF) is a machine learning technique in which a “reward model” is trained with direct human feedback, then used to optimize the performance of an artificial intelligence agent through reinforcement learning.
语音识别+
https://developer.nvidia.com/blog/essential-guide-to-automatic-speech-recognition-technology/
Over the past decade, AI-powered speech recognition systems have slowly become part of our everyday lives, from voice search to virtual assistants in contact centers, cars, hospitals, and restaurants.
语音合成+
https://www.ibm.com/think/topics/text-to-speech
Text to speech (TTS) is a type of technology that converts text on a digital interface into natural-sounding audio.
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
缓存+
https://hackernoon.com/the-system-design-cheat-sheet-cache
The cache is a layer that stores a subset of data, typically the most frequently accessed or essential information, in a location quicker to access than its primary storage location.
https://www.youtube.com/watch?v=bP4BeUjNkXc
Caching strategies, Distributed Caching, Eviction Policies, Write-Through Cache and Least Recently Used (LRU) cache are all important terms when it comes to designing an efficient system with a caching layer.
https://www.youtube.com/watch?v=dGAgxozNWFE
还有更多 •••
相关职位
校招研发类
1、计算机科学、电子工程、人工智能、统计学等相关领域专业; 2、具备良好的数学基础,熟悉至少一种深度学习框架,并有一定编程实践经验; 3、对多模态领域有一定的了解,包括文本、图像、语音等多模态数据处理
更新于 2026-08-18深圳
校招AI/算法类
1. 拥有计算机科学、人工智能、机器学习或相关领域的硕士或博士学位; 2. 扎实的编程基础,熟练掌握Python、C++等编程语言,有TensorFlow、PyTorch等深度学习框架的使用经验; 3
更新于 2026-03-06北京|深圳

校招ai 算法类
- 计算机、数学、统计学等相关专业硕士及以上学历。 - 熟悉深度学习框架(如 TensorFlow、PyTorch),对多模态模型(如 CLIP、OFA 等)有深入了解。 - 掌握多模态数据处理、特征
杭州