腾讯企业微信-多模态大模型算法工程师 -音频方向
社招全职3年以上企业微信SaaS技术地点:广州 | 成都状态:招聘
任职要求
1.计算机科学、人工智能、计算机视觉或相关专业硕士及以上学历; 2.扎实的编程基础,熟练掌握Python和PyTorch/TensorFlow; 3.在计算机视觉(CV)和自然语言处理(NLP)其中一个领域有深厚积累,并有多模态学习项目经验; 4.熟悉主流的多模态模型架构(如Transformer-based VL models),有相关模型的训练、微调或部署经验; 5.对技术创新有强烈兴趣,具备优秀的工程实现能力,能将算法模型应用于大规模实际场景。 加分项 1.计算机、信号处理、电子工程等相关专业,硕士及以上学历,3年以…
登录查看完整任职要求
微信扫码,1秒登录
工作职责
1.负责企业微信音频 AI 相关算法的研究与落地,包括但不限于语音识别(ASR)、语音合成(TTS)、声纹识别、音色转换等方向; 2.负责热词定制、领域自适应、说话人分离等场景化能力的算法设计与优化; 3.探索音频大模型在企业办公场景的创新应用,推动模型训练、微调及端侧部署落地; 4.跟进语音/音频领域前沿技术进展(Whisper、SpeechGPT 等),持续提升核心指标与用户体验; 5.与客户端、后台团队协作,完成算法从原型验证到工程化落地的全链路交付。
包括英文材料
OpenCV+
https://learnopencv.com/getting-started-with-opencv/
At LearnOpenCV we are on a mission to educate the global workforce in computer vision and AI.
https://opencv.org/university/free-opencv-course/
This free OpenCV course will teach you how to manipulate images and videos, and detect objects and faces, among other exciting topics in just about 3 hours.
学历+
Python+
https://liaoxuefeng.com/books/python/introduction/index.html
中文,免费,零起点,完整示例,基于最新的Python 3版本。
https://www.learnpython.org/
a free interactive Python tutorial for people who want to learn Python, fast.
https://www.youtube.com/watch?v=K5KVEU3aaeQ
Master Python from scratch 🚀 No fluff—just clear, practical coding skills to kickstart your journey!
https://www.youtube.com/watch?v=rfscVS0vtbw
This course will give you a full introduction into all of the core concepts in python.
PyTorch+
https://datawhalechina.github.io/thorough-pytorch/
PyTorch是利用深度学习进行数据科学研究的重要工具,在灵活性、可读性和性能上都具备相当的优势,近年来已成为学术界实现深度学习算法最常用的框架。
https://www.youtube.com/watch?v=V_xro1bcAuA
Learn PyTorch for deep learning in this comprehensive course for beginners. PyTorch is a machine learning framework written in Python.
TensorFlow+
https://www.youtube.com/watch?v=tpCFfeUEGs8
Ready to learn the fundamentals of TensorFlow and deep learning with Python? Well, you’ve come to the right place.
https://www.youtube.com/watch?v=ZUKz4125WNI
This part continues right where part one left off so get that Google Colab window open and get ready to write plenty more TensorFlow code.
NLP+
https://www.youtube.com/watch?v=fNxaJsNG3-s&list=PLQY2H8rRoyvzDbLUZkbudP-MFQZwNmU4S
Welcome to Zero to Hero for Natural Language Processing using TensorFlow!
https://www.youtube.com/watch?v=R-AG4-qZs1A&list=PLeo1K3hjS3uuvuAXhYjV2lMEShq2UYSwX
Natural Language Processing tutorial for beginners series in Python.
https://www.youtube.com/watch?v=rmVRLeJRkl4&list=PLoROMvodv4rMFqRtEuo6SGjY4XbRIVRd4
The foundations of the effective modern methods for deep learning applied to NLP.
Transformer+
https://huggingface.co/learn/llm-course/en/chapter1/4
Breaking down how Large Language Models work, visualizing how data flows through.
https://poloclub.github.io/transformer-explainer/
An interactive visualization tool showing you how transformer models work in large language models (LLM) like GPT.
https://www.youtube.com/watch?v=wjZofJX0v4M
Breaking down how Large Language Models work, visualizing how data flows through.
算法+
https://roadmap.sh/datastructures-and-algorithms
Step by step guide to learn Data Structures and Algorithms in 2025
https://www.hellointerview.com/learn/code
A visual guide to the most important patterns and approaches for the coding interview.
https://www.w3schools.com/dsa/
大模型+
https://www.youtube.com/watch?v=xZDB1naRUlk
You will build projects with LLMs that will enable you to create dynamic interfaces, interact with vast amounts of text data, and even empower LLMs with the capability to browse the internet for research papers.
https://www.youtube.com/watch?v=zjkBMFhNj_g
GPT+
https://www.youtube.com/watch?v=kCc8FmEb1nY
We build a Generatively Pretrained Transformer (GPT), following the paper "Attention is All You Need" and OpenAI's GPT-2 / GPT-3.
CVPR+
https://cvpr.thecvf.com/
还有更多 •••
相关职位
社招3年以上企业微信SaaS
1.负责多模态大模型(如音视频理解、视觉问答、图像生成等)的技术研究、应用落地与性能优化; 2.研发和优化基于大模型的多模态应用; 3.收集和构建高质量的多模态数据集,并进行模型的训练、微调和提示工程(Prompt Engineering); 4.将多模型算法高效地集成到企业微信客户端,与客户端团队合作解决端侧部署和推理的挑战; 5.紧跟多模态领域(如CLIP, BLIP, Stable Diffusion, Sora等)的技术前沿,推动技术创新在产品中落地。
更新于 2026-01-09成都
社招3年以上企业微信SaaS
1.负责企业微信音频 AI 相关算法的研究与落地,包括但不限于语音识别(ASR)、语音合成(TTS)、声纹识别、音色转换等方向; 2.负责热词定制、领域自适应、说话人分离等场景化能力的算法设计与优化; 3.探索音频大模型在企业办公场景的创新应用,推动模型训练、微调及端侧部署落地; 4.跟进语音/音频领域前沿技术进展(Whisper、SpeechGPT 等),持续提升核心指标与用户体验; 5.与客户端、后台团队协作,完成算法从原型验证到工程化落地的全链路交付。
更新于 2026-03-31成都
社招1年以上企业微信SaaS
1.负责AI agent的研发与维护工作,完成各AI agent的功能实现、性能维护; 2.与产品、算法、测试等团队紧密配合,搭建AI agent服务,打通企业自有数据和服务,丰富企业微信AI能力; 3.持续进行系统的技术架构优化与性能优化,提升产品竞争力。
更新于 2026-07-28广州