阿里云阿里云智能-PolarDB for AI研发工程师-杭州
社招全职3年以上云智能集团地点:杭州状态:招聘
任职要求
1)硕士及以上,计算机、软件工程、人工智能、数学等相关专业; 2)3-5年以上AI系统研发或高性能计算相关经验; 3)深入理解CPU/GPU异构计算架构,具备资源调度与优化经验;精通至少一种主流推理框架(TensorRT、ONNX Runtime、vLLM、TGI等); 4)熟练掌握模型量化(INT8/FP16/BF16)、剪枝、蒸馏等压缩技术;了解小模型(BERT、ResNet…
登录查看完整任职要求
微信扫码,1秒登录
工作职责
1)构建数据库内AI推理系统架构,高效整合CPU,GPU等资源,设计数据迁移管理机制,优化模型(包括大模型和小模型)的核心性能指标;对PostgreSQL/MySQL等任一数据库有深入理; 2)研发具备负载自适应能力的推理框架,开发高精度指标采集模块,实现基于实时负载特征的动态参数调优功能; 3)研究在精度可接受范围内,多种近似推理及轻量化技术,包括采用模型压缩(如量化、剪枝)或近似算法(如近似最近邻搜索)降低计算开销; 4)收集、识别、分析客户需求,并确定技术方案的目标、范围和交付成果; 5)基于需求分析,进行技术可行性分析和方案评审,选择合适的技术选型、功能设计、技术架构、数据架构和开发流程等; 6) 对开发中和部署后的程序进行必要的维护和迭代,包括值班oncall、升级工单处置、bug排查、问题诊断、产品体验改善、性能和成本优化等。
包括英文材料
TensorRT+
https://docs.nvidia.com/deeplearning/tensorrt/latest/getting-started/quick-start-guide.html
This TensorRT Quick Start Guide is a starting point for developers who want to try out the TensorRT SDK; specifically, it demonstrates how to quickly construct an application to run inference on a TensorRT engine.
ONNX+
https://github.com/onnx/tutorials
Open Neural Network Exchange (ONNX) is an open standard format for representing machine learning models.
[英文] Introduction to ONNX
https://onnx.ai/onnx/intro/
This documentation describes the ONNX concepts (Open Neural Network Exchange).
vLLM+
https://www.newline.co/@zaoyang/ultimate-guide-to-vllm--aad8b65d
vLLM is a framework designed to make large language models faster, more efficient, and better suited for production environments.
https://www.youtube.com/watch?v=Ju2FrqIrdx0
vLLM is a cutting-edge serving engine designed for large language models (LLMs), offering unparalleled performance and efficiency for AI-driven applications.
TGI+
https://huggingface.co/docs/text-generation-inference/en/index
Text Generation Inference (TGI) is a toolkit for deploying and serving Large Language Models (LLMs).
https://learn.ritual.net/examples/tgi_inference_with_mistral_7b
In this tutorial, we will use Huggingface's TGI (Text Generation Interface) API to query a Large Language Model (LLM) and enable users to requests jobs from it, both on-chain and off-chain.
https://www.sandgarden.com/learn/text-generation-inference-tgi
Text Generation Inference (TGI) is the process by which a trained AI model generates new text based on an input prompt, focusing on producing this text efficiently in terms of speed and computational resources.
还有更多 •••
相关职位