
商汤Senior Data Engineer@
任职要求
• AI First Developer (>80% coding with AI Tools),understand how to control and maximize the AI programming tools capabilities, super-charged developer • Experience: 5+ years in data engineering or software engineering, with 3+ years architecting cloud scale data platforms. • Flexible and have the passion of creation impactful product • Fluent English communication skills (spoken and written) are requiredprefered • Technical Skills: Expert in SQL and one of Python/Scala/Java; deep hands on knowledge of (Spark), Kafka, and orchestration (Airflow, dbt, Prefect). • Ability to understand business scenarios and design data acquisition solutions that directly impact AI model perfo…
工作职责
• Understand business scenarios and design targeted data acquisition solutions, ensuring data is relevant, high-quality, and aligned with project goals. • Architect, design, and maintain enterprise-grade databases, data warehouses, and lakehouse systems to support analytical, operational, and AI workloads. • Model and optimize schema design, storage layouts, data partitioning, clustering, and indexing strategies for large-scale datasets. • Implement and maintain ETL/ELT pipelines feeding data warehouses (e.g., Snowflake, BigQuery, Redshift, Databricks, or open-lakehouse environments). • Design, collect, and maintain high-quality datasets for AI inferencing and LLM model optimization, fine-tuning, and testing, ensuring data is formatted and preprocessed to meet model requirements. • Collaborate with AI application engineers to understand model performance requirements and translate them into targeted data collection and preparation strategies. • Develop and implement automated data pipelines for efficient data processing, including data cleaning, labeling, augmentation, and transformation. • Proactively identify data gaps based on model performance metrics, design solutions to acquire, clean, and optimize data for enhanced model accuracy and efficiency. • Build, clean, and manage diverse data sources, ensuring compliance with data security and privacy standards. • Conduct exploratory data analysis to discover data patterns, anomalies, and optimization opportunities, directly impacting model performance. • Continuously learn and adapt to the latest advancements in data engineering, AI, and large language model (LLM) technologies.
AI Agent Engineering • Design, develop, and deploy production-grade AI agent systems, including multi-agent orchestration, tool-use frameworks, memory management, and API integration — ensuring reliability, scalability, and maintainability • Build and optimize Retrieval-Augmented Generation (RAG) pipelines: document ingestion, chunking strategy, embedding, vector search, and re-ranking to maximize LLM grounding quality • Support LLM adaptation to WWGS business domains through prompt engineering, context injection, fine-tuning signal curation, and systematic prompt evaluation frameworks • Develop automated knowledge base construction and real-time data access capabilities (Data Agent, MCP server/client) to connect AI agents with live business data • Design and implement LLM evaluation pipelines to systematically assess agent output quality, hallucination risk, and business impact Data Engineering • Design and implement end-to-end data pipelines (batch and streaming) for data collection, transformation, and storage — supporting both AI application and analytics use cases • Build and maintain integration layer data models that serve as a unified, AI-ready data foundation across WWGS domains • Develop automated data quality monitoring, alerting, and observability tooling to ensure pipeline reliability and data trustworthiness • Integrate multi-source data (seller behavior, transaction logs, off-platform signals, AI outputs) into a coherent, governed data layer • Establish data standardization and governance policies ensuring consistency, accuracy, and compliance across AI and BI consumption layers Technical Leadership • Provide technical guidance on AI-data architecture decisions; define best practices for the team's AI agent and data engineering stack • Collaborate cross-functionally with Product, Operations, and Science teams to translate business requirements into scalable technical solutions • Mentor junior engineers and conduct design reviews; raise the technical bar across the team
• Design and implement end-to-end data pipelines (ETL) to ensure efficient data collection, cleansing, transformation, and storage, supporting both real-time and offline analytics needs. • Develop automated data monitoring tools and interactive dashboards to enhance business teams’ insights into core metrics (e.g., user behavior, AI model performance). • Collaborate with cross-functional teams (e.g., Product, Operations, Tech) to align data logic, integrate multi-source data (e.g., user behavior, transaction logs, AI outputs), and build a unified data layer. • Establish data standardization and governance policies to ensure consistency, accuracy, and compliance. • Provide structured data inputs for AI model training and inference (e.g., LLM applications, recommendation systems), optimizing feature engineering workflows. • Explore innovative AI-data integration use cases (e.g., embedding AI-generated insights into BI tools). • Provide technical guidance and best practice on data architecture that meets both traditional reporting purpose and modern AI Agent requirements.
AI Agent Engineering • Design, develop, and deploy production-grade AI agent systems, including multi-agent orchestration, tool-use frameworks, memory management, and API integration — ensuring reliability, scalability, and maintainability • Build and optimize Retrieval-Augmented Generation (RAG) pipelines: document ingestion, chunking strategy, embedding, vector search, and re-ranking to maximize LLM grounding quality • Support LLM adaptation to WWGS business domains through prompt engineering, context injection, fine-tuning signal curation, and systematic prompt evaluation frameworks • Develop automated knowledge base construction and real-time data access capabilities (Data Agent, MCP server/client) to connect AI agents with live business data • Design and implement LLM evaluation pipelines to systematically assess agent output quality, hallucination risk, and business impact Data Engineering • Design and implement end-to-end data pipelines (batch and streaming) for data collection, transformation, and storage — supporting both AI application and analytics use cases • Build and maintain integration layer data models that serve as a unified, AI-ready data foundation across WWGS domains • Develop automated data quality monitoring, alerting, and observability tooling to ensure pipeline reliability and data trustworthiness • Integrate multi-source data (seller behavior, transaction logs, off-platform signals, AI outputs) into a coherent, governed data layer • Establish data standardization and governance policies ensuring consistency, accuracy, and compliance across AI and BI consumption layers Technical Leadership • Provide technical guidance on AI-data architecture decisions; define best practices for the team's AI agent and data engineering stack • Collaborate cross-functionally with Product, Operations, and Science teams to translate business requirements into scalable technical solutions • Mentor junior engineers and conduct design reviews; raise the technical bar across the team
• Providing Ethernet and routing expertise to customers during project delivery to design, architect and test Ethernet networking solutions. • Work on multi-functional teams to provide Ethernet network expertise to server infrastructure builds, accelerated computing workloads and GPU enabled AI applications. • Crafting and evaluating DevOps automation scripts for network operations, crafting network architectures, and developing switch fabric configurations. • Implementing tasks related to network configuration and validation for data centers. • Create Methods of Procedure and deployment documents. • Use software tools to validate and monitor network performance.