滴滴Senior Data Scientist(J251202006)
任职要求
1. 计算机、数学、统计学等相关专业本科及以上,3 年以上数据仓库或大数据平台经验。 2. 精通 Hive/Flink/Spark/Kafka,具备 TB 级以上数据建模与调优能力;熟悉 StarRocks/Doris/ClickHouse 等 OLAP 引擎优先。 3. 熟悉维度建模理论,能独立完…
工作职责
【数据研发/数据分析方向】 1. 负责国际化出行(拉美、亚太、中东等)用户增长、用户运营、广告投放,支持全球多国家/多时区/多币种数据模型。 2. 熟悉数仓架构(ODS→CDM→ADS),确保数据质量、一致性及可扩展性,丰富业务应用层集市,构建业务数据门户,提升数据分析、业务决策的效率。 3. 与增长产品经理、策略运营深度协同,将埋点、用户行为、营销花费、补贴效果等数据沉淀为可复用的增长数据资产。 4. 构建实时+离线数据体系,支持用户画像的实时获取,以及实时计算用户平台价值。 5. 制定并落地数据治理规范(口径、生命周期、成本、权限、合规),满足 GDPR/CCPA/当地数据出境要求。 6. 建立数据异常检测与自动修复机制,保障核心指标(ROI、CAC、留存、LTV)的准确性与及时性。
• Collaborate with BIE,DE, PM, CSM to research, design, develop, and evaluate generative AI solutions to address Global Selling challenges. • Interact with stakeholders directly to understand their business problems, aid them in implementation of generative AI solutions, brief stkaholders and guide them on adoption patterns and paths to production • Create and deliver best practice recommendations, tutorials, blog posts, sample code, and presentations adapted to technical, business, and executive stakeholder
• Ship features with PM & Engineering. Co‑own scenario goals; translate product requirements into scientific plans and productionized solutions that meet quality/latency/cost targets. • Model development & optimization. Design, fine‑tune, and evaluate models for LLM‑based authoring, summarization, reasoning, voice/chat, and personalization (e.g., SFT, alignment, prompt/tool use, safety filtering, multilingual & multimodal). • Data & evaluation at scale. Build/extend data pipelines for curation/labeling/feature stores; author offline eval harnesses; run online A/Bs and interleavings; define guardrails and success metrics; author scorecards and decision memos. • Production ML engineering. contribute to service code and configs; add monitoring, tracing, dashboards, and auto‑scaling; participate in on‑call and postmortems to improve live‑site reliability. • Responsible AI. Produce review artifacts, document mitigations for safety/privacy/fairness, support red‑teaming and sensitive‑use checks, and align with Microsoft’s Responsible AI Standard. • Collaboration & mentoring. Partner across PM/ENG/Design/CE/ORA/CELA; share methods and code, review PRs, improve reproducibility and documentation; mentor junior scientists.
AI Agent Engineering • Design, develop, and deploy production-grade AI agent systems, including multi-agent orchestration, tool-use frameworks, memory management, and API integration — ensuring reliability, scalability, and maintainability • Build and optimize Retrieval-Augmented Generation (RAG) pipelines: document ingestion, chunking strategy, embedding, vector search, and re-ranking to maximize LLM grounding quality • Support LLM adaptation to WWGS business domains through prompt engineering, context injection, fine-tuning signal curation, and systematic prompt evaluation frameworks • Develop automated knowledge base construction and real-time data access capabilities (Data Agent, MCP server/client) to connect AI agents with live business data • Design and implement LLM evaluation pipelines to systematically assess agent output quality, hallucination risk, and business impact Data Engineering • Design and implement end-to-end data pipelines (batch and streaming) for data collection, transformation, and storage — supporting both AI application and analytics use cases • Build and maintain integration layer data models that serve as a unified, AI-ready data foundation across WWGS domains • Develop automated data quality monitoring, alerting, and observability tooling to ensure pipeline reliability and data trustworthiness • Integrate multi-source data (seller behavior, transaction logs, off-platform signals, AI outputs) into a coherent, governed data layer • Establish data standardization and governance policies ensuring consistency, accuracy, and compliance across AI and BI consumption layers Technical Leadership • Provide technical guidance on AI-data architecture decisions; define best practices for the team's AI agent and data engineering stack • Collaborate cross-functionally with Product, Operations, and Science teams to translate business requirements into scalable technical solutions • Mentor junior engineers and conduct design reviews; raise the technical bar across the team
• Design and implement end-to-end data pipelines (ETL) to ensure efficient data collection, cleansing, transformation, and storage, supporting both real-time and offline analytics needs. • Develop automated data monitoring tools and interactive dashboards to enhance business teams’ insights into core metrics (e.g., user behavior, AI model performance). • Collaborate with cross-functional teams (e.g., Product, Operations, Tech) to align data logic, integrate multi-source data (e.g., user behavior, transaction logs, AI outputs), and build a unified data layer. • Establish data standardization and governance policies to ensure consistency, accuracy, and compliance. • Provide structured data inputs for AI model training and inference (e.g., LLM applications, recommendation systems), optimizing feature engineering workflows. • Explore innovative AI-data integration use cases (e.g., embedding AI-generated insights into BI tools). • Provide technical guidance and best practice on data architecture that meets both traditional reporting purpose and modern AI Agent requirements.