微软Senior Data & Applied Scientist
任职要求
Required Qualifications: • Doctorate in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, or related field AND 1+ year(s) data-science experience (e.g., managing structured and unstructured data, applying statistical techniques and reporting results)• OR Master's Degree in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, or related field AND 3+ years data-science experience (e.g., managing structured and unstructured data, applying statistical techn • OR Bachelor's Degree in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, or related field AND 5+ years data-science experience (e.g., managing structured and unstructured data, applying statistical tec • OR equivalent experience. • 2+ years customer-facing, project-delivery experience, professional services, and/or consulting experience • 4+ years applied ML/NLP experience delivering models and features to production at scale. • Software engineering excellence:• Proficiency in Python and PyTorch (or equivalent DL framework). • Solid SDLC practices: unit/integration testing, CI/CD, code reviews, version control, performance profiling, and reliability hardening. • Ability to write clean, maintainable, efficient code for production services and clients. • Experimentation & evaluation: sound experimental design, metric design (quality, safety, latency, cost), and statistical analysis; experience running online A/B tests. • Proven collaboration with PM & Engineering to integrate ML into shipped product (APIs/services/clients) and to drive measurable user or business impact. Preferred Qualifications: • Doctorate in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science,• OR related field AND 3+ years data-science experience (e.g., managing structured and unstructured data, applying statistical techniques and reporting results) • OR Master's Degree in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, • OR related field AND 5+ years data-science experience (e.g., managing structured and unstructured data, applying statistical techniques and reporting results) • OR Bachelor's Degree in Data Science, Mathematics, Statistics, Econometrics, Economics, Operations Research, Computer Science, • OR related field AND 7+ years data-science experience (e.g., managing structured and unstructured data, applying statistical techniques and reporting results) • OR e…
工作职责
• Ship features with PM & Engineering. Co‑own scenario goals; translate product requirements into scientific plans and productionized solutions that meet quality/latency/cost targets. • Model development & optimization. Design, fine‑tune, and evaluate models for LLM‑based authoring, summarization, reasoning, voice/chat, and personalization (e.g., SFT, alignment, prompt/tool use, safety filtering, multilingual & multimodal). • Data & evaluation at scale. Build/extend data pipelines for curation/labeling/feature stores; author offline eval harnesses; run online A/Bs and interleavings; define guardrails and success metrics; author scorecards and decision memos. • Production ML engineering. contribute to service code and configs; add monitoring, tracing, dashboards, and auto‑scaling; participate in on‑call and postmortems to improve live‑site reliability. • Responsible AI. Produce review artifacts, document mitigations for safety/privacy/fairness, support red‑teaming and sensitive‑use checks, and align with Microsoft’s Responsible AI Standard. • Collaboration & mentoring. Partner across PM/ENG/Design/CE/ORA/CELA; share methods and code, review PRs, improve reproducibility and documentation; mentor junior scientists.
AI Agent Engineering • Design, develop, and deploy production-grade AI agent systems, including multi-agent orchestration, tool-use frameworks, memory management, and API integration — ensuring reliability, scalability, and maintainability • Build and optimize Retrieval-Augmented Generation (RAG) pipelines: document ingestion, chunking strategy, embedding, vector search, and re-ranking to maximize LLM grounding quality • Support LLM adaptation to WWGS business domains through prompt engineering, context injection, fine-tuning signal curation, and systematic prompt evaluation frameworks • Develop automated knowledge base construction and real-time data access capabilities (Data Agent, MCP server/client) to connect AI agents with live business data • Design and implement LLM evaluation pipelines to systematically assess agent output quality, hallucination risk, and business impact Data Engineering • Design and implement end-to-end data pipelines (batch and streaming) for data collection, transformation, and storage — supporting both AI application and analytics use cases • Build and maintain integration layer data models that serve as a unified, AI-ready data foundation across WWGS domains • Develop automated data quality monitoring, alerting, and observability tooling to ensure pipeline reliability and data trustworthiness • Integrate multi-source data (seller behavior, transaction logs, off-platform signals, AI outputs) into a coherent, governed data layer • Establish data standardization and governance policies ensuring consistency, accuracy, and compliance across AI and BI consumption layers Technical Leadership • Provide technical guidance on AI-data architecture decisions; define best practices for the team's AI agent and data engineering stack • Collaborate cross-functionally with Product, Operations, and Science teams to translate business requirements into scalable technical solutions • Mentor junior engineers and conduct design reviews; raise the technical bar across the team
Partner with Product, Engineering, Design, and Business leads to define priorities, shape measurement frameworks, and translate insights into growth strategies for iCloud. Deliver data‑informed narratives for senior leadership that directly inform product and business roadmaps. Drive organizational alignment by clarifying analytical outcomes and implications for diverse stakeholders. Design performance metrics that capture product health, user engagement, and business impact across iCloud services. Lead controlled experiments, apply causal frameworks and marketing mix models to evaluate initiatives at scale. Conduct deep behavioral analysis to uncover engagement, retention, and churn drivers, then translate findings into actionable strategies. Collaborate with Data Engineering to define instrumentation and build reliable data pipelines across ASE. Develop and productionize ML and statistical models that enable predictive and diagnostic insights for product decisions. Leverage generative AI to automate analytical workflows and expand measurement capabilities at scale.
• Collaborate with BIE,DE, PM, CSM to research, design, develop, and evaluate generative AI solutions to address Global Selling challenges. • Interact with stakeholders directly to understand their business problems, aid them in implementation of generative AI solutions, brief stkaholders and guide them on adoption patterns and paths to production • Create and deliver best practice recommendations, tutorials, blog posts, sample code, and presentations adapted to technical, business, and executive stakeholder
• Design and implement end-to-end data pipelines (ETL) to ensure efficient data collection, cleansing, transformation, and storage, supporting both real-time and offline analytics needs. • Develop automated data monitoring tools and interactive dashboards to enhance business teams’ insights into core metrics (e.g., user behavior, AI model performance). • Collaborate with cross-functional teams (e.g., Product, Operations, Tech) to align data logic, integrate multi-source data (e.g., user behavior, transaction logs, AI outputs), and build a unified data layer. • Establish data standardization and governance policies to ensure consistency, accuracy, and compliance. • Provide structured data inputs for AI model training and inference (e.g., LLM applications, recommendation systems), optimizing feature engineering workflows. • Explore innovative AI-data integration use cases (e.g., embedding AI-generated insights into BI tools). • Provide technical guidance and best practice on data architecture that meets both traditional reporting purpose and modern AI Agent requirements.