微软Principal Software Engineer--M365 Storage Team
任职要求
Required Qualifications: • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python or equivalent experience. • Extensive industry experience building large-scale cloud services, distributed systems, storage platforms, or database systems. • Solid understanding of operating systems, memory management, threading, synchronization, networking, and I/O subsystems. • Expert knowledge of distributed system design and architecture. • Proven experience designing or operating large-scale storage and database platforms. • Solid understanding of storage engines, database internals, replication mechanisms, transaction processing, consistency models, fault tolerance, high availability architectures, and data durability strategies. • Experience with large-scale data platforms supporting mission-critical workloads. • Demonstrated success building systems that operate at hyperscale. • Experience designing and supporting high-concurrency architectures, low-latency services, and high-throughput systems, with expertise in service resiliency, performance tuning, capacity management, and production operations. • Ability to diagnose and resolve problems across complex distributed environments. • Solid interest and demonstrated experience applying AI to engineering workflows and product architectures. • Ability to critically evaluate emerging AI technologies and identify transformational opportunities. • Experience using AI-assisted development, design, diagnostics, automation, or operational workflows. • Passion for reimagining platforms and engineering systems through AI-driven innovation. • Exceptional analytical and problem-solving skills. • Ability to navigate ambiguity and drive clarity in highly complex technical domains. • Solid architecture review and design evaluation skills. • Proven influence across organizations without direct authority. • Excellent communication and collaboration skills with engineers, architects, product leaders, and executives. Other Requirements: Ability to meet Microsoft, customer and/or government security screening requirements are required for thi…
工作职责
OverviewAI-Native Systems, Storage & Data Platform Architecture We are seeking a highly technical Principal Software Engineer to drive the next generation of AI-native architecture across critical storage, database, and distributed system components. This role will lead the modernization and transformation of foundational platform technologies that power large-scale cloud services and AI workloads. The ideal candidate possesses deep expertise in systems programming, storage architecture, distributed databases, and high-performance service design. They combine strong critical thinking with an AI-native mindset, leveraging AI not only as a productivity tool, but as a catalyst to fundamentally rethink system architecture, engineering workflows, performance optimization, reliability, and operational excellence. This role requires both hands-on technical depth and strategic influence, shaping technologies that operate at hyperscale while delivering industry-leading reliability, scalability, efficiency, and latency characteristics. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. ResponsibilitiesSystem Architecture & Platform Leadership• Lead the architecture, design, and evolution of large-scale distributed storage, database, and data platform systems. • Drive modernization of critical platform components to support next-generation AI and Copilot workloads. • Re-architect legacy services and infrastructure using AI-native design principles. • Define long-term technical strategy and architectural direction across multiple services and teams. • Identify and eliminate architectural bottlenecks impacting scalability, reliability, performance, and operational efficiency. Systems Programming & Performance Engineering• Design and implement highly efficient systems-level software in languages such as C++, Rust, C#, Go, or similar. • Drive end-to-end performance optimization across storage engines, networking, caching, concurrency control, and data access layers. • Solve complex challenges involving high concurrency, low latency, throughput optimization, and resource efficiency. • Lead root-cause analysis and resolution of difficult production issues involving distributed systems and storage infrastructure. • Establish engineering standards and best practices for performance-critical software development. Storage & Database Innovation• Design and optimize storage engines, database architectures, replication technologies, indexing systems, and data management frameworks. • Lead innovation in areas such as: • High availability and disaster recovery • Data durability and consistency • Replication and synchronization • Metadata management • Query optimization • Storage efficiency and cost optimization • Intelligent caching and tiering • Drive architectural improvements that enable large-scale AI and retrieval-driven workloads. AI-Native Transformation• Champion AI-native engineering practices across architecture, development, testing, and operations. • Apply AI-assisted approaches to system design, performance analysis, reliability engineering, and operational automation. • Identify opportunities to redesign products and platforms around emerging AI capabilities instead of incrementally enhancing legacy models. • Build frameworks and workflows that integrate AI into engineering decision-making and operational management. • Influence the organization's AI transformation strategy through thought leadership and technical innovation.
OverviewWe are looking for a Principal Software Engineer to help lead the strategic evolution of the Experimentation and Configuration Service (ECS), Microsoft’s platform for safe, scalable, intelligent, and resilient configuration management. As a Principal Software Engineer in ECS, you will be accountable for broad technical direction across a major platform area or cross-team initiative. You will lead complex architecture from ideation to production, create clarity across teams, drive multi-release technical strategy, and influence engineering practices across ECS and partner organizations. This role is for a technical leader who can operate in ambiguous problem spaces, broker difficult cross-team architecture decisions, and change the momentum of current plans toward higher-impact outcomes. You will help ECS evolve into an agentic-native Continuous Configuration platform for Microsoft services, combining distributed systems expertise, operational excellence, AI-native development, and customer-focused platform strategy. About the ECS Team The ECS team builds the Experimentation and Configuration platform used by Microsoft services to manage safe configuration delivery, experimentation, feature rollout, policy enforcement, telemetry, reliability, authenticated access, and operational governance at scale. ECS is transforming from a traditional configuration platform into an intelligent, agentic-native infrastructure layer. The team is investing in AI-assisted engineering, MCP interfaces, agent-powered troubleshooting, autonomous validation, self-healing workflows, configuration intelligence, and compliance automation to help Microsoft services move faster while reducing operational risk. Why Join Us This is a great opportunity to shape a foundational Microsoft platform at the moment when configuration management, AI-native engineering, service reliability, and agentic automation are converging. As a Principal Software Engineer in ECS, you will help define the next generation of Continuous Configuration for Microsoft, influencing how services safely ship, validate, diagnose, and recover at global scale. You will work on deeply technical, high-impact problems with broad customer reach, influence cross-team strategy, and help build an engineering system where agents and humans collaborate to deliver safer, faster, and more intelligent platform outcomes.Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. Responsibilities• Strategic Technical Leadership: Define and drive the architecture, roadmap, and execution strategy for a major ECS platform area or cross-team initiative with multi-release impact. • Cross-Team Architecture Ownership: Broker complex architecture decisions across ECS, partner teams, PM, security, compliance, and service owners; create shared direction and durable technical alignment. • Agentic-Native Platform Transformation: Lead the design and adoption of agentic-native engineering systems, including AI coding agents, MCP interfaces, SKILLs, intelligent automation, and AI-native engineering practices such as harness engineering, loop engineering, and evaluation systems. • Technical Vision & Long-Term Direction: Translate ambiguous customer, business, security, reliability, and operational needs into clear technical bets, phased execution plans, and measurable outcomes. • Platform Quality & Engineering Excellence: Define quality metrics, coding patterns, validation strategy, telemetry, observability, and operational readiness standards that raise the engineering bar across ECS. • Complex Problem Resolution: Resolve the hardest technical and product-system problems in ECS, including issues spanning distributed systems, scale, reliability, security, compliance, and cross-service dependencies. • Customer & Partner Influence: Build trusted technical relationships with major partner teams and use customer insights, production signals, and platform adoption data to improve ECS strategy and design. • Organizational Impact: Change the momentum of existing plans where needed, identify higher-impact directions, and help teams pivot toward better customer and business outcomes. • Mentorship & Capability Building: Mentor senior engineers and tech leads, build durable technical capability across the team, and create an inclusive engineering environment where people can do their best work. • Operational Accountability: Ensure ECS platform investments are designed for production excellence, measurable customer impact, cost awareness, reliability, security, and long-term maintainability.
• Own the design and implementation of core infrastructure components, ensuring scalability, security, and high performance across systems. • Lead architectural decisions, set engineering standards, and drive long-term technical strategy for complex solutions. • Produce and review high-quality, secure, and maintainable code while championing best practices and modern patterns, including AI-driven approaches. • Mentor engineers across teams, fostering growth in coding, design, testing, and operational excellence. • Define and enforce robust test strategies, integrate automation, and ensure security and reliability in all deliverables. • Oversee telemetry, incident response, and operational readiness to improve system stability and supportability. • Collaborate with stakeholders to validate requirements, incorporate feedback, and uphold compliance, privacy, and accessibility standards.
- Keep up to date with and utilize the latest developments in LLM system optimization.- Take the lead in designing innovative system optimization solutions for internal LLM workloads.- Optimize LLM inference workloads through innovative kernel, algorithm, scheduling, and parallelization technologies.- Continuously develop and maintain internal LLM inference infrastructure.- Discover new LLM system optimization needs and innovations.
•  Work with a team of passionate engineers to deliver success for customers•  Design, implement, test, and operate data services.•  Release features on time, with high quality, meeting functional, performance, scalability, and compliance requirements.•  Drive quality right from the design phase, incorporating best practices and engineering for testability.•  Solve problems relating to mission critical services and create solutions to prevent problem recurrence.•  Participate in product live site and operations.