Microsoft Logo

Microsoft

Principal Software Engineering Manager

Reposted 3 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in United States
143K-304K Annually
Senior level
Remote
Hiring Remotely in United States
143K-304K Annually
Senior level
Lead and grow a software engineering team to design, build, and operate high-performance, scalable, observable networking systems for Azure AI/HPC infrastructure. Drive architecture, reliability, testing, automation, and cross-team collaboration to support large-scale distributed training and inference workloads.
The summary above was generated by AI
Overview

The HPC/AI (High-Performance Computing and Artificial Intelligence) organization is on a mission to build the next generation of distributed AI supercomputers - systems that deliver unprecedented computational power, scalability, and reliability to accelerate breakthroughs in artificial intelligence. Our teams design and develop world-class AI infrastructure that enables large-scale model training and inference, forming the backbone of Microsoft’s AI innovation.

As a Principal Software Engineering Manager, you will lead a team building foundational components of Azure’s AI networking infrastructure—powering some of the largest and most complex distributed training systems in the world. This is a rare opportunity to work at the intersection of AI, cloud infrastructure, and high-performance networking, driving innovation across hardware and software boundaries. With the explosive growth of generative AI and the demand for low-latency, high-bandwidth systems, your work will directly impact the scale, performance, and reliability of Microsoft’s AI platforms.
You will lead the design, development, and deployment of high-performance, scalable, and observable networking systems that connect AI accelerators at massive scale. The role requires deep technical acumen, strategic thinking, and a passion for engineering excellence. You’ll collaborate across Microsoft teams to define architecture, deliver solutions to complex infrastructure challenges, and ensure our systems meet the evolving needs of AI workloads.
If you’re passionate about building large-scale distributed systems, pushing the boundaries of AI infrastructure, and leading teams that shape the future of supercomputing, we invite you to join us on this journey to define the next era of AI at Microsoft.
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. 


Responsibilities
  • Hire, manage, and grow a high-performing team of software engineers, fostering a culture of excellence, inclusion, and innovation.
  • Lead the design and development of large-scale distributed systems and services that power Azure’s AI infrastructure.
  • Drive engineering planning and execution while ensuring alignment with organizational OKRs and long-term strategy.
  • Establish lean, scalable, and efficient processes that promote innovation and engineering rigor.
  • Deliver best-in-class engineering by ensuring services and components are modular, secure, reliable, diagnosable, observable, and reusable.
  • Improve test coverage, automation, and integration testing to proactively identify and resolve reliability gaps.
  • Ensure live-site reliability and service health through robust monitoring, telemetry, and automation.
  • Collaborate across Microsoft and partner organizations to deliver cohesive, end-to-end infrastructure solutions.
  • Apply data-driven insights to optimize performance, scalability, and customer satisfaction.

Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience. 

Other Requirements: 

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:  
    • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter. 
Preferred Qualifications:
  • Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
    • OR equivalent experience.
  • 4+ years people management experience.
  • 1+ years experience building and operating networking infrastructure for hyperscale datacenters or AI clusters.
  • 1+ years hands-on experience with networking technologies in AI-specific hardware (e.g., InfiniBand, ROCE, MRC, NVLink, UALink).
  • 10+ years of professional software design and development experience in large-scale distributed systems.
#azurecorejobs

Software Engineering M5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar Jobs

3 Days Ago
Remote
United States
143K-304K Annually
Expert/Leader
143K-304K Annually
Expert/Leader
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead and grow an engineering team that designs, builds, and operates hyperscale distributed systems and cloud-native platform services for Azure. Drive technical strategy, architecture, cross-team initiatives, operational excellence, and stakeholder collaboration to deliver scalable, reliable, secure, and observable infrastructure solutions.
Top Skills: AzureC#C++Cloud-Native ServicesDistributed SystemsGoJavaPython
4 Days Ago
Remote
United States
143K-304K Annually
Senior level
143K-304K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead and grow a software engineering team to design, build, and operate large-scale distributed cloud services that automate infrastructure maintenance, improve reliability, and support GPU-based AI platforms. Drive technical vision, partner across engineering and operations, modernize platforms, and use data-driven decisions to improve availability, performance, and operational efficiency at global scale.
Top Skills: AzureCC#C++Distributed SystemsGpu ClustersJavaJavaScriptPython
8 Days Ago
Remote
United States
143K-331K Annually
Senior level
143K-331K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead and grow an engineering team building core VM and container platform capabilities for Azure. Drive technical strategy, execution, incident resolution, cross-team coordination, and scalable, secure platform innovations.
Top Skills: AzureCC#C++Confidential ComputingContainersDevice DriversDistributed SystemsEdgeFirmwareJavaJavaScriptOperating SystemsPythonVirtual MachinesVirtualizationWindows

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account