Microsoft Logo

Microsoft

Principal Software Engineer, AI Infra Management and Ops

Posted 11 Hours Ago
Remote
Hiring Remotely in United States
120K-304K Annually
Senior level
Remote
Hiring Remotely in United States
120K-304K Annually
Senior level
Lead the design and evolution of cloud-native platforms, distributed systems, and infrastructure services at large scale. Define architecture across compute, networking, storage, datacenter, and AI infrastructure environments. Improve automation, lifecycle management, observability, reliability, security, and developer productivity. Guide technical strategy across engineering organizations while contributing to GPU, HPC, bare-metal, networking, and distributed computing platforms. Mentor engineers and foster technical excellence through architecture reviews and cross-organizational collaboration.
The summary above was generated by AI
Overview
Help shape the future of cloud infrastructure and artificial intelligence platforms at unprecedented scale. Our team is building the foundational platform capabilities that power large-scale distributed systems, accelerated computing environments, and next-generation cloud services. You will work at the intersection of software, infrastructure, networking, and datacenter technologies to enable reliable, secure, and scalable platforms that support a broad range of business-critical workloads. You will collaborate with engineering teams across Microsoft to advance platform capabilities, accelerate innovation, and improve operational excellence across diverse infrastructure environments.
 
As a Principal Software Engineer, you will lead the design and evolution of cloud-native platforms, distributed systems, and infrastructure services that support service lifecycle management, resource orchestration, observability, reliability, and automation. You will help guide technical strategy across multiple engineering investments, contribute to architecture decisions involving compute, networking, storage, and platform services, and collaborate across organizations to deliver capabilities that improve platform efficiency and developer productivity. This opportunity will allow you to deepen your expertise in distributed systems and cloud infrastructure, expand your influence across engineering organizations, and contribute to innovations in Artificial Intelligence (AI) infrastructure, graphics processing unit (GPU) platforms, high-performance networking, and large-scale datacenter systems.
 
Microsoft's mission is to empower every person and every organization on the planet to achieve more. We cultivate a culture grounded in respect, integrity, accountability, inclusion, collaboration, and continuous learning, enabling individuals and teams to grow, contribute, and make meaningful impact.
 
 

Responsibilities
  • Lead the design, development, and evolution of cloud-native services, distributed systems, and platform capabilities that support large-scale production environments.
  • Collaborate across engineering organizations to help define technical direction and architecture for platform infrastructure, cloud services, networking, storage, compute, and operational excellence investments.
  • Guide the development of platform engineering capabilities that improve deployment automation, service lifecycle management, observability, reliability, scalability, and developer productivity.
  • Advance software engineering practices including Continuous Integration and Continuous Delivery (CI/CD), Infrastructure as Code (IaC), testing, security, monitoring, and operational readiness throughout the engineering lifecycle.
  • Contribute to the design and operation of Artificial Intelligence (AI) infrastructure, High Performance Computing (HPC) platforms, bare-metal infrastructure, graphics processing unit (GPU) environments, and large-scale distributed computing systems.
  • Collaborate on datacenter architecture and networking solutions, including rack-scale systems, network topology design, Ethernet fabrics, InfiniBand fabrics, Remote Direct Memory Access (RDMA), Smart Network Interface Cards (SmartNICs), Data Processing Units (DPUs), and accelerated computing platforms.
  • Foster a culture of technical excellence through architecture collaboration, mentorship, knowledge sharing, design review participation, and support for engineering excellence across the broader organization.

Qualifications

Required Qualifications: 

  • Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Other Requirements:

Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. This includes passing the Microsoft Cloud background check upon hire/transfer and every two years thereafter. 

Preferred Qualifications: 

  • Recognized for delivering technical solutions that span multiple engineering teams, organizations, or business areas.
  • Demonstrated ability to simplify complex technical challenges and translate them into practical, scalable, and maintainable solutions.
  • Proven ability to influence technical direction through collaboration, architectural leadership, technical mentoring, and cross-organizational partnerships.
  • Knowledge of large-scale distributed systems, cloud-native platforms, or infrastructure services supporting business-critical production environments.
  • Familiarity with Artificial Intelligence (AI) infrastructure, High Performance Computing (HPC) environments, accelerated computing platforms, or large-scale graphics processing unit (GPU) deployments.
  • Knowledge of datacenter architecture, rack-scale infrastructure, network topology design, and modern datacenter networking technologies.
  • Familiarity with InfiniBand fabrics, NVLink, NVSwitch, Ethernet networking, Remote Direct Memory Access (RDMA), Smart Network Interface Cards (SmartNICs), or Data Processing Units (DPUs) supporting large-scale compute environments.
  • Understanding of bare-metal infrastructure, hardware lifecycle management, fleet operations, infrastructure telemetry, or platform operations at scale.
  • Proficiency in one or more modern programming languages such as Go, Rust, C#, Java, or Python.

#AIINFRA


Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar Jobs

A Minute Ago
In-Office or Remote
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software • Defense
Own platform reliability, observability, incident response, scaling, capacity planning, and deployment automation. Monitor system health, debug and resolve incidents, build logging and monitoring tools, improve CI/CD pipelines, develop self-service automation, and strengthen high-availability delivery systems. The role requires clear incident communication, post-incident learning, and collaboration across engineering teams in secure, high-side environments.
Top Skills: AWSAws GovcloudBashCi/CdDatadogDockerElasticsearchOpensearchPlg StackPulumiPythonSQLTerraform
2 Minutes Ago
In-Office or Remote
160K-380K Annually
Senior level
160K-380K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Own and grow enterprise relationships with AI startups in the Bay Area ecosystem. Identify and win high-potential accounts, serve as the primary business and technical contact, manage renewals and upsells, advocate for customers internally, and collaborate with marketing, product, engineering, support, and customer success teams. Translate customer feedback into product insights while expanding DigitalOcean’s cloud infrastructure and AI business.
Top Skills: Artificial Intelligence (Ai)Cloud Infrastructure
11 Minutes Ago
Remote or Hybrid
United States
71K-88K Annually
Senior level
71K-88K Annually
Senior level
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Develops strategy and operations for the Loyalty & VIP Growth organization by building roadmaps, scalable processes, KPI frameworks, competitive intelligence, and data-driven recommendations. Identifies AI and automation opportunities, builds productivity tools and agents, and partners cross-functionally to execute strategic initiatives and improve business performance.
Top Skills: AIAutomationClaudeGoogle ApplicationsMicrosoft ApplicationsSnowflakeSQL

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account