Microsoft Logo

Microsoft

Infrastructure Data & Analytics

Reposted 7 Days Ago
Remote
Hiring Remotely in United States
143K-331K Annually
Senior level
Remote
Hiring Remotely in United States
143K-331K Annually
Senior level
Lead infrastructure analytics for compute, storage, and networking: design scalable data pipelines, define core metrics, build dashboards and APIs, ensure data quality and governance, drive instrumentation fixes, and advise executives on capacity, utilization, and readiness.
The summary above was generated by AI
Overview

We are seeking experienced Infrastructure Data & Analytics Engineers to join our Microsoft AI team and own the end-to-end technical vision and execution for infrastructure analytics, turning raw telemetry into trusted, decision-quality insights on utilization, capacity, readiness, and efficiency. This role is critical to helping the Microsoft AI, SuperIntelligence leadership make informed investment and planning decisions at scale.


Microsoft AI
This role is part of Microsoft AI. Our Superintelligence team is a startup-like organization within Microsoft, dedicated to pushing the boundaries of artificial intelligence while maintaining a strong commitment to safety, responsibility, and human values.
Our mission is to build AI that amplifies human potential and empowers people around the world. We strive to deliver breakthroughs that advance science, education, productivity, and global well-being.
We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models!

MAI employees are expected to work from a designated Microsoft office at least four days a week if they live within 50 miles (U.S.) or 25 miles (non-U.S., country-specific) of that location. This expectation is subject to local law and may vary by jurisdiction.   


Responsibilities
  • Act as the technical lead and owner for infrastructure analytics across compute, storage, and networking.
  • Design and build durable, scalable data pipelines that ingest telemetry from clusters, schedulers, health systems, and capacity trackers into Data Warehouse
  • Define and standardize core metrics and semantics (e.g., utilization, occupancy, MFU, goodput, capacity readiness, delivery-to-production).
  • Architect and maintain self-service dashboards and APIs for fleet, cluster, and squad-level visibility.
  • Partner closely with stakeholders across Supercomputing Infra, Researchers, Strategy and Executives to ensure metrics reflect operational and business reality.
  • Implement robust and fault-tolerant systems for data ingestion and processing.
  • Lead data architecture and engineering decisions, applying strong technical judgment to proactively shape executive-level discussions and decisions.
  • Identify data gaps and instrumentation issues; drive fixes by influencing upstream engineering teams.
  • Establish data quality, validation, documentation, and governance so metrics are trusted and repeatable.


Qualifications

Required Qualifications:

  • Bachelor’s degree in computer science, or related technical field AND 8+ years technical engineering experience with data engineering, analytics, or data science, with increasing technical ownership in startup environment AND 6+ years experience with distributed data processing frameworks and large-scale data systems
    • OR equivalent experience.

Preferred Qualifications:

  • Master's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with technical engineering experience with data engineering, analytics, or data science, with increasing technical ownership in startup environment AND 10+ years experience with distributed data processing frameworks and large-scale data systems
    • OR equivalent experience.
  • Proven technical leadership in data engineering, analytics platforms, or large-scale telemetry systems.
  • Hands-on experience with ETL orchestration frameworks such as Airflow, Dagster, or similar.
  • Strong communication skills; can explain complex systems clearly to senior leader.


Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay

Software Engineering IC6 - The typical base pay range for this role across the U.S. is USD $165,600 - $296,400 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $220,800 - $331,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

Similar Jobs

An Hour Ago
Remote or Hybrid
120K-225K Annually
Expert/Leader
120K-225K Annually
Expert/Leader
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Leads Compensation Operations for monthly commission and compensation processing serving approximately 1,500 agents and managers. Oversees an eight-person team, ensures accurate and timely monthly close, drives process improvement and automation, promotes AI adoption, reconciles transactions across finance systems and carrier statements, partners with compensation design and technology teams, and advises organizational leadership on compensation operations.
Top Skills: Artificial IntelligenceCompensation Technology
An Hour Ago
Remote
US
Senior level
Senior level
Consumer Web • eCommerce • Machine Learning • Software • Sports • Analytics
Design and develop scalable cloud-based backend architectures for PSA’s consumer-facing Set Registry platform. Own features through architecture, implementation, testing, deployment, and production release. Build APIs for web and mobile applications, improve databases and existing applications, troubleshoot complex systems, and collaborate with product and engineering teams. Contribute to technical design reviews, documentation, monitoring, security practices, and automated testing.
Top Skills: Amazon Api GatewayAmazon EcsAmazon EventbridgeAmazon SnsAmazon SqsAws CdkAws CloudformationAws LambdaBatch ProcessingDockerDynamoDBEvent-Driven ArchitecturesKubernetesMicroservicesPostgresPythonRest ApisTerraform
2 Hours Ago
Remote or Hybrid
United States
Senior level
Senior level
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Develop autonomy evaluation metrics, statistical and machine-learning analyses, dashboards, and pipelines for simulation and on-road autonomous vehicle testing. Analyze perception, prediction, and planning performance; identify anomalous behavior and critical scenarios; apply VLMs and LLMs where appropriate; and support release gating and safety decisions. The role requires production Python, C++ debugging, technical leadership, cross-functional collaboration, and rigorous software engineering practices.
Top Skills: C++Computational GeometryLarge Language ModelsLinear AlgebraMachine LearningNumpyPandasPythonPyTorchRosScipySQLStatistical ModelingTime-Series AnalysisVision-Language Models

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account