Andromeda (andromeda.ai) Jobs

Site Reliability Engineer - AI Infrastructure

Andromeda (andromeda.ai)

Site Reliability Engineer - AI Infrastructure

Reposted 3 Days Ago

In-Office or Remote

Hiring Remotely in United States

Senior level

In-Office or Remote

Hiring Remotely in United States

Senior level

The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.

The summary above was generated by AI

Site Reliability Engineer - AI Infrastructure

Location: Global Remote / San Francisco · Full-Time

About Andromeda

Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.

We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.

Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.

Our long-term vision is to build the liquidity layer for global AI compute — a marketplace that moves the infrastructure and workloads powering AGI not dissimilar to the flows of capital in the world's financial markets.

We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.

What You’ll Do

Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers.
Build automation and tooling to streamline cluster deployments and integrations.
Debug customer issues across networking, storage, scheduling, and system layers.
Improve reliability and scalability of both training and inference infrastructure.
Design and implement monitoring, alerting, and observability for critical systems.
Collaborate with engineering and product teams to plan and deliver infrastructure for new services.
Participate in on-call and incident response, leading postmortems and reliability improvements.
What We’re Looking For

5+ years experience in SRE, DevOps, or infrastructure engineering roles.
Strong Linux systems and networking fundamentals.
Deep experience with Kuber

Kubernetes and container orchestration at scale.

Proficiency with Infrastructure-as-Code (Terraform, Helm, Ansible, etc.).
Strong automation and scripting skills (Python, Go, or Bash).
Experience with observability stacks (Prometheus, Grafana, Loki, Datadog, etc.).
Track record of operating production systems and leading incident response.

Nice to Have

Exposure to ML/AI infrastructure or GPU-based systems (CUDA, Slurm, Triton, etc.).
Familiarity with high-performance networking (InfiniBand, NVLink) or distributed storage (VAST, Weka, Ceph).
Customer-facing support or consulting experience.

Why You’ll Love It Here

This is a builder’s role. You’ll have ownership and autonomy to shape how our systems run, working directly with customers and providers while building the foundation for reliable, scalable AI infrastructure.

Similar Jobs

Andromeda (andromeda.ai)

Senior Site Reliability Engineer

3 Days Ago

In-Office or Remote

United States

Senior level

Artificial Intelligence • Cloud • Information Technology • Software

Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.

Top Skills: AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform

UL Solutions

Field Business Manager - Northwest Region

2 Hours Ago

Remote or Hybrid

128K-145K Annually

Expert/Leader

128K-145K Annually

Expert/Leader

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy

Manage regional field inspection and audit operations to meet budgets and customer requirements. Oversee staffing, performance management, continuous improvement, client technical engagement, and adherence to UL policies and procedures. Drive customer satisfaction, subcontractor relationships, and implementation of preventive/corrective actions.

Square

Software Engineer

4 Hours Ago

Remote or Hybrid

CA, USA

121K-213K Annually

Mid level

121K-213K Annually

Mid level

eCommerce • Fintech • Hardware • Payments • Software • Financial Services

Build and operate ingestion and reconciliation pipelines that match card-network and partner reports against internal transactions. Develop reporting products, produce standardized journals for accounting, support tax and regulatory filing systems, handle sensitive PII for compliance, debug production issues, and collaborate with finance, product, and data teams to automate and scale reconciliation and reporting.

Top Skills: Ai ToolsAirflowAWSBigQueryCi/CdDelta LakeGoHadoopIso-8583JavaKafkaKubernetesPysparkPythonSnowflakeSparkSQLTemporalTerraform

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center