Top Senior Site Reliability Engineer Jobs in Austin, TX

Reposted 5 Hours AgoSaved
Remote
Austin, TX
Internship
Internship
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills: BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Reposted 10 Days AgoSaved
In-Office
Austin, TX
208K-269K Annually
Senior level
208K-269K Annually
Senior level
Artificial Intelligence • Software
Own and scale the GPU compute fleet: build metrics, alerting, observability, and repair automation; design GPU qualification/burn-in pipelines; own Redfish/BMC tooling and firmware telemetry; run incidents, eliminate toil, and deliver end-to-end reliability and orchestration for Kubernetes and bare-metal compute at hyperscale.
Top Skills: Agentic FrameworksBare MetalBmcCadenceClaude CodeCursorFirmwareGoGpuGrafanaIpmiKubernetesLlm ApisMcp ServersPrometheusPythonRedfishTemporal
Reposted YesterdaySaved
Remote
Austin, TX
115K-135K Annually
Mid level
115K-135K Annually
Mid level
Aerospace • Manufacturing
As a Site Reliability Engineer, you'll build and manage observability platforms for satellite communications, define SLOs/SLIs, and collaborate on incident response and deployment automation.
Top Skills: ArgocdAWSElkGCPGoGrafanaIstioJaegerKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
Reposted YesterdaySaved
Remote
Austin, TX
140K-165K Annually
Senior level
140K-165K Annually
Senior level
Hardware • Machine Learning • Security • Software
Design and maintain developer experience tooling and CI/CD pipelines, manage self-service deployment tools (secrets, rollbacks), partner with cloud teams to troubleshoot production systems, participate in on-call rotation, and write production-grade automation and tests in Go and TypeScript to improve developer velocity and reliability.
Top Skills: AWSGithub ActionsGoHelmKubernetesTerraformTypescript
2 Days AgoSaved
Remote
Austin, TX
Entry level
Entry level
Cybersecurity
Assist with monitoring system performance, uptime, and reliability; support incident response, troubleshooting, and root cause analysis; monitor SRE alerts in Slack and report issues; learn cloud platforms; and collaborate with development and DevOps teams.
Top Skills: AWSAzureDevOpsDevsecopsGCPGrafanaKubernetesPrometheusSlack
11 Days AgoSaved
In-Office
Austin, TX
Senior level
Senior level
Blockchain • Fintech • Cryptocurrency
Lead design, implementation, and operation of reliable hybrid cloud and on-prem infrastructure. Deliver complex initiatives, build immutable infrastructure with IaC, improve observability and incident response, mentor engineers, and ensure reliability, scalability, and compliance in change-controlled environments.
Top Skills: AnsibleAWSDatadogGithub ActionsHelmKubernetesKustomizeMakefilesPythonTerraformVirtualization
Reposted 7 Days AgoSaved
Remote or Hybrid
Austin, TX
Senior level
Senior level
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills: .NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
3 Days AgoSaved
In-Office or Remote
Austin, TX
138K-171K Annually
Junior
138K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application and network security, reliability, speed, and capacity; automate cloud deployments; monitor and troubleshoot services to meet SLAs; analyze logs and events; and resolve infrastructure issues across large-scale cloud infrastructure.
Top Skills: AnsibleAWSBashBitbucketDatabricksDatadogDigicertGradleJenkinsNode.jsPythonSplunkTenableTerraformThreatmetrix
13 Days AgoSaved
In-Office
Austin, TX
112K-191K Annually
Senior level
112K-191K Annually
Senior level
Artificial Intelligence • Big Data • Information Technology • Security • Software
Design, build, and maintain scalable infrastructure and tooling to improve availability, reliability, performance, and security. Implement SRE practices (SLI/SLO/SLA), eliminate toil, monitor telemetry, participate in incident response and on-call rotation, and collaborate cross-functionally to support global production systems.
Top Skills: AnsibleAnycastBgpC/C++CdnDnsDockerElasticsearchGitGitlabGoGrafanaHTTPJenkinsKafkaKubernetesLinuxNoSQLPrometheusPythonRdbmsRedisSaltstackTcpTls/SslUdp
Reposted 3 Days AgoSaved
Remote or Hybrid
Austin, TX
150K-225K Annually
Senior level
150K-225K Annually
Senior level
Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Lead architecture and implementation of reliability platforms and SRE practices for a production SaaS. Build self-service reliability tooling, drive AIOps automation, advance observability (monitoring, tracing, profiling), lead incident response and postmortems, mentor engineers, and embed production readiness across teams to achieve 99.99% uptime.
Top Skills: AWSAzureContinuous ProfilingDatadogDnsElkGCPGoGrafanaHttp/SKubernetesLoad BalancingOpentelemetryPrometheusPythonTcp/Ip
3 Days AgoSaved
In-Office or Remote
Austin, TX
Mid level
Mid level
Artificial Intelligence • Fintech • Machine Learning • Software • App development • Conversational AI • Generative AI
Own and improve production infrastructure reliability, deployments, Infrastructure-as-Code, Kubernetes environments, automation, CI/CD, monitoring, alerting, and observability. Investigate incidents, optimize system performance, maintain documentation and runbooks, and support DNS, WAF, CDN, and caching infrastructure. The role requires strong Linux administration, Bash scripting, networking, Git, and containerization skills, with independent ownership and collaboration across development and operations teams.
Top Skills: AkamaiAmqpAnsibleAWSBashCdnCloudflareDnsDockerGCPGitGitlab CiGrafanaHttp/HttpsKubernetesLinuxPodmanPrometheusPythonRabbitMQTerraformVictoriametricsWafZabbix
Reposted 3 Days AgoSaved
Remote
Austin, TX
Senior level
Senior level
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills: AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 3 Days AgoSaved
In-Office or Remote
Austin, TX
200K-200K Annually
Mid level
200K-200K Annually
Mid level
Cloud • Software
The Site Reliability Engineer will ensure reliable cloud operations by applying Python for infrastructure automation, managing OpenStack and Kubernetes, and practicing devsecops in a fast-paced environment.
Top Skills: KubernetesLinuxOpenstackPython
Reposted 3 Days AgoSaved
In-Office or Remote
Austin, TX
95K-171K Annually
Junior
95K-171K Annually
Junior
Cloud • Security • Software • Cybersecurity
As a Site Reliability Engineer II, you'll automate tasks, monitor AI workloads, enhance dashboards, support CI/CD processes, and collaborate with engineering teams on complex issues while participating in on-call rotations.
Top Skills: GoGrafanaKubernetesLinuxPrometheusPythonSaltstackTerraform
Reposted 3 Days AgoSaved
In-Office or Remote
Austin, TX
165K-215K Annually
Senior level
165K-215K Annually
Senior level
Software • Cybersecurity
This role involves managing Kubernetes clusters, cloud infrastructure, and CI/CD pipelines. The engineer will enhance system reliability and efficiency while troubleshooting production issues.
Top Skills: AlertmanagerAWSAzureBashCi/CdDockerElastic StackElasticsearchGCPGoGrafanaHelmKafkaKubernetesLokiMongoDBOciPrometheusPythonRedisSparkTerraform
Reposted 3 Days AgoSaved
Remote
Austin, TX
Senior level
Senior level
Software • Web3
Lead reliability practices across teams: embed early in projects, define SLIs/SLOs, build multi-cloud paved roads with Terraform, run on-call, drive org-wide incident maturity and tooling.
Top Skills: AWSAzureGCPRuby On RailsTerraformTypescriptWebcontainers
Reposted 4 Days AgoSaved
Remote
Austin, TX
Mid level
Mid level
Blockchain • Software
Build, operate, and scale production Kubernetes infrastructure using GitOps and declarative IaC. Design CI/CD workflows, observability, and secure-by-default systems. Troubleshoot networking/storage, participate in on-call rotations, automate operational workflows, and drive postmortems and reliability improvements.
Top Skills: ArbitrumArgocdArgocd ApplicationsetsAWSAzureBashCloudwatchCodebuildGCPGithub ActionsGitopsGoGrafanaK9SKubernetesLinuxLokiMimirPrometheusPrysmPythonTerraformYamlZerodev
Reposted 4 Days AgoSaved
In-Office or Remote
Austin, TX
140K-150K Annually
Mid level
140K-150K Annually
Mid level
Healthtech
Build, operate, and scale AWS cloud infrastructure and Kubernetes workloads using Terraform and Helm. Improve observability, define SLIs/SLOs, automate deployments and incident response, support on-call rotation, and implement security and compliance (HIPAA, SOC 2) best practices while partnering with product and engineering teams.
Top Skills: AWSCi/CdEvent SourcingHelmKubernetesLinuxMonitoring/Logging/TracingNetworkingTerraform
Reposted 4 Days AgoSaved
Remote
Austin, TX
120K-165K Annually
Senior level
120K-165K Annually
Senior level
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills: AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
5 Days AgoSaved
In-Office or Remote
Austin, TX
142K-268K Annually
Expert/Leader
142K-268K Annually
Expert/Leader
Automotive
Leads SRE engineering leaders and engineers while defining enterprise observability, reliability, and platform strategy across GCP, on-premise, manufacturing, distribution, and campus environments. Oversees vendor-agnostic tooling, OpenTelemetry integrations, CI/CD observability, SRE maturity models, and Agentic AI initiatives. Drives adoption of SRE practices, develops technical roadmaps, partners with senior leadership and operational teams, and maintains hands-on architectural and technical credibility.
Top Skills: Agentic AiAWSAzureCi/CdDatadogDynatraceGCPNew RelicOpentelemetryOtel Genai Semantic ConventionsSource Control PlatformsSplunkTerraform
Reposted 5 Days AgoSaved
Remote
Austin, TX
130K-160K Annually
Senior level
130K-160K Annually
Senior level
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills: AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Reposted 5 Days AgoSaved
In-Office or Remote
Austin, TX
76K-136K Annually
Mid level
76K-136K Annually
Mid level
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills: Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
Reposted 5 Days AgoSaved
Remote
Austin, TX
180K-224K Annually
Senior level
180K-224K Annually
Senior level
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills: Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
15 Days AgoSaved
In-Office
Austin, TX
Senior level
Senior level
Artificial Intelligence • Big Data • Information Technology • Security • Software
Design, build, and maintain cloud infrastructure and CI/CD for a high-availability telecommunications product. Define SLOs/SLIs, manage incident response and on-call rotations, implement observability, perform performance and capacity planning, run blameless postmortems, and collaborate with security teams to ensure compliance and access control.
Top Skills: AnsibleAWSDatadogDockerGCPGitlabHelmJavaJenkinsKubernetesNoSQLTerraform
6 Days AgoSaved
Remote
Austin, TX
38K-58K Annually
Mid level
38K-58K Annually
Mid level
Information Technology • Consulting
Designs and maintains AWS cloud infrastructure, CI/CD pipelines, Docker workloads, and IaC provisioning. Implements monitoring, logging, alerting, observability, security hardening, and code quality checks. Troubleshoots production issues, supports incident response and root-cause analysis, improves system reliability, and automates operational tasks. Collaborates with distributed engineering teams to establish infrastructure, deployment, monitoring, and SRE best practices.
Top Skills: AWSAws IamCi/CdCloudFormationCloudwatchDatadogDockerGrafanaInfrastructure As Code (Iac)KubernetesPrometheusSonarqubeTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account