Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Austin, TX
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills:
AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Logistics • Mobile • Productivity • Software • Transportation
The Senior Site Reliability Engineer will manage the reliability of Zello's data tier, contribute to monitoring and incident response while improving cloud infrastructure and database performance.
Top Skills:
BashDockerElasticsearchGoKubernetesLokiMongoDBMySQLPrometheusPythonRedisScylladbTempo
Reposted 10 Days AgoSaved
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
The Staff Engineer will define reliability architecture, automate foundational utilities, develop observability tools, ensure environment integrity, and mentor colleagues.
Top Skills:
AnsibleChefDhcpKubernetesLinuxNtpPxe
Reposted 12 Days AgoSaved
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead architecture and implementation of reliability improvements across CrowdStrike's cloud-native platform. Build shared libraries and services, drive observability and SLO practices, perform performance and cost optimization, run resilience engineering and chaos experiments, automate infrastructure-as-code, mentor engineers, and embed with product teams to deliver scalable, highly reliable distributed systems at organizational scale.
Top Skills:
AIAlertingAWSCassandraElasticsearchGCPGoInfrastructure-As-CodeJavaKafkaKotlinKubernetesNode.jsObservability (TracingOciOpensearchProfilingProtobufPythonScalaSlos)
Reposted 12 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Reposted 14 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Reposted 16 Days AgoSaved
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills:
AWSGoKubernetesPuppetPythonTerraform
Big Data • Real Estate • Software
Senior SRE responsible for reliability, observability, and operational excellence of a large AWS/Kubernetes platform. Duties include maintaining EKS/Fargate infrastructure, monitoring SLIs/SLOs, implementing observability with NewRelic, driving cost optimization and FinOps practices, executing chaos engineering and incident response, contributing automation and IaC, and supporting security/compliance and developer experience.
Top Skills:
Apollo GraphqlArgo CdAWSAws Secrets ManagerCircleCICloudFormationCloudfrontCloudwatchDatadogDockerEc2EcsEksFargateGithub ActionsGitopsGoGrafanaHelmIamIstioJavaJenkinsKubernetesKustomizeLambdaNewrelicOpsgeniePagerdutyPrometheusPythonRdsRoute53S3ServicenowSplunkTerraformTyk GatewayVaultVpc
Reposted 11 Days AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Computer Vision • Hardware • Machine Learning • Robotics • Software
The role involves maintaining cloud infrastructure, collaborating with engineering teams, troubleshooting issues, deploying solutions, and ensuring system reliability.
Top Skills:
AnsibleC++GrafanaHelmKubernetesPagerdutyPythonTerraformTypescript
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own Garner’s cloud reliability strategy across AWS and Kubernetes, including SLOs, observability, incident response, infrastructure automation, cost optimization, and security compliance. Lead complex incident resolution, architect Terraform-based infrastructure, establish deployment and monitoring standards, mentor engineers, and use AI tools to automate operational work. Support high-scale AI/ML workloads while setting technical direction for platform reliability and production quality.
Top Skills:
AWSClaudeDatadogGitlabGoIstioKubernetesNatsPostgresPythonTerraformTypescript
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead the design and roadmap for global Active Directory and identity infrastructure, implement Identity-as-Code and GitOps automation, own incident escalation and observability, define delegation/tiered administration, integrate applications with Okta and cloud identity, mentor teams, and publish identity architecture and security best practices.
Top Skills:
Active Directory Domain Services (Ad Ds)AnsibleAWSAws Directory ServiceAzureAzure Active Directory (Entra Id)Azure SentinelCertificate ServicesChefDhcpDnsGCPGitopsGroup Policy Objects (Gpo)New RelicOktaPowershellPowershell DscPythonTerraform
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Build Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with Wiz, lead major incidents, coordinate remediation, communicate status to stakeholders, and create incident-management playbooks and runbooks. The role also mentors SREs and supports platforms subject to PCI-DSS and SOC 2 compliance.
Top Skills:
AWSAzureCi/CdCnappCspmGoGoogle Cloud PlatformKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Reposted 20 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Artificial Intelligence • Cloud • Information Technology • Consulting
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
Top Skills:
BashCi/CdDockerGrafanaLangchainLlmLokiMakefileNestjsOpentelemetryOracle CloudPgbackrestPostgresql 15PrometheusPythonRedisTempoTerraformTraefik
Edtech • Kids + Family • Sports
Audit infrastructure, deployment pipelines, monitoring, alerting, incident response, on-call practices, and internal tools. Produce actionable audit reports, implement code and configuration fixes, improve SLOs and reliability practices, advise on scalable architecture, and partner with engineers on implementation and handoff. The role is a fully remote, part-time consulting engagement with potential for full-time conversion.
Top Skills:
Ai Coding ToolsAWSCi/CdDatadogGrafanaPrometheus
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills:
DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills:
.NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
Healthtech • Software
Manage and optimize a multi-account AWS environment supporting production healthcare applications and analytics platforms. Responsibilities include AWS infrastructure administration, CI/CD and infrastructure automation, observability, incident response, on-call support, disaster recovery validation, HIPAA/HiTrust compliance, IAM and security controls, and support for containerized Java and Python applications. The role also contributes to Kubernetes and EKS modernization initiatives and partners with developers to resolve complex production issues.
Top Skills:
Amazon EksApi GatewayArgocdAuroraAWSAws CloudformationAws CodepipelineAws Security HubBashCloudfrontCloudwatchCortex CloudDatadogDockerEc2EcsFargateGitGuarddutyHelmIamJavaJenkinsKubernetesLambdaLinuxPrismaPythonRdsS3Spring BootUbuntuZabbix
Software
Owns reliability, observability, performance, and security for a multi-region SaaS platform. Responsibilities include managing Datadog, implementing APM and tracing, defining SLOs, developing automation, expanding infrastructure as code and CI/CD, automating operational workflows, maintaining security controls, participating in incident response, documenting procedures, and mentoring engineers.
Top Skills:
ApmAzure DevopsAzure Kubernetes ServiceAzure SqlBashBicepCi/CdCosmos DbDatadogDistributed TracingHelmInfrastructure As CodeKey VaultKubernetesKustomizeManaged IdentitiesAzureMicrosoft Entra IdPowershellPythonRedisService BusTerraform
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Austin, TX Companies Hiring SRE Engineers
See AllPopular Austin, TX Engineering Job Searches
Engineering Jobs in Austin, TX
.NET Developer Jobs in Austin, TX
Android Developer Jobs in Austin, TX
Application Engineer Jobs in Austin, TX
Automation Engineer Jobs in Austin, TX
Backend Engineer Jobs in Austin, TX
C# Jobs in Austin, TX
C++ Jobs in Austin, TX
Cloud Engineer Jobs in Austin, TX
Controls Engineer Jobs in Austin, TX
CTO Jobs in Austin, TX
Design Engineer Jobs in Austin, TX
DevOps Engineer Jobs in Austin, TX
DevOps Jobs in Austin, TX
Director of Engineering Jobs in Austin, TX
Electrical Engineering Jobs in Austin, TX
Embedded Software Engineer Jobs in Austin, TX
Engineering Manager Jobs in Austin, TX
Field Engineer Jobs in Austin, TX
Front End Developer Jobs in Austin, TX
Full Stack Developer Jobs in Austin, TX
Golang Jobs in Austin, TX
Hardware Engineer Jobs in Austin, TX
Infrastructure Engineer Jobs in Austin, TX
iOS Developer Jobs in Austin, TX
Java Developer Jobs in Austin, TX
Javascript Jobs in Austin, TX
Linux Jobs in Austin, TX
Manufacturing Engineer Jobs in Austin, TX
Mechanical Design Engineer Jobs in Austin, TX
Mechanical Engineering Jobs in Austin, TX
Network Engineer Jobs in Austin, TX
PHP Developer Jobs in Austin, TX
Platform Engineer Jobs in Austin, TX
Principal Software Engineer Jobs in Austin, TX
Process Engineer Jobs in Austin, TX
Product Engineer Jobs in Austin, TX
Project Engineer Jobs in Austin, TX
Python Jobs in Austin, TX
QA Engineer Jobs in Austin, TX
QA Jobs in Austin, TX
Robotics Engineer Jobs in Austin, TX
Ruby Jobs in Austin, TX
Salesforce Developer Jobs in Austin, TX
Scala Jobs in Austin, TX
Security Engineer Jobs in Austin, TX
Software Engineer Jobs in Austin, TX
Software Engineering Manager Jobs in Austin, TX
Software Test Engineer Jobs in Austin, TX
Solutions Architect Jobs in Austin, TX
Solutions Engineer Jobs in Austin, TX
SRE Engineer Jobs in Austin, TX
Staff Software Engineer Jobs in Austin, TX
Systems Engineer Jobs in Austin, TX
Web Developer Jobs in Austin, TX
All Filters
Total selected ()
No Results
No Results
.png)












.png)
















