Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Austin, TX
Gaming • Information Technology • Mobile • Software • Esports
Lead design, build, and operation of multi-cloud hybrid infrastructure and Kubernetes platforms. Drive observability, SLI/SLOs, incident response, automation, CI/CD hardening, secrets/policy-as-code, and promote SRE practices across studios.
Top Skills:
1PasswordAnsibleArgocdAWSAws Secrets ManagerAws Systems ManagerBare MetalCiliumDatadogEksFluxGCPGithub ActionsGkeGoGrafanaHelmIstioJenkinsKubernetesOpa/GatekeeperOpentelemetryPasswordstatePrometheusPulumiPuppetPythonTerraformTerragruntTypescriptVMware
Big Data • Real Estate • Software
Senior SRE responsible for reliability, observability, and operational excellence of a large AWS/Kubernetes platform. Duties include maintaining EKS/Fargate infrastructure, monitoring SLIs/SLOs, implementing observability with NewRelic, driving cost optimization and FinOps practices, executing chaos engineering and incident response, contributing automation and IaC, and supporting security/compliance and developer experience.
Top Skills:
Apollo GraphqlArgo CdAWSAws Secrets ManagerCircleCICloudFormationCloudfrontCloudwatchDatadogDockerEc2EcsEksFargateGithub ActionsGitopsGoGrafanaHelmIamIstioJavaJenkinsKubernetesKustomizeLambdaNewrelicOpsgeniePagerdutyPrometheusPythonRdsRoute53S3ServicenowSplunkTerraformTyk GatewayVaultVpc
Computer Vision • Hardware • Machine Learning • Robotics • Software
The role involves maintaining cloud infrastructure, collaborating with engineering teams, troubleshooting issues, deploying solutions, and ensuring system reliability.
Top Skills:
AnsibleC++GrafanaHelmKubernetesPagerdutyPythonTerraformTypescript
Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Lead reliability efforts for cloud-native production systems: design and operate infrastructure, define SLOs/SLIs, lead incident response, build IaC and CI/CD, improve observability and automate toil, and mentor SRE engineers.
Top Skills:
AWSAzureCassandraCdnCloudFormationDnsEcsElkGCPGithub ActionsGitopsGoGrafanaJavaJenkinsKubernetesLinuxMySQLNewrelicOraclePagerdutyPostgresPrometheusPythonRedisSplunkTcp/IpTerraform
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Lead SRE work to ensure high readiness of hybrid on-prem/cloud test benches. Implement SLOs/SLIs, automate provisioning (DHCP/PXE/NTP), build observability and alerting, stabilize data ingestion paths, eliminate manual toil through IaC and automation, and mentor peers to improve resilience and recovery.
Top Skills:
AnsibleChefDhcpGoKubernetesLinuxNtpPxePythonTerraform
Reposted 4 Days AgoSaved
Cloud • Software
Responsible for maintaining FedRAMP-compliant infrastructure, collaborating with software engineers, and ensuring system availability and security. Duties include infrastructure design, automation, monitoring, and incident response.
Top Skills:
AWSGoKubernetesPuppetPythonTerraform
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
In this role, you'll ensure the reliability and scalability of the NG-SIEM platform, manage incident responses, and collaborate across teams to enhance system performance.
Top Skills:
BashC++GoJavaKafkaPythonRust
Reposted YesterdaySaved
Software • Defense
Own reliability, scalability, and security for on-prem and AWS deployments. Build observability (Prometheus/Loki/Grafana/ELK), define SLOs/SLIs, lead incident response and postmortems, automate infrastructure (Terraform/Ansible), operate Kubernetes clusters, embed security/compliance controls, eliminate operational toil, and mentor teams.
Top Skills:
AlloyAnsibleAWSAws GovcloudBashCloudFormationDatadogElkGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubernetesLokiPrometheusPythonRmfStigsTerraform
Reposted YesterdaySaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Big Data • Cloud • Software • Database
Seeking a Site Reliability Engineer with expertise in networking and distributed systems for building secure multi-cloud infrastructure. Responsibilities include maintaining network architecture and ensuring reliable service-to-service communication, involving a 24/7 on-call rotation.
Top Skills:
AWSAzureBgpDnsGCPIpv6KubernetesLoad BalancingMtlsService MeshTcp/IpTlsVpcsVpns
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Insurance
As a Site Reliability Engineer II, you will build, test, and maintain the technology infrastructure for Openly's insurance platform, focusing on automation, monitoring, incident response, and operational decisions.
Top Skills:
Aiven DebeziumArcgisBigQueryCircleCICloud FunctionsCloud RunCloudsqlComposer/AirflowDatadogFivetranGcp GcsGitGoGCPJupyter NotebooksKafkaKubernetesNuxtPostgresPub/SubPythonRSQLTailwindTerraformVuejsWebpack
Information Technology • Insurance • Software
Own and operate production services end-to-end to ensure reliability, scalability, performance, and operational health. Define SLIs/SLOs, perform incident response and root cause analysis, build automation and self-healing, manage production changes, and collaborate with engineering, product, and operations teams to improve system design and observability.
Top Skills:
.NetAWSC#Ci/CdInfrastructure As CodeJavaKubernetesLinuxPythonReactRelational DatabasesWindows
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills:
Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
Artificial Intelligence • Cloud • Consumer Web • Productivity • Software • App development • Data Privacy
The role involves defining reliability strategies, leading initiatives across teams, enhancing monitoring and incident response, and mentoring engineers at Dropbox.
Top Skills:
Ai TechnologiesDebuggingDistributed SystemsIncident ResponseObservabilityReliability Risk ManagementSlasSlos
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead long-term strategy and architecture for cloud and on‑prem platform infrastructure, driving Kubernetes and multi‑cloud reliability, IaC/GitOps automation, observability, SLO/SLI/error‑budget practices, incident leadership, AI‑augmented tooling adoption, and mentorship of senior engineers to improve platform resilience and developer experience.
Top Skills:
Amazon Elastic Kubernetes Service (Eks)AutoscalingAWSCapacity PlanningCi/CdGitopsGoGoogle Cloud PlatformGoogle Kubernetes Engine (Gke)Identity And Access ManagementInfrastructure As CodeKubernetesLinuxNetworkingObservabilityOperatorsPulumiPythonRke2StorageTerraform
Reposted 17 Days AgoSaved
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
The Staff Engineer will define reliability architecture, automate foundational utilities, develop observability tools, ensure environment integrity, and mentor colleagues.
Top Skills:
AnsibleChefDhcpKubernetesLinuxNtpPxe
Logistics • Mobile • Productivity • Software • Transportation
The Senior Site Reliability Engineer will manage the reliability of Zello's data tier, contribute to monitoring and incident response while improving cloud infrastructure and database performance.
Top Skills:
BashDockerElasticsearchGoKubernetesLokiMongoDBMySQLPrometheusPythonRedisScylladbTempo
Reposted 17 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Reposted 10 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 11 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
As a Cloud Cost Utilization SRE at GitLab, you'll manage cloud spending, improve tracking and optimization of cloud usage, and collaborate with finance and engineering teams to enhance cost efficiency across AWS and GCP.
Top Skills:
AnsibleAWSElkGCPGrafanaLokiMimirPrometheusTempoTerraform
Reposted 21 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
eCommerce • Legal Tech • Professional Services • Software • Data Privacy
The Site Reliability Engineer will ensure systems run smoothly, work with automation tools, resolve issues, and drive operational improvements.
Top Skills:
AWSAzureCloudFormationDockerGCPGrafanaKubernetesMemcachedNew RelicOpentelemetryPostgresPrometheusPulumiRedisSentryTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Austin, TX Engineering Job Searches
Engineer Jobs in Austin, TX
.NET Developer Jobs in Austin, TX
Android Developer Jobs in Austin, TX
Application Engineer Jobs in Austin, TX
Automation Engineer Jobs in Austin, TX
Backend Engineer Jobs in Austin, TX
C# Jobs in Austin, TX
C++ Jobs in Austin, TX
Cloud Engineer Jobs in Austin, TX
Controls Engineer Jobs in Austin, TX
CTO Jobs in Austin, TX
Design Engineer Jobs in Austin, TX
DevOps Engineer Jobs in Austin, TX
DevOps Jobs in Austin, TX
Director of Engineering Jobs in Austin, TX
Electrical Engineer Jobs in Austin, TX
Embedded Software Engineer Jobs in Austin, TX
Engineering Manager Jobs in Austin, TX
Field Engineer Jobs in Austin, TX
Front End Developer Jobs in Austin, TX
Full Stack Developer Jobs in Austin, TX
Golang Jobs in Austin, TX
Hardware Engineer Jobs in Austin, TX
Infrastructure Engineer Jobs in Austin, TX
iOS Developer Jobs in Austin, TX
Java Developer Jobs in Austin, TX
Javascript Jobs in Austin, TX
Linux Jobs in Austin, TX
Manufacturing Engineer Jobs in Austin, TX
Mechanical Design Engineer Jobs in Austin, TX
Mechanical Engineer Jobs in Austin, TX
Network Engineer Jobs in Austin, TX
PHP Developer Jobs in Austin, TX
Platform Engineer Jobs in Austin, TX
Principal Software Engineer Jobs in Austin, TX
Process Engineer Jobs in Austin, TX
Product Engineer Jobs in Austin, TX
Project Engineer Jobs in Austin, TX
Python Jobs in Austin, TX
QA Engineer Jobs in Austin, TX
QA Jobs in Austin, TX
Robotics Engineer Jobs in Austin, TX
Ruby Jobs in Austin, TX
Salesforce Developer Jobs in Austin, TX
Scala Jobs in Austin, TX
Security Engineer Jobs in Austin, TX
Software Engineer Jobs in Austin, TX
Software Engineering Manager Jobs in Austin, TX
Software Test Engineer Jobs in Austin, TX
Solutions Architect Jobs in Austin, TX
Solutions Engineer Jobs in Austin, TX
SRE Engineer Jobs in Austin, TX
Staff Software Engineer Jobs in Austin, TX
Systems Engineer Jobs in Austin, TX
Web Developer Jobs in Austin, TX
All Filters
Total selected ()
No Results
No Results









.png)






















