Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Austin, TX
Software
As a Site Reliability Engineer, you'll enhance system reliability, collaborate on production readiness, define SLIs/SLOs, and improve incident response.
Top Skills:
AWSDatadogGrafanaKubernetesOpentelemetryPrometheusTypescript
Digital Media • Edtech
Drive reliability and observability of Epic's GCP-based platform. Own cloud infrastructure, container platform (Kubernetes/GKE), CI/CD, observability, and IaC (Terraform). Define SLOs/SLIs, reduce toil, manage security and compliance practices, participate in on-call rotations, lead incident response and post-mortems, and partner with product and data teams to troubleshoot and improve platform reliability.
Top Skills:
ArgocdBashCloud MonitoringDockerGceGCPGcsGithub ActionsHelmIamJenkinsKubernetes (Gke)New RelicPythonTerraformVpc
Information Technology • Cybersecurity
Lead global SRE team to design, implement, and operate scalable, highly available cloud infrastructure. Drive observability, automation (IaC/CI-CD), cost optimization (FinOps), security hygiene, and post-incident improvements while remaining hands-on and exploring AI tooling for reliability.
Top Skills:
Ai ToolingAlloyAWSAzureCi/CdCloudFormationContainer OrchestrationDatadogEdge ComputingFinopsGCPGrafanaGrafana IrnInfrastructure-As-CodeKubernetesLokiPrometheusPulumiServerlessSplunkSpot InstancesTerraform
Logistics • Software
Own and operate scalable infrastructure on GCP (GKE, Cloud Run, AlloyDB); author Terraform modules; manage containerized workloads; build observability in Datadog; design CI/CD in GitHub Actions; automate operational workflows; lead incident response and post-mortems; partner with engineers to improve reliability, cost, and automation.
Top Skills:
AlloydbClickhouseCloud RunCloudflare WorkersDatadogDockerGCPGithub ActionsGkeGoGrafanaIamKafkaKubernetesNetworkingPostgresPrometheusPub/SubPythonRedisRedpandaTerraformTypescript
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills:
AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Develop and maintain code and automation for highly scalable M365 sovereign cloud services. Operate live sites, participate in on-call rotations, troubleshoot incidents, improve observability, design automation for deployments, and collaborate with product and security stakeholders to ensure reliability and compliance.
Top Skills:
AzureCC#C++CopilotExchange Online ProtectionExchange TransportGenerative AiJavaJavaScriptMicrosoft 365Microsoft Defender For OfficeOffice 365OnedrivePurviewPythonSharepointTeams
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead qualification, performance validation, and production readiness for new Azure Storage hardware and firmware. Drive test planning, automation frameworks, large-scale telemetry and benchmark analysis, root-cause investigations across software/hardware/firmware, and partner with engineering and vendors to resolve reliability and performance issues.
Top Skills:
Automation FrameworksAzureAzure StorageBenchmarkingFirmwareNetworkingSsdsTelemetry
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, develop, deploy, manage, and monitor Azure infrastructure and services. Serve as on-call DRI, analyze metrics, automate production deployments, ensure security/compliance, collaborate with partner teams, respond to incidents, and run postmortems. Support physical infrastructure including GPUs and InfiniBand.
Top Skills:
AzureGpusInfiniband
Insurance
Lead reliability and observability for the financial data platform: define SLOs/SLIs, build metrics pipelines, extend instrumentation across Velocity, Redpanda, MuleSoft, Snowflake, Fabric, and AWS; implement incident management (Datadog -> Incident.io -> ServiceNow), scale automation and remediation, and design AI-assisted SRE agents using Cursor for triage and root-cause analysis.
Top Skills:
Ai Coding AssistantsAWSCursorD365DatadogFabricGrafanaIncident.IoKafkaLlmsMulesoftOpentelemetryPower AppsPrometheusRedpandaServicenowSnowflakeVelocity
Information Technology • Security
Lead technical strategy and architecture for SimSpace's infrastructure, evolving CI/CD and multi-cluster Kubernetes platforms using Jsonnet and Grafana Tanka. Define SLIs/SLOs, build observability with the Grafana stack, embed security and compliance into pipelines, enable self-service developer tooling, command major incidents, and mentor engineering teams to improve reliability and scalability across cloud, on-prem, VMware, and air-gapped deployments.
Top Skills:
ArgocdCi/CdGithub ActionsGitopsGoGrafanaGrafana TankaJsonnetKubernetesKustomizePythonVMware
Software
Join as the company's first SRE to design reliability processes, build observability/CI/CD/IaC tooling, define SLOs, embed with product teams, run incident response, and support compliance for patient data.
Top Skills:
Access ControlsAlerting)Audit TrailsCi/CdGoHipaaInfrastructure As CodeLoggingObservability (MetricsPythonSlosSoc 2TracingTypescript
Healthtech • Telehealth
Lead the creation of SRE practices across six product teams: define SLOs/SLIs, implement observability, reduce outages, automate toil with IaC and tooling, run production readiness reviews, and train teams to improve incident detection and reliability.
Top Skills:
AWSAws EksDatadogGrafanaKubernetesNode.jsPrometheusPythonRdsReactTerraformTypescript
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Automotive • Information Technology • Logistics • Software
Drive reliability across cloud production infrastructure: define SLOs/SLIs, lead incident response, build IaC and CI/CD pipelines, operate Kubernetes/ECS, improve observability, automate toil, mentor SREs, and embed security best practices.
Top Skills:
AWSAzureCassandraCdnCloudFormationDnsEcsElkFinopsGCPGithub ActionsGitopsGoGrafanaIamJavaJenkinsKubernetesLinuxMySQLNewrelicOraclePagerdutyPostgresPrometheusPythonRedisSplunkTcp/IpTerraform
Other • Social Impact
The Senior Site Reliability Engineer is responsible for maintaining Wikimedia's infrastructure, improving reliability, automating processes, and collaborating with teams. The role involves troubleshooting, managing deployments, and leading incident responses while working remotely.
Top Skills:
AnsibleBashCassandraDebianGoGrafanaHhvmKubernetesMariadbMemcachedPHPPrometheusPuppetPythonRedisRubyShell
Information Technology • Security • Cybersecurity
Design, build, and scale Kubernetes-based, multi-tenant infrastructure and CI/CD systems. Own AI tooling infrastructure (MCP servers) and secure AI access patterns. Optimize CI/CD, streaming analytics (Kafka, Flink, ClickHouse), observability, and incident response. Implement IaC (Terraform, Helm, Pulumi), GitOps (Argo CD), progressive delivery, automated testing, and mentor engineering teams.
Top Skills:
Ai AgentsAi/Llm ToolingAksArgo CdBashClickhouseDatadogEksFlinkGithub ActionsGitlab CiGitopsGkeGoGrafanaHelmJenkinsKafkaKubernetesLangfuseLangsmithMcp ServersMlopsOpentelemetryPrometheusPulumiPythonTerraform
Real Estate • Financial Services • PropTech
Support and optimize products migrated to AWS, implement cloud best practices, maintain operational coverage, enhance automation, observability, CI/CD/GitOps, and security. Collaborate with development and platform teams to scale, troubleshoot, and ensure reliable SaaS operations.
Top Skills:
AmisArgocdAWSAws Elastic BeanstalkAws Transfer FamilyAzure DevopsBashCloudwatchCurlDockerEc2EksFluxcdGitGitopsHTTPIstioKubernetesLinkerdLoad BalancerPowershellPythonRdsSQLTerraformWget
Artificial Intelligence • Software • Cybersecurity
Design, build, and maintain scalable, highly available cloud infrastructure for an AI-native cybersecurity platform. Automate deployments and incident response, optimize performance for AI workloads, manage IaC across cloud environments, lead incident management and post-mortems, and collaborate with engineering and security teams to embed reliability.
Top Skills:
AWSAzureDatadogDistributed DatabasesEksElkEvent-Driven SystemsGCPGkeGrafanaKubernetesMicroservicesNetworkingPrometheusPulumiStorageTerraform
Cloud • Security • Software • Cybersecurity
Lead and mentor SRE teams; partner with engineering, operations and product; apply statistical analysis and networking expertise to diagnose performance and reliability issues; define and implement data feeds; influence technical decisions and investments; build tooling to automate analytical workflows and increase platform reliability.
Top Skills:
CCloudDistributed SystemsDnsEdgeHTTPJavaPerlPythonRSQLTcpTls
Legal Tech • Software
Lead observability and incident management efforts: define SLIs/SLOs, build monitoring/alerting, dashboards, logging, and tracing. Drive incident response, postmortems, and reliability improvements to reduce MTTD/MTTR. Integrate observability into CI/CD, maintain AWS and Kubernetes infrastructure, automate operations, and mentor engineers on SRE best practices.
Top Skills:
AWSBashCi/CdDatadogDistributed TracingDynatraceGrafanaKubernetesNew RelicOpentelemetryPowershellPrometheusPython
Reposted 16 Days AgoSaved
Financial Services
Own reliability and scalability of on-prem observability platforms (ELK, Grafana); handle production escalations, capacity planning, SLOs, onboarding, automation, IaC (Terraform/Helm/Ansible), upgrades, security hardening, and platform modernization.
Top Skills:
AnsibleApm InstrumentationBashBeatsChefElasticsearchElk StackFluent BitFluentdGrafanaHelmKibanaLinuxLogstashNew RelicOpentelemetryPrometheusPuppetPythonShell Scripting/Linux ShellSolarwindsTerraform
eCommerce
Ensure reliability and availability of Tradeweb's global AWS platform through IaC automation, observability and SLO definition, incident triage and resolution, on-call duties, collaboration with development teams, and security-focused platform improvements.
Top Skills:
ArgocdAWSAws LambdaEksGitsecopsInfrastructure As Code (Iac)Kubernetes (K8S)KustomizeLgtmLinux/UnixPulumiPythonSmsSns
Blockchain
The Blockchain Site Reliability Engineer is responsible for maintaining blockchain nodes' reliability, monitoring, incident response, and building automation tools to enhance operations.
Top Skills:
DockerElkGoGrafanaJavaScriptKubernetesLinuxPrometheusPythonRustShell
Artificial Intelligence • Cloud • Information Technology • Software
As a Staff SRE, you will ensure the reliability and performance of Andromeda's GPU infrastructure, lead incident responses, build observability systems, and mentor engineers, while collaborating closely with engineering and customers.
Top Skills:
AnsibleCudaGoHelmKubernetesLinuxNcclNvidiaPythonRustSlurmTerraform
Information Technology • Security
The Staff Site Reliability Engineer will lead the architecture and security of the SimSpace cyber range platform, focusing on reliability, automation, and observability across diverse deployment environments while mentoring engineers and driving infrastructure initiatives.
Top Skills:
ArgocdGithub ActionsGoGrafana TankaJsonnetKubernetesPython
Big Data
You will manage AWS infrastructure, automate deployments, debug application issues, and improve the operational health of Metabase Cloud.
Top Skills:
AWSDatadogGoGrafanaKubernetesPrometheusPythonTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Austin, TX Companies Hiring SRE Engineers
See AllPopular Austin, TX Engineering Job Searches
Engineer Jobs in Austin, TX
.NET Developer Jobs in Austin, TX
Android Developer Jobs in Austin, TX
Application Engineer Jobs in Austin, TX
Automation Engineer Jobs in Austin, TX
Backend Engineer Jobs in Austin, TX
C# Jobs in Austin, TX
C++ Jobs in Austin, TX
Cloud Engineer Jobs in Austin, TX
Controls Engineer Jobs in Austin, TX
CTO Jobs in Austin, TX
Design Engineer Jobs in Austin, TX
DevOps Engineer Jobs in Austin, TX
DevOps Jobs in Austin, TX
Director of Engineering Jobs in Austin, TX
Electrical Engineer Jobs in Austin, TX
Embedded Software Engineer Jobs in Austin, TX
Engineering Manager Jobs in Austin, TX
Field Engineer Jobs in Austin, TX
Front End Developer Jobs in Austin, TX
Full Stack Developer Jobs in Austin, TX
Golang Jobs in Austin, TX
Hardware Engineer Jobs in Austin, TX
Infrastructure Engineer Jobs in Austin, TX
iOS Developer Jobs in Austin, TX
Java Developer Jobs in Austin, TX
Javascript Jobs in Austin, TX
Linux Jobs in Austin, TX
Manufacturing Engineer Jobs in Austin, TX
Mechanical Design Engineer Jobs in Austin, TX
Mechanical Engineer Jobs in Austin, TX
Network Engineer Jobs in Austin, TX
PHP Developer Jobs in Austin, TX
Platform Engineer Jobs in Austin, TX
Principal Software Engineer Jobs in Austin, TX
Process Engineer Jobs in Austin, TX
Product Engineer Jobs in Austin, TX
Project Engineer Jobs in Austin, TX
Python Jobs in Austin, TX
QA Engineer Jobs in Austin, TX
QA Jobs in Austin, TX
Robotics Engineer Jobs in Austin, TX
Ruby Jobs in Austin, TX
Salesforce Developer Jobs in Austin, TX
Scala Jobs in Austin, TX
Security Engineer Jobs in Austin, TX
Software Engineer Jobs in Austin, TX
Software Engineering Manager Jobs in Austin, TX
Software Test Engineer Jobs in Austin, TX
Solutions Architect Jobs in Austin, TX
Solutions Engineer Jobs in Austin, TX
SRE Engineer Jobs in Austin, TX
Staff Software Engineer Jobs in Austin, TX
Systems Engineer Jobs in Austin, TX
Web Developer Jobs in Austin, TX
All Filters
Total selected ()
No Results
No Results































