Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Austin, TX
Artificial Intelligence • Big Data • Information Technology • Security • Software
Design, build, and maintain scalable infrastructure and tooling to improve availability, reliability, performance, and security. Implement SRE practices (SLI/SLO/SLA), eliminate toil, monitor telemetry, participate in incident response and on-call rotation, and collaborate cross-functionally to support global production systems.
Top Skills:
AnsibleAnycastBgpC/C++CdnDnsDockerElasticsearchGitGitlabGoGrafanaHTTPJenkinsKafkaKubernetesLinuxNoSQLPrometheusPythonRdbmsRedisSaltstackTcpTls/SslUdp
Artificial Intelligence • Logistics • Software • Defense
Operate and harden production logistics decision systems for classified and cloud environments. Own availability, monitoring, incident response, CI/CD and IaC automation, and compliance (RMF/STIG/ATO). Partner with engineers and government stakeholders, document runbooks, and support deployments in air-gapped and multi-cloud environments while traveling frequently to customer sites.
Top Skills:
AnsibleAWSAzureAzure Government (Gcc High)Ci/CdDatadogDockerElkGovcloudGrafanaKubernetesLinuxPrometheusTerraform
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application and network security, reliability, speed, and capacity; automate cloud deployments; monitor and troubleshoot services to meet SLAs; analyze logs and events; and resolve infrastructure issues across large-scale cloud infrastructure.
Top Skills:
AnsibleAWSBashBitbucketDatabricksDatadogDigicertGradleJenkinsNode.jsPythonSplunkTenableTerraformThreatmetrix
Cloud • Security • Software • Cybersecurity
Deploy and operate scalable, highly available cloud systems; improve application and network security, reliability, speed, and capacity; automate cloud deployments; monitor and troubleshoot services to meet SLAs; analyze logs and events; and resolve infrastructure issues across large-scale distributed systems.
Top Skills:
Cloud ComputingDistributed SystemsHTTPJavaLinux/UnixPerlPythonTcp/IpTls/Ssl
Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Lead architecture and implementation of reliability platforms and SRE practices for a production SaaS. Build self-service reliability tooling, drive AIOps automation, advance observability (monitoring, tracing, profiling), lead incident response and postmortems, mentor engineers, and embed production readiness across teams to achieve 99.99% uptime.
Top Skills:
AWSAzureContinuous ProfilingDatadogDnsElkGCPGoGrafanaHttp/SKubernetesLoad BalancingOpentelemetryPrometheusPythonTcp/Ip
Artificial Intelligence • Fintech • Machine Learning • Software • App development • Conversational AI • Generative AI
Own and improve production infrastructure reliability, deployments, Infrastructure-as-Code, Kubernetes environments, automation, CI/CD, monitoring, alerting, and observability. Investigate incidents, optimize system performance, maintain documentation and runbooks, and support DNS, WAF, CDN, and caching infrastructure. The role requires strong Linux administration, Bash scripting, networking, Git, and containerization skills, with independent ownership and collaboration across development and operations teams.
Top Skills:
AkamaiAmqpAnsibleAWSBashCdnCloudflareDnsDockerGCPGitGitlab CiGrafanaHttp/HttpsKubernetesLinuxPodmanPrometheusPythonRabbitMQTerraformVictoriametricsWafZabbix
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills:
AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
Cloud • Software
The Site Reliability Engineer will ensure reliable cloud operations by applying Python for infrastructure automation, managing OpenStack and Kubernetes, and practicing devsecops in a fast-paced environment.
Top Skills:
KubernetesLinuxOpenstackPython
Cloud • Security • Software • Cybersecurity
As a Site Reliability Engineer II, you'll automate tasks, monitor AI workloads, enhance dashboards, support CI/CD processes, and collaborate with engineering teams on complex issues while participating in on-call rotations.
Top Skills:
GoGrafanaKubernetesLinuxPrometheusPythonSaltstackTerraform
Software • Cybersecurity
This role involves managing Kubernetes clusters, cloud infrastructure, and CI/CD pipelines. The engineer will enhance system reliability and efficiency while troubleshooting production issues.
Top Skills:
AlertmanagerAWSAzureBashCi/CdDockerElastic StackElasticsearchGCPGoGrafanaHelmKafkaKubernetesLokiMongoDBOciPrometheusPythonRedisSparkTerraform
Software • Web3
Lead reliability practices across teams: embed early in projects, define SLIs/SLOs, build multi-cloud paved roads with Terraform, run on-call, drive org-wide incident maturity and tooling.
Top Skills:
AWSAzureGCPRuby On RailsTerraformTypescriptWebcontainers
Information Technology • Security • Cybersecurity
Lead design, build, and scale of Kubernetes-based, multi-tenant infrastructure and AI tooling. Own CI/CD, IaC, GitOps, and streaming analytics (Kafka/Flink/ClickHouse). Improve observability, SLOs, automated testing, progressive delivery, incident response, and mentor teams on reliability, security, and automation.
Top Skills:
Ai/Llm ToolingAksArgo CdBashClickhouseDatadogEksFlinkGithub ActionsGitlab CiGitopsGkeGoGrafanaHelmJenkinsKafkaKubernetesMcp ServersOpentelemetryPrometheusPulumiPythonTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills:
AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
Automotive
Leads SRE engineering leaders and engineers while defining enterprise observability, reliability, and platform strategy across GCP, on-premise, manufacturing, distribution, and campus environments. Oversees vendor-agnostic tooling, OpenTelemetry integrations, CI/CD observability, SRE maturity models, and Agentic AI initiatives. Drives adoption of SRE practices, develops technical roadmaps, partners with senior leadership and operational teams, and maintains hands-on architectural and technical credibility.
Top Skills:
Agentic AiAWSAzureCi/CdDatadogDynatraceGCPNew RelicOpentelemetryOtel Genai Semantic ConventionsSource Control PlatformsSplunkTerraform
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills:
AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Artificial Intelligence • Cloud • Information Technology • Software
Design and operate large-scale GPU infrastructure for distributed AI training, ensuring reliability, performance, and efficient customer partnerships.
Top Skills:
AnsibleCudaDeepspeedFsdpGpuHelmInfinibandKubernetesLinuxMegatronNcclNvidia A100Nvidia B200Nvidia H100NvlinkPyTorchRoceTerraform
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills:
AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills:
Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills:
Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
Artificial Intelligence • Big Data • Information Technology • Security • Software
Design, build, and maintain cloud infrastructure and CI/CD for a high-availability telecommunications product. Define SLOs/SLIs, manage incident response and on-call rotations, implement observability, perform performance and capacity planning, run blameless postmortems, and collaborate with security teams to ensure compliance and access control.
Top Skills:
AnsibleAWSDatadogDockerGCPGitlabHelmJavaJenkinsKubernetesNoSQLTerraform
Software
The role involves managing compute infrastructure for decentralized applications, requiring critical thinking, documentation skills, and experience in Kubernetes and blockchain management.
Top Skills:
BlockchainGitopsInfrastructure-As-CodeKubernetesProgramming Languages
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills:
DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
Reposted 17 Days AgoSaved
Other • Social Impact
As a Senior Site Reliability Engineer, you will design, develop, and maintain reliable infrastructure for Wikimedia's API services, ensuring performance and availability while driving reliability engineering practices and improving developer experience.
Top Skills:
AnsibleArgocdAWSAzureGCPGitlabGoKubernetesOpentelemetryPrometheusPythonTerraform
Angel or VC Firm • Blockchain • Fintech • Cryptocurrency
Apply to join Galaxy Ventures' invite-only Talent Network for DevOps, SRE, QA, and Security professionals. Upon acceptance, your profile may be discreetly shared with portfolio companies for relevant roles, and you'll receive invitations to exclusive networking events. Participation is confidential and non-binding.
Artificial Intelligence • Information Technology • Software
The Senior SRE will manage multi-cloud infrastructure, ensuring reliability and scalability. Responsibilities include building CI/CD pipelines, defining SLOs, and implementing automation.
Top Skills:
Ai-Assisted DevelopmentAWSAzureClaude CodeDatadogGCPGrafanaKubernetesTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Austin, TX Companies Hiring SRE Engineers
See AllPopular Austin, TX Engineering Job Searches
Engineering Jobs in Austin, TX
.NET Developer Jobs in Austin, TX
Android Developer Jobs in Austin, TX
Application Engineer Jobs in Austin, TX
Automation Engineer Jobs in Austin, TX
Backend Engineer Jobs in Austin, TX
C# Jobs in Austin, TX
C++ Jobs in Austin, TX
Cloud Engineer Jobs in Austin, TX
Controls Engineer Jobs in Austin, TX
CTO Jobs in Austin, TX
Design Engineer Jobs in Austin, TX
DevOps Engineer Jobs in Austin, TX
DevOps Jobs in Austin, TX
Director of Engineering Jobs in Austin, TX
Electrical Engineering Jobs in Austin, TX
Embedded Software Engineer Jobs in Austin, TX
Engineering Manager Jobs in Austin, TX
Field Engineer Jobs in Austin, TX
Front End Developer Jobs in Austin, TX
Full Stack Developer Jobs in Austin, TX
Golang Jobs in Austin, TX
Hardware Engineer Jobs in Austin, TX
Infrastructure Engineer Jobs in Austin, TX
iOS Developer Jobs in Austin, TX
Java Developer Jobs in Austin, TX
Javascript Jobs in Austin, TX
Linux Jobs in Austin, TX
Manufacturing Engineer Jobs in Austin, TX
Mechanical Design Engineer Jobs in Austin, TX
Mechanical Engineering Jobs in Austin, TX
Network Engineer Jobs in Austin, TX
PHP Developer Jobs in Austin, TX
Platform Engineer Jobs in Austin, TX
Principal Software Engineer Jobs in Austin, TX
Process Engineer Jobs in Austin, TX
Product Engineer Jobs in Austin, TX
Project Engineer Jobs in Austin, TX
Python Jobs in Austin, TX
QA Engineer Jobs in Austin, TX
QA Jobs in Austin, TX
Robotics Engineer Jobs in Austin, TX
Ruby Jobs in Austin, TX
Salesforce Developer Jobs in Austin, TX
Scala Jobs in Austin, TX
Security Engineer Jobs in Austin, TX
Software Engineer Jobs in Austin, TX
Software Engineering Manager Jobs in Austin, TX
Software Test Engineer Jobs in Austin, TX
Solutions Architect Jobs in Austin, TX
Solutions Engineer Jobs in Austin, TX
SRE Engineer Jobs in Austin, TX
Staff Software Engineer Jobs in Austin, TX
Systems Engineer Jobs in Austin, TX
Web Developer Jobs in Austin, TX
All Filters
Total selected ()
No Results
No Results




























