Ziff Davis Logo

Ziff Davis

Site Reliability Engineer

Posted 2 Days Ago
Remote
Hiring Remotely in United States
90K-100K Annually
Mid level
Remote
Hiring Remotely in United States
90K-100K Annually
Mid level
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
The summary above was generated by AI

Site Reliability Engineer 

The Opportunity

 

We are looking for a highly capable engineer to join our Platform and Site Reliability engineering team. You will be responsible for building, maintaining and operating the infrastructure platform on which all Ookla services are built. In this role, you will build, maintain, and support a massive-scale dynamic infrastructure that is relied on by  hundreds of millions of users around the world. You will obsess over systems performance, scalability, reliability, observability, and security. Most importantly, you will help deliver critical application functionality and help make the internet experience better for our users to help us achieve our goal of better connectivity for all.

 

We are committed to providing you a flexible work environment where individuality, fun, and talent are all valued equally. If you consider yourself innovative, adept at collaboration, and you care deeply about the work you do, we want to talk!

 

Key Responsibilities

  • Maintaining a distributed, global ecosystem of thousands of cloud instances, containerized workflows, serverless applications, Linux servers, and associated infrastructure supporting billions of requests daily.

  • Maintaining transactional database infrastructure using MySQL, PostgreSQL, and managed services such as RDS/Aurora.

  • Supporting the use of NoSQL data storage engines such as DynamoDb and MongoDB.

  • Building and supporting data stream processing with Kinesis or Kafka.

  • Supporting data engineering and big data toolchains such as Spark.

  •  Supporting production systems in a 24x7x365 environment, including on-call responsibilities.

  • Providing architectural and operational support to software engineers in a wide variety of focus areas.

  • Support software and data engineering teams and guiding operational best practices.

  • Implementation and oversight of security programs including vulnerability remediation, patch management, IDS/IPS, penetration testing, and interfacing with our corporate InfoSec team.

  • Supporting the development to production code deploy pipeline for a range of production applications.

  • Providing the tooling and guidance for the software and data engineering team to implement our monitoring and observability best practices.

  • Assisting development teams with troubleshooting.

 

Job Qualifications

We are looking for the right person, not the exact list of requirements. If you believe your life experience has prepared you for similar challenges, we’d like to hear from you.

  • 4+ Years Systems/Platform engineering experience

  • Experience building globally-distributed systems

  • Strong understanding of security best practices 

  • Infrastructure as Code: Terraform, Cloudformation

  • Branching and Merge based Source Code Configuration Management: Git, Github

  • Configuration management systems such as Chef or Ansible

  • Container-based architectures including Docker, Kubernetes

  • Proficiency in one or more high level programming languages such as Typescript, Go, Python, PHP, Ruby, Java, etc.

  • Experience with AWS and other Cloud infrastructure platforms

  • Comfort writing SQL queries and analyzing query performance

  • Comfortable learning and working with new technologies in an ever-changing environment

  • Strong verbal and written communication skills 

  • Strong time management skills and a self-driven work ethic

  

About 

Ookla, an Accenture company, is a global leader in connectivity intelligence that brings together the trusted expertise of Speedtest®, Downdetector®, Ekahau®, and RootMetrics® to deliver unmatched network and connectivity insights. By combining multi-source data with industry-leading expertise, we transform network performance metrics into strategic, actionable insights.

 Our solutions empower service providers, enterprises, and governments with the critical data and insights needed to optimize networks, enhance digital experiences, and help close the digital divide. At the same time, we amplify the real-world experiences of individuals and businesses that rely on connectivity to work, learn, and communicate. From measuring and analyzing connectivity to driving industry innovation, Ookla helps the world stay connected.

 About Accenture

Accenture helps the world’s leading enterprises reinvent by building their digital core and unleashing the power of AI to create value at speed for organizations across industries. Our strategy is to be the reinvention partner of choice for our clients and lead in the safe, widespread adoption of AI, and to be the most client focused, AI-enabled, great place to work in the world. We bring together the talent of our approximately 799,000 people with proprietary assets and platforms, deep process and industry expertise, and leading ecosystem relationships to deliver end-to-end solutions and measurable outcomes at scale. Through our Reinvention Services, we offer broad expertise across Cybersecurity, Digital Core, Finance, Industry and Enterprise, Song, Supply Chain and Engineering, and Talent, with advanced capabilities in AI and Data, Industry and Process, and Technology. We serve approximately 9,000 clients and generated approximately $70 billion in FY25 revenue. Visit us at accenture.com.

 Compensation Range 

Ookla provides a range for the base pay. Factors that may be used to determine your actual pay may include your specific job related knowledge, skills, experience, and geographic location. The salary compensation for this role is $90,000 - $100,000. Individual pay within the compensation range for this business unit specific role is determined based on a variety of factors including experience, scope of the role, capabilities to perform the role, education and training, as well as business and company performance.


Similar Jobs

3 Days Ago
In-Office or Remote
135K-231K Annually
Expert/Leader
135K-231K Annually
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Leads AI-assisted site reliability engineering across Azure and AWS. Designs observability, incident response, automation, resiliency testing, disaster recovery, chaos engineering, and recovery-validation capabilities. Establishes OpenTelemetry, SLI, SLO, error-budget, and reliability-scorecard standards; improves alert quality and operational insights; creates human-in-the-loop mitigation workflows; and mentors engineers while driving cross-functional reliability improvements.
Top Skills: AnsibleAWSAzureDatadogGrafanaHelmKubernetesLlmsOpentelemetryPrometheusPulumiRagTerraform
10 Days Ago
Easy Apply
Remote
USA
Easy Apply
241K-270K Annually
Senior level
241K-270K Annually
Senior level
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own Garner’s cloud reliability strategy across AWS and Kubernetes, including SLOs, observability, incident response, infrastructure automation, cost optimization, and security compliance. Lead complex incident resolution, architect Terraform-based infrastructure, establish deployment and monitoring standards, mentor engineers, and use AI tools to automate operational work. Support high-scale AI/ML workloads while setting technical direction for platform reliability and production quality.
Top Skills: AWSClaudeDatadogGitlabGoIstioKubernetesNatsPostgresPythonTerraformTypescript
13 Days Ago
Remote
US
125K-174K Annually
Expert/Leader
125K-174K Annually
Expert/Leader
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead the design and roadmap for global Active Directory and identity infrastructure, implement Identity-as-Code and GitOps automation, own incident escalation and observability, define delegation/tiered administration, integrate applications with Okta and cloud identity, mentor teams, and publish identity architecture and security best practices.
Top Skills: Active Directory Domain Services (Ad Ds)AnsibleAWSAws Directory ServiceAzureAzure Active Directory (Entra Id)Azure SentinelCertificate ServicesChefDhcpDnsGCPGitopsGroup Policy Objects (Gpo)New RelicOktaPowershellPowershell DscPythonTerraform

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account