CAPS Payroll Logo

CAPS Payroll

Staff DevOps Engineer

Posted 8 Hours Ago
Remote
Hiring Remotely in United States of America
190K-235K Annually
Senior level
Remote
Hiring Remotely in United States of America
190K-235K Annually
Senior level
Own and evolve CI/CD pipelines, AWS EKS infrastructure, Terraform-based infrastructure, GitOps practices, and developer platform tooling. Improve reliability through observability, incident response, automation, and systemic fixes. Establish DevOps standards, lead architecture reviews, mentor engineers, maintain documentation, and guide platform investments across a multi-team engineering organization.
The summary above was generated by AI

About Us

At Cast & Crew, we’ve empowered creativity and supported the global entertainment industry for decades. Together with our family of brands - Backstage, CAPS, Checks & Balances, Final Draft, Media Services, Sargent-Disc, and The TEAM Companies – we operate as a combined entertainment technology and services provider offering industry standard screenwriting accounting software, digital payroll products, data & reporting, and a host of creative tools.  The industry continues to move faster than ever, and the need for our expertise, our technology, and our people has never been greater.  We are a production’s best ally every step of the way. #OneCastOneCrew

Position Overview

We are looking for a Staff DevOps Engineer to serve as a technical anchor for our platform engineering practice. In this role you will own the design and evolution of our CI/CD pipelines, Kubernetes infrastructure on AWS EKS, and the developer experience tooling that hundreds of engineers depend on daily. Staff-level engineers at this organization are expected to operate with significant autonomy, identify and resolve systemic problems before they become incidents, and raise the technical bar across the teams they partner with.

Core ResponsibilitiesPlatform & Infrastructure
  • Architect and continuously improve CI/CD pipelines in Azure DevOps, including pipeline-as-code standards, templating strategies, and artifact promotion workflows across environments.

  • Own the health and evolution of our AWS EKS clusters — node lifecycle, autoscaling, networking (VPC/CNI), RBAC, and cluster upgrades with minimal service disruption.

  • Design and enforce Infrastructure-as-Code practices using Terraform or equivalent tooling; champion GitOps patterns across engineering teams.

  • Drive platform reliability improvements informed by observability data from New Relic, working closely with SRE to translate dashboards and alerts into actionable platform changes.

Developer Experience
  • Define and maintain golden-path templates for containerized workloads — Dockerfile standards, Helm chart libraries, and local development parity with production.

  • Partner with engineering teams to accelerate onboarding of new services onto the platform and reduce toil through automation.

Incident & Operational Excellence
  • Act as an escalation point for complex infrastructure incidents coordinated through PagerDuty; participate in on-call rotation and lead post-incident reviews for platform-layer failures.

  • Identify recurring failure modes and drive systemic fixes that reduce page volume and MTTR across the platform.

  • Maintain and improve runbooks and platform documentation in Confluence, ensuring knowledge is accessible and current.

Technical Leadership
  • Define and socialize DevOps standards — pipeline design, container hygiene, secret management, and deployment safety — across a multi-team engineering organization.

  • Conduct architecture reviews and provide technical guidance on infrastructure-impacting decisions made by product engineering teams.

  • Mentor senior and mid-level engineers; grow internal platform capability through pairing, code review, and structured knowledge sharing.

  • Identify tooling gaps and build the business case for platform investments, working with engineering leadership to prioritize roadmap items.

Key Qualifications
  • 8+ years of DevOps or platform engineering experience, with at least 2 years operating at a Staff or Principal level in an organization of 100+ engineers.

  • Deep, hands-on expertise with Kubernetes — EKS specifically preferred — including troubleshooting workloads, networking, storage, and cluster operations at scale.

  • Strong command of Azure DevOps Pipelines, including YAML pipeline authoring, library management, service connections, and environment promotion gates.

  • Proven track record designing and maintaining CI/CD systems for microservice architectures with multiple independent teams as consumers.

  • Experience operating observability platforms (New Relic, Datadog, or similar) to drive proactive reliability improvements, not just reactive alerting.

  • Proficiency in at least one scripting language (Python, Bash, or Go) and Infrastructure-as-Code tooling (Terraform, Pulumi, or CDK).

  • Familiarity with feature flag patterns and operational considerations around progressive delivery (Unleash or equivalent is a plus).

  • Excellent written communication skills — you default to documentation and can translate complex infrastructure decisions into guidance engineers actually read.

Preferred Qualifications
  • Experience with data engineering or ML infrastructure workloads on Kubernetes (Spark on EKS, Argo Workflows, Airflow).

  • Background contributing to or maintaining internal developer portals (Backstage or similar).

  • Familiarity with FinOps practices and tooling for AWS cost attribution and optimization across shared Kubernetes clusters.

  • Experience in SRE-adjacent roles; comfort with SLO/SLI definition and error budget policy.

Special Work Conditions    
  • Sedentary – Involves sitting most of the time but may involve walking or standing for brief periods of time. Some positions may entail exerting up to 15 lbs. of force occasionally and/or a negligible amount of force to lift, carry, push, or pull.


We take care of our people.
When you join Cast & Crew, you're backed by a benefits package built around what matters most. From comprehensive medical, dental, and vision coverage, to a 401(k) match, generous PTO, paid parental leave, health and wellness programs, tuition reimbursement, and employee discounts. We’re committed to helping you thrive, professionally, and personally.

Benefits are subject to eligibility requirements.


Cast & Crew is an equal opportunity employer committed to hiring a diverse workforce and sustaining an inclusive culture. It is our policy to provide equal employment opportunities to all individuals based on job-related qualifications and ability to perform a job, without regard to age, gender, gender identity, sexual orientation, race, color, religion, creed, national origin, disability, genetic information, veteran status, citizenship or marital status, and to maintain a non-discriminatory environment free from intimidation, harassment or bias based upon these grounds.


CA residents
Your personal information may be collected in connection with certain services provided by Cast & Crew or its affiliated companies.  A summary of your California privacy rights can be found at: https://www.castandcrew.com/privacy-policy/


Compensation is commensurate with various factors including, but not limited to, relevant experience, qualifications, skills, training, licensure, certifications, geographic cost of labor, and other business and organizational needs. Compensation range for candidates in other locations may differ based on the cost of labor in that location. The compensation range for this position is: $190,000.00 - $235,000.00 per year.

Similar Jobs

10 Hours Ago
Remote
United States of America
190K-235K Annually
Senior level
190K-235K Annually
Senior level
Digital Media • Events • News + Entertainment
Own and evolve CI/CD pipelines, AWS EKS Kubernetes infrastructure, Infrastructure-as-Code, GitOps practices, and developer experience tooling. Improve platform reliability through observability, lead incident response and post-incident reviews, establish DevOps standards, and reduce operational toil. The role provides architectural guidance, mentors engineers, supports service onboarding, maintains documentation, and drives platform investments across a large engineering organization.
Top Skills: AirflowArgo WorkflowsAWSAws CdkAws EksAzure DevopsAzure Devops PipelinesBackstageBashCi/CdConfluenceDatadogDockerGitopsGoHelmInfrastructure As CodeKubernetesNew RelicPagerdutyPulumiPythonRbacSparkTerraformUnleashVpc/Cni
13 Days Ago
In-Office or Remote
206K-343K Annually
Expert/Leader
206K-343K Annually
Expert/Leader
Artificial Intelligence • Information Technology • Software
Advance the reliability, scalability, and automation of Horizon Cloud’s global infrastructure. Build CI/CD pipelines, manage Azure and Kubernetes environments, implement observability with Elastic Stack, Prometheus, Grafana, and OpenTelemetry, and automate operations with Python and Terraform. Apply SRE, incident response, infrastructure security, and compliance practices while troubleshooting distributed systems. Provide technical leadership, mentor team members, and drive improvements to resilience, developer productivity, and self-healing capabilities.
Top Skills: ArgocdAzureElastic StackFirewallsGithub ActionsGitopsGrafanaHelmJenkinsKubernetesOpentelemetryPrometheusPythonTerraformWaf
2 Days Ago
Remote
USA
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Healthtech
Designs and operates secure, scalable Azure cloud infrastructure supporting AI initiatives. Builds Infrastructure as Code, CI/CD pipelines, self-service provisioning, containerized environments, observability, monitoring, and reliability practices. Establishes platform standards, evaluates emerging technologies, drives cross-team architecture decisions, mentors engineers, and ensures governance, security, compliance, data quality, and auditability. The role is fully remote with up to 10% travel.
Top Skills: AlertingAWSAzureBashCi/CdCloud InfrastructureDockerGithub ActionsGitlab CiGoIncident ResponseInfrastructure As CodeKubernetesMonitoringMulti-CloudObservabilityPulumiPythonTerraform

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account