Some of the world’s most innovative global software and technology companies struggle to find engineering partners capable of stepping into complex environments and immediately driving meaningful outcomes. These teams need more than additional hands—they need senior engineers who can quickly understand an environment, identify the path forward, and execute without constant direction.
Enter EverOps – the premier Embedded Service Provider. We partner directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges.
The ChallengeEverOps is looking for a Lead Kubernetes Platform Engineer with deep Amazon EKS experience at very large scale to lead a modernization discovery, and the engineering program that follows it, for a high-scale consumer mobile platform.
You’ll be stepping into a multi-cluster production EKS estate where the largest clusters run tens of thousands of pods at peak and compute demand roughly doubles between overnight lows and daytime highs. Upgrading the estate is slow and largely manual, so the team is perpetually behind the Kubernetes release cycle. Compute is one of the largest lines in the business, and the platform backs real-time, safety-critical features where failing to scale at peak is not an option.
This role requires someone who has run Kubernetes at this scale before, can form a clear point of view on an unfamiliar estate quickly, and can turn that into a sequenced, costed plan that engineering leadership can commit to, then lead the team that delivers it.
The MissionAs a Lead Kubernetes Platform Engineer, you will join our U.S.-Based Virtual Operating Center and lead an embedded TechPod as its Pod Leader: a player-coach who sets technical direction, serves as the primary point of contact for the customer’s infrastructure leadership, and stays hands-on in the work.
Your immediate priority is leading a two-month EKS modernization discovery. You’ll baseline the production estate, map Kubernetes version and support status cluster by cluster, analyze blast radius and failure domains, and deliver the upgrade automation design, compute and capacity economics model, and prioritized roadmap needed to commit the program.
As discovery closes, your focus shifts to leading execution of the modernization program across three phases: reducing blast radius through a multi-cluster target architecture, automating the upgrade cycle so the platform stays current without heroics, and right-sizing compute, including a measured, capacity-aware move to Graviton (ARM64).
Success is measured in the customer’s own terms: a lower cost to serve per active user, and a real reduction in the share of platform engineering time consumed by maintenance and keep-the-lights-on work. You will be expected to tie every recommendation back to those outcomes and explain the tradeoffs clearly to engineers and executives alike.
What You’ll DoEstate Discovery: Build a complete baseline of a multi-cluster production EKS estate, covering cluster inventory, topology, workload placement, ownership, per-cluster cost, and Kubernetes version and support status.
Blast Radius & Architecture: Analyze failure domains, isolation boundaries, and dependency concentration in very large clusters, and design a multi-cluster target architecture with clear workload placement and tenancy models.
Upgrade Automation: Design and implement an upgrade approach (in-place vs. blue/green clusters, Karpenter drift-based node rotation, add-on and API deprecation management) that turns a months-long manual cycle into a repeatable, largely automated process.
Compute & Capacity Strategy: Lead instance sizing and workload-fit analysis across instance families, generations, and node sizes, accounting for DaemonSet overhead, bin-packing, network limits, and headroom for rapid scale-up.
Graviton Migration: Plan and drive a phased ARM64 migration across Java, Go, PHP, and Python workloads, including multi-architecture builds, native dependency remediation, and per-service rightsizing against latency SLOs.
Karpenter Engineering: Configure NodePools, weights, and node overlays so scheduling favors price-performance rather than hourly price, and design safe Spot patterns that protect interruption-sensitive and stream-processing workloads.
Cost Engineering: Model compute, support, and commitment economics (Savings Plans, On-Demand, Spot suitability by workload tier, extended support exposure) and quantify savings ranges by lever.
Platform Posture: Assess and improve ingress (including migration off Ingress NGINX toward Gateway API), service mesh, CNI, GitOps, and infrastructure-as-code posture across the estate.
Operability & Toil Reduction: Quantify maintenance and toil burden using the customer’s own engineering data, set reduction targets, and build the automation that gives time back to platform teams.
AI-Native Operations: Define the scope and success criteria for an AI-assisted DevOps agent workstream, and assess platform readiness for AI-accelerated development demand.
Technical Leadership: Lead the TechPod as a player-coach, run the working cadence with the customer’s infrastructure owner, and partner with AWS specialists on capacity planning and architecture decisions.
Documentation & Readouts: Produce the estate baseline, architecture designs, upgrade plans, roadmap, and executive readout, and present findings and recommendations to engineering leadership.
Experience: 8+ years in DevOps, SRE, Platform, or Infrastructure Engineering, including 4+ years operating production Kubernetes and prior experience in a technical lead, staff, or principal-level role.
EKS at Scale: Deep production experience with Amazon EKS at large scale (multiple production clusters, thousands of nodes, or tens of thousands of pods), with a working understanding of where control plane, scheduling, networking, and API limits start to bite.
Cluster Upgrades: Hands-on ownership of EKS version upgrades across multiple production clusters, including API deprecations, add-on compatibility, node rotation, and rollback planning, plus a working knowledge of the EKS standard and extended support lifecycle.
Karpenter: Advanced production experience with Karpenter (v1+), including NodePools, disruption and consolidation, weighting, and instance-type flexibility.
Compute Fundamentals: Strong understanding of EC2 instance families and generations, CPU architecture differences, network performance limits, and how workload bottlenecks shift when the underlying compute changes.
ARM64 / Graviton: Experience migrating production workloads to Graviton or other ARM64 platforms, including multi-arch container builds and performance validation.
Cloud Cost Management: Experience modeling compute costs and commitments (Savings Plans, Reserved Instances, Spot) and building savings cases grounded in real usage data.
Kubernetes Networking: Solid knowledge of the AWS VPC CNI, ingress controllers, Gateway API, service mesh, and load balancing on EKS.
Infrastructure as Code: Advanced proficiency with Terraform, plus experience with Helm and GitOps tooling such as Argo CD or Flux.
Observability: Comfortable using Datadog, Prometheus, Grafana, or comparable tooling to evaluate cluster health, capacity, and workload performance.
Automation: Strong scripting ability using Python, Go, or Bash to automate infrastructure and operational workflows.
Discovery & Assessment: Demonstrated ability to enter an unfamiliar environment, surface what isn’t documented, and produce a defensible current-state picture and roadmap in weeks rather than quarters.
Communication: Ability to translate technical findings into cost, risk, and capacity terms that engineering leadership and executives can act on.
Consumer Scale: Experience with high-traffic B2C platforms with sharp daily peaks, or with real-time, safety-critical services where availability at peak is non-negotiable.
Multi-Cluster Design: Experience designing cell-based or multi-cluster architectures to reduce blast radius, including fleet management and workload placement across clusters.
Price-Performance Benchmarking: Experience running performance-per-dollar comparisons across instance generations and architectures and feeding the results into scheduling policy.
Spot at Scale: Experience operating Spot capacity for large production fleets, including interruption handling for Kafka consumers and other stateful or stream-processing workloads.
JVM & Runtime Tuning: Experience tuning JVM heap, garbage collection, and Kubernetes requests/limits when moving workloads to new instance generations or architectures.
Container-Optimized OS: Experience running Bottlerocket or similar, including node boot-time optimization.
AI-Assisted Operations: Experience building AI agents or LLM-driven tooling for infrastructure operations, upgrades, or toil reduction.
Consulting: Experience leading assessment or discovery engagements that end in a committed, funded program of work.
Certifications: CKA, CKS, AWS Certified Solutions Architect – Professional, AWS Certified DevOps Engineer – Professional, or similar advanced certifications.
100% Remote Workplace: We’ve been remote since Day 1!
Unlimited Paid Time Off.
Equity: Become a true owner of the company.
401K with company contribution and sponsored healthcare.
Professional Growth: Access to training and certification programs to accelerate your career.
Similar Jobs
What you need to know about the Austin Tech Scene
Key Facts About Austin Tech
- Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
- Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
- Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
- Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center



