Some of the world’s most innovative global software and technology companies struggle to find engineering partners capable of stepping into complex environments and immediately driving meaningful outcomes. These teams need more than additional hands—they need senior engineers who can quickly understand an environment, identify the path forward, and execute without constant direction.
Enter EverOps – the premier Embedded Service Provider. We partner directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges.
The ChallengeEverOps is looking for a Lead Observability Engineer with deep hands-on experience across metrics, logs, and traces at very large scale to lead an observability maturity assessment, and the platform consolidation that follows it, for a high-scale consumer mobile platform.
The current estate spans multiple commercial observability vendors alongside self-managed Prometheus, Thanos, Grafana, Vector, and ELK. It carries tens of millions of active time series, tens of thousands of scrape targets, well over a thousand dashboards, and more than ten thousand alert definitions. Much of the log and trace data is sampled or dropped for cost reasons, ownership tagging is sparse, and the self-managed components carry an operational load that crowds out improvement.
The direction under evaluation is consolidation onto AWS-native observability services built on OpenTelemetry. This role requires someone who knows these systems well enough to price that move honestly, prove or disprove performance parity, and then lead the migration if the answer is go.
The MissionAs a Lead Observability Engineer, you will join our U.S.-Based Virtual Operating Center and lead an embedded TechPod as its Pod Leader: a player-coach who sets technical direction, serves as the primary point of contact for the customer’s engineering leadership, and stays hands-on in the work.
Your immediate priority is leading a two-month Observability Maturity Assessment. You’ll validate telemetry volumes, retention, sampling, and cardinality; build a total cost of ownership model; design an AWS-native target architecture; and deliver a migration plan and commercial recommendation that leadership can make a go/no-go decision on.
As discovery closes, your focus shifts to leading the platform migration: standing up the target collection pipeline and storage tiers, porting log transforms to OpenTelemetry, rebuilding dashboards and alerts, and establishing policy-driven retention and ownership tagging that make observability both less expensive and more useful.
The customer’s observability team is capable but stretched thin. You will be expected to add capacity rather than consume it: pull the data yourself, ask targeted questions, keep recommendations unbiased, and leave the team with fewer systems to run and better telemetry than they have today.
What You’ll DoEstate Assessment: Build a complete picture of the observability estate across vendors, agents, collectors, query surfaces, data volumes, and operating model, grounded in live measurement rather than questionnaires alone.
Telemetry Analysis: Validate metric cardinality, active series, scrape target health, log volumes, trace sampling, and retention, and identify where fidelity is being lost today and why.
Cost Modeling: Build a total cost of ownership comparison between the current state and two to three costed target states, covering vendor spend, self-managed infrastructure, data transfer and egress, and AWS Pricing Calculator estimates.
Target Architecture: Design the AWS-native target across collection (ADOT / OpenTelemetry Collector), pipeline (Amazon Data Firehose), storage tiering (CloudWatch, S3, Amazon Managed Service for Prometheus), query surfaces (Amazon Managed Grafana, CloudWatch, Athena, OpenSearch), and alerting.
Retention & Data Classification: Partner with Security, Legal, and Engineering to classify telemetry data and design policy-driven retention, including long-term compliance archives.
Capability & Gap Analysis: Map current capabilities to AWS-native equivalents and make clear recommendations where no equivalent exists, such as continuous profiling.
Performance Parity: Define and run tests that show whether the target state holds query performance, alert latency, and data fidelity against success criteria agreed with engineering leadership.
Migration Planning & Execution: Size and sequence the migration, then lead it, including pipeline cutover, porting log transforms to OpenTelemetry, rebuilding dashboards, and deduplicating and rebuilding alerts.
OpenTelemetry Standards: Define the OpenTelemetry conventions (semantic conventions, resource attributes, collector topology, sampling strategy) that application teams will adopt as instrumentation moves over.
Ownership & Cost Attribution: Rebuild service ownership tagging so telemetry cost can be attributed back to the teams generating it.
Commercial Analysis: Reconcile platform spend, analyze licensing and commit structures, and inform vendor renewal strategy with a clear, unbiased recommendation.
Technical Leadership: Lead the TechPod as a player-coach, run the working cadence with the customer’s observability leadership, and keep the engagement light-touch on a busy internal team.
Documentation & Readouts: Produce the assessment report, target architecture, cost model, migration plan, and executive summary, and present findings and tradeoffs to engineering leadership.
Experience: 8+ years in SRE, DevOps, Observability, or Platform Engineering, including 4+ years owning production observability platforms and prior experience in a technical lead, staff, or principal-level role.
Observability at Scale: Deep experience running metrics, logging, and tracing platforms at large scale (millions of active series, multiple terabytes of logs per day) and making them cheaper and more reliable over time.
Prometheus Ecosystem: Advanced production experience with Prometheus and a long-term storage layer such as Thanos, Cortex, or Mimir, including PromQL, recording rules, remote write, and cardinality management.
Commercial Platforms: Hands-on experience with Datadog or a comparable commercial platform, including how its pricing is built across hosts, custom metrics, log indexing, and APM.
AWS Observability: Production experience with CloudWatch, CloudWatch Logs, Container Insights, X-Ray or Application Signals, Amazon Managed Service for Prometheus, and Amazon Managed Grafana.
OpenTelemetry: Production experience with the OpenTelemetry Collector (or ADOT), including pipeline design, processors, tail and head sampling, and multi-backend export.
Log Pipelines: Experience designing and operating log pipelines with Vector, Fluent Bit, Logstash, or Firehose, including transforms, routing, and tiered storage in S3.
Kubernetes: Strong production experience with EKS or Kubernetes, including DaemonSet agent sizing, kube-state-metrics, and monitoring very large clusters.
Infrastructure as Code: Advanced proficiency with Terraform; experience with Terragrunt or Atmos is a plus.
Alerting & Incident Response: Experience designing SLO-based alerting, reducing alert sprawl, and integrating with incident management tooling such as PagerDuty.
Cost Modeling: Ability to build defensible TCO models from usage data and pricing, and to explain the assumptions behind every number.
Automation: Strong scripting ability using Python, Go, or Bash to pull usage data, analyze telemetry, and automate migration work.
Discovery & Assessment: Demonstrated ability to enter an unfamiliar environment, measure it directly, and produce a defensible current-state picture and recommendation in weeks rather than quarters.
Communication: Ability to explain technical, cost, and compliance tradeoffs to engineers, Security and Legal stakeholders, and executive leadership.
Vendor Migration: Experience migrating off Datadog, Splunk, New Relic, or similar platforms onto AWS-native or open-source observability stacks.
Dashboards & Alerts as Code: Experience automating large-scale dashboard and alert migrations using Grafana provisioning, Grafonnet, Terraform providers, or conversion tooling.
Grafana Ecosystem: Experience with Grafana Alloy, Beyla, or other eBPF-based instrumentation.
Continuous Profiling: Experience with Pyroscope, Parca, or comparable profiling tools.
Data Governance: Experience designing telemetry retention and data handling to meet SOC 2, GDPR, privacy, or similar compliance requirements.
Analytics Platforms: Familiarity with Athena, OpenSearch, Databricks, or similar platforms used for log analytics and long-term telemetry queries.
Consumer Scale: Experience with high-traffic B2C platforms where telemetry volume tracks tens of millions of users.
Consulting: Experience leading assessment or discovery engagements that end in a committed, funded program of work.
Certifications: Prometheus Certified Associate, OpenTelemetry Certified Associate, AWS Certified DevOps Engineer – Professional, AWS Certified Solutions Architect – Professional, CKA, or similar.
100% Remote Workplace: We’ve been remote since Day 1!
Unlimited Paid Time Off.
Equity: Become a true owner of the company.
401K with company contribution and sponsored healthcare.
Professional Growth: Access to training and certification programs to accelerate your career.
Similar Jobs
What you need to know about the Austin Tech Scene
Key Facts About Austin Tech
- Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
- Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
- Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
- Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center


