EverOps Logo

EverOps

Lead Observability Engineer

Posted 3 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Lead large-scale observability assessments and migrations across metrics, logs, and traces. Analyze telemetry volume, cardinality, retention, performance, and costs; design AWS-native OpenTelemetry architectures; build TCO models; and lead platform migration execution. Rebuild pipelines, dashboards, alerts, retention policies, ownership tagging, and instrumentation standards. Serve as a player-coach and primary technical contact for customer engineering leadership while delivering architecture, migration, commercial, and executive recommendations.
The summary above was generated by AI
Lead Observability Engineer (Remote)Overview

Some of the world’s most innovative global software and technology companies struggle to find engineering partners capable of stepping into complex environments and immediately driving meaningful outcomes. These teams need more than additional hands—they need senior engineers who can quickly understand an environment, identify the path forward, and execute without constant direction.

Enter EverOps – the premier Embedded Service Provider. We partner directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges.

The Challenge

EverOps is looking for a Lead Observability Engineer with deep hands-on experience across metrics, logs, and traces at very large scale to lead an observability maturity assessment, and the platform consolidation that follows it, for a high-scale consumer mobile platform.

The current estate spans multiple commercial observability vendors alongside self-managed Prometheus, Thanos, Grafana, Vector, and ELK. It carries tens of millions of active time series, tens of thousands of scrape targets, well over a thousand dashboards, and more than ten thousand alert definitions. Much of the log and trace data is sampled or dropped for cost reasons, ownership tagging is sparse, and the self-managed components carry an operational load that crowds out improvement.

The direction under evaluation is consolidation onto AWS-native observability services built on OpenTelemetry. This role requires someone who knows these systems well enough to price that move honestly, prove or disprove performance parity, and then lead the migration if the answer is go.

The Mission

As a Lead Observability Engineer, you will join our U.S.-Based Virtual Operating Center and lead an embedded TechPod as its Pod Leader: a player-coach who sets technical direction, serves as the primary point of contact for the customer’s engineering leadership, and stays hands-on in the work.

Your immediate priority is leading a two-month Observability Maturity Assessment. You’ll validate telemetry volumes, retention, sampling, and cardinality; build a total cost of ownership model; design an AWS-native target architecture; and deliver a migration plan and commercial recommendation that leadership can make a go/no-go decision on.

As discovery closes, your focus shifts to leading the platform migration: standing up the target collection pipeline and storage tiers, porting log transforms to OpenTelemetry, rebuilding dashboards and alerts, and establishing policy-driven retention and ownership tagging that make observability both less expensive and more useful.

The customer’s observability team is capable but stretched thin. You will be expected to add capacity rather than consume it: pull the data yourself, ask targeted questions, keep recommendations unbiased, and leave the team with fewer systems to run and better telemetry than they have today.

What You’ll Do
  • Estate Assessment: Build a complete picture of the observability estate across vendors, agents, collectors, query surfaces, data volumes, and operating model, grounded in live measurement rather than questionnaires alone.

  • Telemetry Analysis: Validate metric cardinality, active series, scrape target health, log volumes, trace sampling, and retention, and identify where fidelity is being lost today and why.

  • Cost Modeling: Build a total cost of ownership comparison between the current state and two to three costed target states, covering vendor spend, self-managed infrastructure, data transfer and egress, and AWS Pricing Calculator estimates.

  • Target Architecture: Design the AWS-native target across collection (ADOT / OpenTelemetry Collector), pipeline (Amazon Data Firehose), storage tiering (CloudWatch, S3, Amazon Managed Service for Prometheus), query surfaces (Amazon Managed Grafana, CloudWatch, Athena, OpenSearch), and alerting.

  • Retention & Data Classification: Partner with Security, Legal, and Engineering to classify telemetry data and design policy-driven retention, including long-term compliance archives.

  • Capability & Gap Analysis: Map current capabilities to AWS-native equivalents and make clear recommendations where no equivalent exists, such as continuous profiling.

  • Performance Parity: Define and run tests that show whether the target state holds query performance, alert latency, and data fidelity against success criteria agreed with engineering leadership.

  • Migration Planning & Execution: Size and sequence the migration, then lead it, including pipeline cutover, porting log transforms to OpenTelemetry, rebuilding dashboards, and deduplicating and rebuilding alerts.

  • OpenTelemetry Standards: Define the OpenTelemetry conventions (semantic conventions, resource attributes, collector topology, sampling strategy) that application teams will adopt as instrumentation moves over.

  • Ownership & Cost Attribution: Rebuild service ownership tagging so telemetry cost can be attributed back to the teams generating it.

  • Commercial Analysis: Reconcile platform spend, analyze licensing and commit structures, and inform vendor renewal strategy with a clear, unbiased recommendation.

  • Technical Leadership: Lead the TechPod as a player-coach, run the working cadence with the customer’s observability leadership, and keep the engagement light-touch on a busy internal team.

  • Documentation & Readouts: Produce the assessment report, target architecture, cost model, migration plan, and executive summary, and present findings and tradeoffs to engineering leadership.

You Have
  • Experience: 8+ years in SRE, DevOps, Observability, or Platform Engineering, including 4+ years owning production observability platforms and prior experience in a technical lead, staff, or principal-level role.

  • Observability at Scale: Deep experience running metrics, logging, and tracing platforms at large scale (millions of active series, multiple terabytes of logs per day) and making them cheaper and more reliable over time.

  • Prometheus Ecosystem: Advanced production experience with Prometheus and a long-term storage layer such as Thanos, Cortex, or Mimir, including PromQL, recording rules, remote write, and cardinality management.

  • Commercial Platforms: Hands-on experience with Datadog or a comparable commercial platform, including how its pricing is built across hosts, custom metrics, log indexing, and APM.

  • AWS Observability: Production experience with CloudWatch, CloudWatch Logs, Container Insights, X-Ray or Application Signals, Amazon Managed Service for Prometheus, and Amazon Managed Grafana.

  • OpenTelemetry: Production experience with the OpenTelemetry Collector (or ADOT), including pipeline design, processors, tail and head sampling, and multi-backend export.

  • Log Pipelines: Experience designing and operating log pipelines with Vector, Fluent Bit, Logstash, or Firehose, including transforms, routing, and tiered storage in S3.

  • Kubernetes: Strong production experience with EKS or Kubernetes, including DaemonSet agent sizing, kube-state-metrics, and monitoring very large clusters.

  • Infrastructure as Code: Advanced proficiency with Terraform; experience with Terragrunt or Atmos is a plus.

  • Alerting & Incident Response: Experience designing SLO-based alerting, reducing alert sprawl, and integrating with incident management tooling such as PagerDuty.

  • Cost Modeling: Ability to build defensible TCO models from usage data and pricing, and to explain the assumptions behind every number.

  • Automation: Strong scripting ability using Python, Go, or Bash to pull usage data, analyze telemetry, and automate migration work.

  • Discovery & Assessment: Demonstrated ability to enter an unfamiliar environment, measure it directly, and produce a defensible current-state picture and recommendation in weeks rather than quarters.

  • Communication: Ability to explain technical, cost, and compliance tradeoffs to engineers, Security and Legal stakeholders, and executive leadership.

Extra Awesome
  • Vendor Migration: Experience migrating off Datadog, Splunk, New Relic, or similar platforms onto AWS-native or open-source observability stacks.

  • Dashboards & Alerts as Code: Experience automating large-scale dashboard and alert migrations using Grafana provisioning, Grafonnet, Terraform providers, or conversion tooling.

  • Grafana Ecosystem: Experience with Grafana Alloy, Beyla, or other eBPF-based instrumentation.

  • Continuous Profiling: Experience with Pyroscope, Parca, or comparable profiling tools.

  • Data Governance: Experience designing telemetry retention and data handling to meet SOC 2, GDPR, privacy, or similar compliance requirements.

  • Analytics Platforms: Familiarity with Athena, OpenSearch, Databricks, or similar platforms used for log analytics and long-term telemetry queries.

  • Consumer Scale: Experience with high-traffic B2C platforms where telemetry volume tracks tens of millions of users.

  • Consulting: Experience leading assessment or discovery engagements that end in a committed, funded program of work.

  • Certifications: Prometheus Certified Associate, OpenTelemetry Certified Associate, AWS Certified DevOps Engineer – Professional, AWS Certified Solutions Architect – Professional, CKA, or similar.

Benefits
  • 100% Remote Workplace: We’ve been remote since Day 1!

  • Unlimited Paid Time Off.

  • Equity: Become a true owner of the company.

  • 401K with company contribution and sponsored healthcare.

  • Professional Growth: Access to training and certification programs to accelerate your career.

Similar Jobs

One Month Ago
In-Office or Remote
Home, TN, USA
107K-284K Annually
Senior level
107K-284K Annually
Senior level
Fitness • Healthtech • Retail • Pharmaceutical
Lead design, build, and operate a scalable observability platform for logs, metrics, and traces across cloud and datacenters. Drive OpenTelemetry adoption, architect high-throughput telemetry pipelines, implement SLOs and platform health checks, automate operations, participate in on-call rotations, and provide technical leadership and mentorship across engineering teams.
Top Skills: Argo CdCloudFormationDockerGoGrafanaHelmJavaKubernetesKustomizeLokiMimirMySQLOpentelemetryOtlpPostgresTempoTerraform
118K-201K Annually
Expert/Leader
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Leads supplier quality engineering for high-reliability printed wiring boards. Conducts supplier audits, capability assessments, source inspections, First Article Inspections, and process evaluations. Analyzes defects, drives root-cause corrective actions, reviews PPAP and FAI documentation, and develops supplier improvement plans using quality metrics. Requires extensive PCB/PWB manufacturing and supplier quality experience, knowledge of aerospace and defense standards, and approximately 50% travel to suppliers across the United States.
Top Skills: ApqpAs9100As9102Asme Y15.1Black BeltControl PlansCorrective And Preventive Action (Capa)First Article Inspection (Fai)Green BeltIpc-6012Ipc-6013Ipc-6018Ipc-A-600Ipc-A-610Ipc-Tm-650Lean ManufacturingLean Six SigmaMil-Prf-31032Mil-Prf-38534Mil-Prf-55110Mil-Std-883Pcba/CcaPfmeaPpapPrinted Circuit Boards (Pcb)Printed Wiring Boards (Pwb)Product Production Line ValidationRoot Cause AnalysisSource InspectionSupplier Scorecards
7 Minutes Ago
Remote or Hybrid
136K-231K Annually
Senior level
136K-231K Annually
Senior level
Aerospace • Hardware • Information Technology • Security • Software • Cybersecurity • Defense
Leads optimization and enhancement of ServiceNow HRSD and Employee Relations capabilities. Configures Virtual Agent, workflows, scripts, portals, integrations, and Employee Center components, including PeopleSoft HCM 9.2 integrations. Gathers stakeholder requirements, oversees testing and deployments, supports production issues, and develops the HRSD roadmap. Partners with enterprise ServiceNow and IT teams on governance, security, architecture, automation, and platform standards.
Top Skills: Application RepositoryAutomated Test FrameworkAzure DevopsBusiness RulesClient ScriptsGitGlide ScriptingJavaScriptNluPeoplesoft Hcm 9.2Performance AnalyticsPowershellRest ApisServicenow Employee CenterServicenow Employee RelationsServicenow Flow DesignerServicenow Hr Case ManagementServicenow HrsdServicenow Integration HubServicenow Knowledge ManagementServicenow Lifecycle EventsServicenow Service PortalServicenow Ui BuilderServicenow Update SetsServicenow Virtual AgentSoap ApisUi Policies

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account