At Everbridge, we’re building a resilient, scalable, and secure cloud platform that powers critical services used around the world. We’re looking for a Senior Platform Site Reliability Specialist to own, operate, and evolve our enterprise observability platform.
In this role, you will be responsible for the up-keep, reliability, scalability, and strategic growth of Everbridge’s observability stack, EKS, and supporting services, ensuring our engineering teams have deep visibility into system health, performance, and reliability across a large-scale, cloud-native environment. You will also be working with other cloud technologies within the AWS and GCP areas.
We’re looking for someone who shows up for the team, not just themselves. This role works best for a person who communicates clearly, collaborates easily, and treats interactions with other teams with respect and professionalism. You should be comfortable being involved, offering support, and helping move work forward without ego. We value people who build trust, keep things running smoothly, and make the teams around them better.
What you'll do:
- Head the design, operation, and evolution of Everbridge’s observability stack
- Build and maintain a highly available, scalable observability platform
- Standardize instrumentation, dashboards, alerts, and SLOs
- Support incident response, root cause analysis, and capacity planning Grafana Stack & Telemetry
- Operate and scale Grafana and technology
- Grafana Loki (logs)
- Grafana Mimir (metrics)
- Grafana Tempo (tracing)
- Grafana Alerting Kubernetes
- Maintain reliability and security of EKS clusters running observability
- Manage cluster lifecycle and upgrades Infrastructure as Code & Automation
- Terraform for infrastructure provisioning
- HashiCorp Packer
- Gitlab CI/CD at Scale
What you'll bring:
- 6+ years in SRE / Platform Engineering
- Strong Grafana ecosystem experience
- Kubernetes and Amazon EKS expertise
- Terraform proficiency
Preferred Qualifications:
- OpenTelemetry experience
- Large-scale observability systems
- Cost optimization experience
Similar Jobs
What you need to know about the Austin Tech Scene
Key Facts About Austin Tech
- Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
- Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
- Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
- Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center


