Akamai Technologies Logo

Akamai Technologies

Senior Site Reliability Engineer

Reposted 9 Hours Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Poland
Senior level
In-Office or Remote
Hiring Remotely in Poland
Senior level
The Senior SRE will be responsible for enhancing reliability, automation, and performance of Akamai's AI platforms and collaborating with engineering teams.
The summary above was generated by AI

Do you enjoy solving complex reliability challenges for cutting-edge technology?

Do you have a passion for automation and building systems that scale?

Join the Akamai Inference Cloud Team

The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications with unmatched performance, compliance, and economics.

Partner with the best

As a Senior SRE, responsibilities include owning reliability workstreams for Akamai's serverless inference platform, building automation and tooling, and contributing to architecture and operational decisions. Opportunities exist to take ownership of critical reliability problems end-to-end, partner with product engineering teams, and develop expertise in GPU infrastructure, Kubernetes at scale, and AI inference workloads.

As a Site Reliability Engineer, you will be responsible for:

  • Building and maintaining observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and driving improvements when targets are missed
  • Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response
  • Integrating AI workloads into Akamai's existing incident management processes, building runbooks, participating in on-call rotations, and conducting blameless post-mortems
  • Building and maintaining CI/CD integrations, deployment safety checks, and rollback automation
  • Collaborating with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases
  • Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure

Do what you love

To be successful in this role you will:

  • Demonstrate expertise in SRE, infrastructure, or platform engineering, managing large-scale distributed systems with extensive operational experience.
  • Demonstrate expertise in Kubernetes and large-scale containerization systems.
  • Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.
  • Demonstrate proficiency in Python or Go for automation, CI/CD pipelines, deployment safety, and infrastructure-as-code like Terraform.
  • Interest in or experience with AI/ML infrastructure, model serving, or GPU workloads
  • Resolve issues independently while maintaining accountability throughout the process.
  • Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.

Work in a way that works for you

FlexBase, Akamai's Global Flexible Working Program, is based on the principles that are helping us create the best workplace in the world. When our colleagues said that flexible working was important to them, we listened. We also know flexible working is important to many of the incredible people considering joining Akamai. FlexBase, gives 95% of employees the choice to work from their home, their office, or both (in the country advertised). This permanent workplace flexibility program is consistent and fair globally, to help us find incredible talent, virtually anywhere. We are happy to discuss working options for this role and encourage you to speak with your recruiter in more detail when you apply.
Learn what makes Akamai a great place to work

Connect with us on social and see what life at Akamai is like!

We power and protect life online, by solving the toughest challenges, together.

At Akamai, we're curious, innovative, collaborative and tenacious. We celebrate diversity of thought and we hold an unwavering belief that we can make a meaningful difference. Our teams use their global perspectives to put customers at the forefront of everything they do, so if you are people-centric, you'll thrive here.

Working for you

At Akamai, we will provide you with opportunities to grow, flourish, and achieve great things. Our benefit options are designed to meet your individual needs for today and in the future. We provide benefits surrounding all aspects of your life:

  • Your health
  • Your finances
  • Your family
  • Your time at work
  • Your time pursuing other endeavors

Our benefit plan options are designed to meet your individual needs and budget, both today and in the future.

About us

Akamai powers and protects life online. Leading companies worldwide choose Akamai to build, deliver, and secure their digital experiences helping billions of people live, work, and play every day. With the world's most distributed compute platform from cloud to edge we make it easy for customers to develop and run applications, while we keep experiences closer to users and threats farther away.

Join us

Are you seeking an opportunity to make a real difference in a company with a global reach and exciting services and clients? Come join us and grow with a team of people who will energize and inspire you!
#LI-Remote

Similar Jobs

17 Days Ago
Easy Apply
Remote
Easy Apply
308K-428K Annually
Senior level
308K-428K Annually
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Lead SRE efforts to improve reliability, observability, and incident lifecycle across engineering teams. Own quarterly goals, mentor engineers, define SLOs, drive incident and change management, build tooling and metrics, and collaborate with product, infra, and analytics to ensure highly available distributed systems and operational readiness.
Top Skills: AWSBashKotlinKubernetesMySQLPython
7 Days Ago
In-Office or Remote
Senior level
Senior level
Software
Own production reliability for an AI Experience Framework: operate Kubernetes deployments, monitor and troubleshoot distributed Node.js and Java/JVM services, run incident response/on-call duties, implement CI/CD and GitOps (Helm/ArgoCD/Flux), build observability with Prometheus/Grafana/Splunk, and collaborate on reliability improvements and root-cause analysis.
Top Skills: ArgocdCi/CdDnsFluxGitopsGlideGrafanaHelmHttp/2JavaJvmJwtKedaKubernetesLinuxLitMtlsNode.jsPrometheusServicenowSplunkTcpV8Web Components
9 Days Ago
In-Office or Remote
Senior level
Senior level
eCommerce • On-Demand • Software • Manufacturing
Lead architecture and operation of cloud infrastructure and Kubernetes (EKS). Drive large-scale automation with Terraform and GitOps, improve reliability and observability (Grafana/Prometheus/Loki/Tempo), participate in on-call incident response, enforce security and cost-optimization practices, and mentor mid-level SREs while partnering with product teams.
Top Skills: ArgocdAtlantisAuroraAWSCiliumEcrEksGitGithub ActionsGrafanaHelmHelm ChartIamImage ScanningJenkinsKubernetesLinuxLokiMimirMongoDBMySQLPostgresPrometheusPythonRdsRedisS3SqsTempoTerraformVpc

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account