Carnival Corporation Logo

Carnival Corporation

Manager, Site Reliability Engineering

Posted One Month Ago
Be an Early Applicant
Hybrid
Fort Lauderdale, FL
Senior level
Hybrid
Fort Lauderdale, FL
Senior level
Leads an SRE and DevOps team supporting high-volume web, mobile backend, and API systems. Establishes SLOs, SLIs, error budgets, observability, and automated deployment practices. Oversees 24/7 incident response, post-mortems, performance improvements, edge infrastructure, bot mitigation, DDoS defense, WAF rules, and CDN caching. Manages operational readiness, cross-functional reliability initiatives, and infrastructure automation while ensuring availability, security, resilience, and performance.
The summary above was generated by AI

One of the best-known names in cruising, Princess is the world’s leading international premium cruise line and tour company, carrying millions of guests each year to hundreds of destinations around the globe.  We give our guests the Medallion Class experience others simply can’t. The Love Boat promises something for everyone. 

 

The Manager, Web and Mobile Site Reliability Engineering (SRE) leads the engineering team responsible for ensuring maximum uptime, high availability, performance, and resilience for enterprise web applications, mobile app backends, and public API endpoints. This role defines reliability standards, oversees 24/7 incident response, manages edge infrastructure and bot mitigation, and drives automated deployment and observability pipelines.

Here’s a summary of what Princess is looking for in a Manager, Web and Mobile Site Reliability Engineering. Is this you? 

Responsibilities: 

  • Team Leadership & SRE Operations: Lead and develop a high-performing team of SRE and DevOps engineers supporting 24/7 high-volume web and mobile systems. Manage on-call rotations, incident command protocols, and operational readiness.

  • Reliability & Observability Governance: Establish Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets. Architect end-to-end monitoring, tracing, and alerting strategies using tools like Datadog, Dynatrace, or Grafana.

  • Incident Management & Remediation: Lead major incident response efforts, drive blameless post-mortems, and collaborate with engineering teams to prioritize root-cause fixes and architectural resiliency improvements.

  • Traffic, Edge & Security Management: Partner with IT Security (PCL IT Security) and CDN providers (Akamai) to implement bot mitigation strategies, DDoS defense, WAF rules, and edge caching for key APIs and digital endpoints.

  • Administrative:  Perform all other administrative and organizational duties as required (time keeping, training, travel, collaboration and correspondence, etc.)

 

Knowledge & Skills:

  • Scope: Direct management of SRE and DevOps engineers. Operational oversight for consumer-facing web platforms, mobile backend APIs, edge routing networks, and cloud deployment pipelines.

  • Problem Solving: Rapidly diagnoses and mitigates complex system outages, performance bottlenecks, traffic anomalies, bot campaigns, and infrastructure failures in high-volume production environments.Resolves highly complex, enterprise-scale operational challenges that impact guest operations, maritime services, revenue-generating systems, regulatory requirements, and technology service availability. Anticipates emerging operational risks, evaluates competing business priorities, establishes governance frameworks, and makes decisions where significant operational, financial, service, and reputational consequences may exist. Develops innovative solutions to improve enterprise resilience, scalability, and operational effectiveness.

  • Impact: Directly ensures continuous operational availability, system security, optimal site performance, and guest trust across web and mobile touchpoints.

  • Leadership: The role requires strong leadership skills. Requires strong incident command leadership, strategic operational decision-making, calm under pressure, and collaborative mentorship. 

  • Knowledge: In-depth understanding of Site Reliability Engineering practices, cloud platforms (AWS/Azure), containerization (Kubernetes, Docker), Akamai/CDN edge routing, bot detection, and CI/CD pipelines (GitLab).

  • Skills: Production incident management, automated infrastructure management (Terraform), performance tuning, distributed tracing, metrics-driven SLI/SLO establishment.

  • Abilities: Ability to lead teams during critical production outages, drive cross-functional engineering accountability for reliability, and automate operational workflows.

 

Essential/Minimum Qualifications: 

  • Bachelor’s degree in Computer Science, Computer Engineering, System Administration, or equivalent experience.

  • 6+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering.

  • 2+ years of leadership or direct engineering management experience.

 

Travel: Less than 25% with shoreside travel likely

Work Conditions:  Work primarily in a climate-controlled environment with minimal safety/health hazard potential.

Physical Demands: Remain in a stationary position at a desk and/or computer for extended periods of time; reasonable accommodations will be offered.


**This position is classified as “hybrid.”  As an in-office role, it requires employees to work from a designated Princess location Mondays through Thursdays.  On Fridays you can work from home.

Princess provides comprehensive and innovative benefits to meet your needs, including:

 

What You Can Expect  

  • Cruise and Travel Privileges for You and Your Family 
  • Health Benefits 
  • 401(k)  
  • Employee Stock Purchase Plan  
  • Training & Professional Development 
  • Tuition & Professional Certification Reimbursement 
  • Rewards & Incentives  

  

Our Culture… Stronger Together 

 

Our highest responsibility and top priority is compliance, environmental protection and the health, safety and well-being of our guests, the people in the communities we touch and serve, and our shipboard and shoreside employees.  Please visit our site to learn more about our Culture Essentials, Corporate Vision Statement and our Core Values at: princess.com/en-us/company-information 

  

Princess is an equal-opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable laws, regulations and ordinances. 

  

Americans with Disabilities Act (ADA) 

Princess will provide reasonable accommodation with the application process, upon your request, as required to comply with applicable laws.  If you have a disability and require assistance in this application process, please contact [email protected]. 

 

#PCL 

#LI-Hybrid

#LI-SF1

Similar Jobs

11 Days Ago
Hybrid
Austin, TX, USA
150K-170K Annually
Expert/Leader
150K-170K Annually
Expert/Leader
Fintech
Leads an SRE Automation Engineering team while remaining hands-on in platform reliability, automation, incident response, and infrastructure architecture. The role establishes SRE practices including SLOs, error budgets, toil reduction, and blameless postmortems; develops Rundeck and Airflow workflows; governs Terraform, ArgoCD, Kubernetes, and AWS environments; and drives high availability, security, cost optimization, and capacity planning. The manager also handles mentoring, performance management, Jira prioritization, stakeholder collaboration, and participation in 24/7 on-call rotations.
Top Skills: Amazon SqsAnsibleApache AirflowArgocdAWSAws CliBashBedrock Agent CoreBoto3DnsFix ApiGitopsGoGrafanaHelmJIRAKafkaKubernetesLinuxLlmsMqPrometheusPublic McpsPythonRundeckTcp/IpTerraform
One Month Ago
Hybrid
Senior level
Senior level
Manufacturing
Leads an SRE and DevOps team supporting high-volume web applications, mobile backends, APIs, cloud deployments, and edge infrastructure. Establishes SLOs, SLIs, error budgets, monitoring, tracing, and alerting strategies. Oversees 24/7 incident response, post-mortems, bot mitigation, DDoS defense, WAF rules, performance tuning, infrastructure automation, and reliability improvements. Requires cross-functional leadership, operational governance, and hands-on expertise with cloud platforms, containers, CI/CD, observability, and infrastructure as code.
Top Skills: AkamaiAWSAzureCdnCi/CdDatadogDdos DefenseDistributed TracingDockerDynatraceGitlabGrafanaKubernetesTerraformWaf
Senior level
Transportation • Travel • Hospitality
Leads an SRE and DevOps team supporting high-volume web applications, mobile backends, APIs, edge infrastructure, and cloud deployment pipelines. Establishes SLOs, SLIs, error budgets, observability, and reliability standards; manages 24/7 incident response, post-mortems, bot mitigation, DDoS defense, WAF rules, performance tuning, and infrastructure automation. Requires cross-functional leadership, operational governance, and resilience improvements for enterprise production systems.
Top Skills: AkamaiAWSAzureCdnCi/CdDatadogDdos DefenseDockerDynatraceGitlabGrafanaKubernetesTerraformWaf

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account