HavocAI Logo

HavocAI

Head of Cloud Operations

Posted Yesterday
Remote
Hiring Remotely in USA
200K-220K Annually
Senior level
Remote
Hiring Remotely in USA
200K-220K Annually
Senior level
Leads SRE and DevOps teams while owning cloud reliability, change and release management, incident response, configuration baselines, compliance evidence, and operational processes. Establishes scalable practices for CI/CD, infrastructure automation, observability, on-call health, deployment traceability, rollback planning, incident command, post-incident remediation, and POA&M execution. Partners with engineering and security stakeholders to maintain dependable, auditable, and compliant production systems.
The summary above was generated by AI
About Us:

Havoc is a leader in all-domain collaborative autonomy. Its software-defined hardware approach powers military and commercial-grade autonomous systems across sea, air, and land to sense, decide, and act together in complex and contested environments. Havoc connects assets, enabling them to share information, adapt in real time, and continue operating even when communications are disrupted or denied. Havoc optimizes mission performance and minimizes human risk.
Havoc was founded in 2024 and headquartered in Providence, Rhode Island. Learn more at Havoc: All-Domain Collaborative Autonomy .

About the Role

HavocAI is seeking a Head of Cloud Operations to own the change, release, incident, and reliability practices that keep our systems dependable, auditable, and compliant as we scale.

Reporting to the Director of Cloud and partnering closely with the ISSO and engineering teams, you will define how production changes are approved and deployed, how releases are coordinated, how incidents are managed, and how we maintain a trustworthy record of what is running across our environments.

You will also lead our SRE and DevOps teams, setting direction across reliability, infrastructure automation, CI/CD, observability, and safe delivery. This role requires someone who can build disciplined processes without creating unnecessary bureaucracy—using automation and engineering practices wherever possible to make the right way of working the easiest way of working.

The ideal candidate combines strong operational leadership with enough technical depth to challenge assumptions, make decisions under pressure, and translate security and compliance requirements into practical engineering processes.

What You’ll DoSRE & DevOps Leadership
  • Lead and manage the SRE and DevOps teams, setting technical and operational direction across reliability, automation, and safe delivery.

  • Hire, coach, develop, and manage performance for engineers across both functions.

  • Own reliability and delivery practices including SLIs, SLOs, error budgets, on-call health, CI/CD, and infrastructure automation.

  • Establish clear ownership and operating expectations across cloud reliability and delivery.

  • Partner with engineering leaders to identify systemic reliability risks and prioritize improvements.

  • Build an engineering culture that balances speed, reliability, security, and operational discipline.

Change & Release Management
  • Define and own change classes—including standard, normal, and emergency changes—with clear approval paths and requirements.

  • Establish and operate an appropriate change approval process, including impact assessments, rollback plans, and approval records.

  • Integrate change management with GitOps workflows, using merged, signed, peer-reviewed pull requests as the foundation of the change record.

  • Own maintenance windows, freeze periods, and emergency-change processes, including retroactive approvals where appropriate.

  • Ensure production changes are traceable to an approved request, approver, and rollback decision.

  • Own the release calendar, versioning strategy, and promotion across environments and tenants.

  • Establish pre-deployment verification requirements covering CI status, security scans, migrations, feature flags, and predefined rollback triggers.

  • Coordinate releases across Cloud Platform, Backend, Autonomy, and Frontend teams to prevent conflicts and manage dependencies.

  • Maintain complete, audit-ready deployment and release records.

Incident & Problem Management
  • Own HavocAI’s incident management framework, including incident declaration, severity levels, escalation paths, and incident command.

  • Establish clear authority and expectations for declaring and managing incidents.

  • Run incident command during significant events, coordinating roles, communications, escalation, and stakeholder or customer notifications.

  • Own on-call health, alert quality, escalation practices, and operational readiness.

  • Lead blameless post-incident reviews and ensure remediation actions are assigned, tracked, and completed.

  • Manage government-sponsor notification obligations and timelines for incidents affecting authorized systems, with company-wide scope beyond the IATT boundary.

  • Establish clear distinctions between incidents and problems and drive analysis of recurring issues.

  • Translate recurring operational issues into technical debt, reliability, and remediation priorities.

Configuration, Baselines & Compliance
  • Maintain the authoritative record of deployed systems, including versions, digests, and dependencies.

  • Keep deployment records synchronized with the ISSO’s system inventory.

  • Establish and maintain system baselines and processes for detecting configuration drift.

  • Produce audit-ready evidence for the ISSO, including change records, deployment logs, incident reports, and post-incident remediation actions.

  • Own execution tracking against POA&M commitments, partnering with the ISSO to ensure remediation dates and engineering commitments are met.

  • Translate security and compliance control language into practical engineering processes and clearly communicate engineering implementation back to security stakeholders.

  • Build automation wherever possible to reduce manual compliance work and improve the reliability of operational evidence.

What We’re Looking For
  • 8+ years of relevant experience across change management, release management, incident management, technical program management, service management, SRE, DevOps, platform engineering, or related disciplines.

  • Demonstrated experience leading or managing SRE, DevOps, or Platform Engineering teams, including hiring, coaching, and performance management.

  • Strong understanding of modern cloud operations, software delivery, infrastructure automation, and production reliability.

  • Experience establishing and operating change, release, and incident management processes in complex technical environments.

  • Proven ability to coordinate complex initiatives across engineering teams and stakeholders you do not directly manage.

  • Technical fluency sufficient to evaluate and challenge engineering impact assessments, deployment strategies, rollback plans, and root-cause analyses.

  • Ability to remain calm, decisive, and directive during active incidents and make sound decisions under pressure.

  • Strong written communication skills, particularly for incident communications, post-incident reports, operational documentation, and executive updates.

  • Ability to create scalable processes that provide appropriate control without unnecessarily slowing engineering teams.

  • Strong ownership, judgment, and comfort operating in a fast-moving and ambiguous environment.

  • Must be a U.S. Citizen and able to obtain and maintain a U.S. Government security clearance.

Nice to Have
  • Prior incident command experience in a regulated, defense, government, or safety-relevant environment.

  • Experience operating cloud systems subject to U.S. Government authorization or compliance requirements.

  • Familiarity with POA&Ms, security authorization processes, configuration baselines, and audit evidence management.

  • Experience implementing GitOps-based change and release processes.

  • Knowledge of ITIL practices or equivalent hands-on experience developing effective service management processes.

  • Experience with tools such as Jira, PagerDuty, status pages, runbook platforms, and incident management systems.

  • Experience with Kubernetes, infrastructure as code, CI/CD platforms, observability systems, and modern cloud infrastructure.

What Success Looks Like

Within your first 12 months, you will have:

  • Established an enforced, practical change and release management process that engineering teams consistently follow.

  • Ensured every production change is traceable to the appropriate approval, deployment record, and rollback decision.

  • Created a consistent incident management framework with clear severity levels, ownership, command structures, escalation paths, and communication standards.

  • Established effective post-incident practices with remediation actions tracked through completion.

  • Improved the health and effectiveness of SRE, DevOps, on-call, and reliability practices.

  • Created an automated and trustworthy record of what is deployed across authorized environments.

  • Kept deployment records synchronized with the ISSO’s inventory and maintained audit-ready operational evidence.

  • Established effective tracking and execution against POA&M remediation commitments.

  • Built operational processes that strengthen reliability and compliance without introducing unnecessary friction for engineering teams.

Benefits:
  • 100% Employer paid Health, Dental and Vision Insurance for you and your families

  • Life Insurance (Employer Paid)

  • Ability to participate in the companies 401k program (Matching)

  • Unlimited PTO policy with an enforced 2 week minimum

  • Equity Package

  • Work / Home Office Stipend

  • Global Entry

  • 16 Week Paid Parental Leave

  • Monthly Health and Wellness Stipend


Our Values:
  • Innovation: We are driven to break new ground. Every day presents an opportunity to challenge the status quo, think boldly, and deliver advanced solutions that transform the future of defense technology.

  • Integrity: We hold ourselves to the highest ethical standards, ensuring transparency, accountability, and trust in all our actions and partnerships.

  • Mission-Driven: We are focused on achieving impactful outcomes that align with our core mission—protecting lives through innovation.

  • Forward-Leaning: We continuously seek out new opportunities and remain at the forefront of technological advancements. We embrace change and anticipate the challenges of tomorrow with confidence and creativity.

  • Ownership of All Tasks: At HavocAI, no problem is too complex or too trivial. We believe that greatness comes from tackling the hardest challenges, but also in handling the smallest, sometimes thankless, tasks with the same level of commitment and care.

  • Servant Leadership: We lead by serving others, whether it’s supporting our employees, partners, or the broader community. Empowering those around us is key to achieving long-term success and making a lasting impact.

HavocAI is an Equal Opportunity Employer and is committed to creating an inclusive and diverse workplace. We welcome applicants from all backgrounds and do not discriminate based on race, color, religion, gender, sexual orientation, age, national origin, disability, veteran status, or any other legally protected status.

Similar Jobs

9 Minutes Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills: AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
9 Minutes Ago
Remote or Hybrid
United States
111K-180K Annually
Senior level
111K-180K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads the architecture, modernization, resilience, security, and performance optimization of enterprise mainframe environments. Responsibilities include z/OS performance tuning, WLM and RACF administration, business continuity planning, automation, technical governance, incident resolution, stakeholder collaboration, and guidance of cross-functional engineering and operations teams. The role also evaluates cloud, DevOps, AI, and hybrid IT technologies for mainframe transformation.
Top Skills: AnsibleCsmGlobal MirrorIbm Z/OsMetro MirrorOpenshiftPr/SmPythonRacfRed Hat Ansible For Ibm Z CollectionsRmfSmfWlmZlinux
An Hour Ago
Remote or Hybrid
Expert/Leader
Expert/Leader
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads post-sales success for strategic enterprise accounts, partnering with C-level executives to drive digital transformation, adoption, renewals, customer satisfaction, and expansion. The role establishes success metrics, mitigates risks, coordinates cross-functional teams and partners, guides Customer Success teams, and develops scalable processes and business transformation strategies.
Top Skills: Artificial IntelligenceCloud ComputingEnterprise SoftwareSaaSServicenow

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account