GE Vernova Logo

GE Vernova

System Reliability Engineering Lead

Posted 4 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in USA
152K-228K Annually
Expert/Leader
Remote
Hiring Remotely in USA
152K-228K Annually
Expert/Leader
Leads system reliability engineering for a global grid software SaaS portfolio. Owns cloud infrastructure, platform standardization, SLOs, release governance, incident command, disaster recovery, FinOps, and capacity planning. Serves as final production deployment authority and customer-facing reliability lead. Builds a reliability enablement center, mentors a distributed SRE team, and drives automation, progressive delivery, observability, compliance, and operational excellence across critical utility applications.
The summary above was generated by AI
Job Description SummaryAs the Tech Lead for System Reliability Engineering within the GridOS SaaS Products organization, you will be the hands-on technical authority on production stability for our global grid software SaaS portfolio. You will bridge the gap between architectural design and real-world operations, driving a culture of high reliability and engineering excellence across a distributed team spanning three geographies. You are the "Gatekeeper" for production environments — owning the Change Management process, holding final authority to approve or halt deployments based on system health, and accountable for meeting SLA/SLO targets for critical infrastructure applications serving major North American utility customers.
This is a player-coach role. You will architect and build alongside your team while setting technical direction, mentoring engineers, and serving as the primary customer-facing SRE point of contact. You will own FinOps for the SaaS platform, driving cloud cost optimization and capacity planning as the customer base scales.

Job DescriptionDay 0 — Strategic Provisioning and Design

Standardized Cloud Infra Provisioning

Architect and implement standardized, secure cloud infrastructure provisioning. Drive extreme automation to reduce account provisioning timelines and accelerate customer onboarding to the SaaS platform.

The Golden Path

Define and build the standardized "Middle-Mile" software delivery platform (IDP) using Backstage, ArgoCD, and GitHub Actions. Eliminate bespoke deployment methodologies and establish a single, repeatable path to production.

Follow-the-Sun Architecture

Design and operate the global handover protocols and 24/7 operational coverage model across US, India, and Mexico time zones. Ensure seamless support continuity without graveyard shifts.

Reliability Targets

Establish and own enterprise-wide Service Level Objectives (SLOs) and Service Level Indicators (SLIs) aligned with critical user journeys for global utility customers. Define error budgets and enforce them.

Day 1 — Release Governance and Deployment

Final Approval Authority

Serve as the final technical authority for all production releases. Enforce rigorous change control and validate that all security and performance quality gates are met before any deployment proceeds.

Progressive Delivery

Implement and operate advanced deployment strategies including Canary and Blue/Green rollouts. Build and verify automated rollback capabilities. Hands-on with deployment tooling and pipeline configuration.

SRE Center for Enablement (C4E)

Build and mature the C4E to provide coaching, standardized templates, and repeatable reliability patterns that uplift practices across all product teams. Act as the go-to technical resource for reliability engineering across the organization.

Day 2 — Operational Excellence and Optimization

Incident Command

Serve as the Lead Incident Commander for high-severity (Sev1/Sev2) events. Lead the technical direction, communication, and containment efforts. Available for P1 escalations around the clock.

Blameless Culture

Own the post-incident lifecycle. Facilitate blameless Root Cause Analysis (RCA) to ensure systemic fixes replace recurring operational risks. Build a team culture where incidents drive improvement, not blame.

Business Continuity

Architect and validate end-to-end Backup and Disaster Recovery (DR) strategies, including cross-region failover and automated recovery testing. Hands-on with DR runbook development and execution.

FinOps and Capacity Planning

Own financial operations for the SaaS platform. Drive cloud cost optimization through reserved instances, right-sizing, and waste elimination. Perform long-term capacity planning based on customer growth trajectory and application scaling requirements.

Customer Engagement and Team Leadership

Customer-Facing Accountability

Serve as the primary SRE point of contact for North American utility customers. Own customer satisfaction and NPS for SaaS reliability. Participate in customer-facing reviews, incident communications, and service health reporting. Must meet customer-mandated background check requirements for access to critical infrastructure data and environments.

Player-Coach Team Leadership

Lead a distributed team of 8 SRE engineers across Hyderabad Technical Center and Querétaro, scaling with SaaS application and customer growth. Set technical direction, assign tasks, own team deliverables, and drive day-to-day execution. Mentor engineers on SRE practices, cloud architecture, and operational discipline. Provide performance feedback to the people leader of record. Foster a culture of high performance and continuous learning.



Required Qualifications

Technical Qualifications

  • Cloud Ecosystem: Deep expertise in AWS core services (EC2, EKS, RDS, S3, IAM) and management tools (CloudTrail, CloudWatch)
  • Orchestration: Advanced mastery of Kubernetes internals and EKS cluster operations across multi-region architectures
  • Continuous Delivery: Expert knowledge of ArgoCD, GitHub Actions, and GitOps-first workflows
  • Automation: Proficiency in Infrastructure as Code (IaC) using Terraform and configuration management via Ansible
  • Observability: Hands-on experience with Prometheus, Grafana, observability platforms (Splunk or Datadog), and OpenTelemetry standard to build comprehensive telemetry pipelines
  • FinOps: Demonstrated experience in cloud cost optimization, reserved instance management, right-sizing, and long-term capacity planning for multi-tenant SaaS platforms

Experience and Leadership

  • Overall Experience: 12+ years in software engineering, cloud operations, or infrastructure roles
  • Domain Depth: 8–10 years of hands-on experience in SRE, Platform Engineering, Cloud Operations, or Production Support for large-scale, distributed SaaS applications
  • Technical Leadership: Proven track record of leading distributed engineering teams as a player-coach — setting technical direction while remaining hands-on with architecture, automation, and incident response
  • Operational Discipline: Exceptional troubleshooting skills under pressure and a "Fire Marshal" mindset toward investigation and proactive inspection
  • Customer Engagement: Experience working directly with enterprise customers on production reliability, incident communication, and service-level reporting
  • Background Check: Must be able to pass customer-mandated background screening for access to critical infrastructure environments


Desired Characteristics

Regulated Environments

  • Practical knowledge of NERC CIP compliance standards in a SaaS context
  • Experience with SOC2, ISO 27001, or IEC 62443 compliance frameworks
  • Familiarity with operating in highly regulated industries such as utilities, financial services, or critical national infrastructure

Certifications

  • AWS Certification: DevOps Engineer — Professional or Solutions Architect — Associate/Professional
  • CKA: Certified Kubernetes Administrator
  • SRE Practitioner Certification
  • AWS FinOps Practitioner or equivalent cloud financial management certification

Additional Information

About the SRE Team: The Grid Software SRE function is a newly established capability supporting the organization's SaaS transformation. The team currently supports Field Damage Assessment (FDA) and Distributed Dynamic Line Rating (DDLR) applications, with the portfolio expanding as GE Vernova Grid Software scales from its initial SaaS customers to a target of 20+ customers by end of 2027. The team operates a follow-the-sun model with engineers based in Hyderabad Technical Center (India) and Querétaro (Mexico).

Why US-Based: North American utility customers operating critical national infrastructure require that production environments and customer data be managed by US-based personnel who have completed customer-mandated background screening. This role exists to meet that requirement while providing hands-on technical leadership to the global SRE team.


Work Schedule

General shift, US business hours. On-call availability required for P1/Sev1 incidents. Follow-the-sun handoff protocols with Hyderabad and Querétaro teams.

Travel Requirements

Up to 10% — customer sites and team locations as needed (estimated 2–4 trips per year)


Additional Information

GE Vernova offers a great work environment, professional development, challenging careers, and competitive compensation. GE Vernova is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, protected veteran status or other characteristics protected by law.

GE Vernova will only employ those who are legally authorized to work in the United States for this opening. Any offer of employment is conditioned upon the successful completion of a drug screen (as applicable).

Relocation Assistance Provided: No

#LI-Remote - This is a remote position

For candidates applying to a U.S. based position, the pay range for this position is between $151,800.00 and $227,700.00. The Company pays a geographic differential of 110%, 120% or 130% of salary in certain areas. The specific pay offered may be influenced by a variety of factors, including the candidate’s experience, education, and skill set.

Bonus eligibility: discretionary annual bonus.

This posting is expected to remain open for at least seven days after it was posted on September 25, 2026.

Available benefits include medical, dental, vision, and prescription drug coverage; access to Health Coach from GE Vernova, a 24/7 nurse-based resource; and access to the Employee Assistance Program, providing 24/7 confidential assessment, counseling and referral services. Retirement benefits include the GE Vernova Retirement Savings Plan, a tax-advantaged 401(k) savings opportunity with company matching contributions and company retirement contributions, as well as access to Fidelity resources and financial planning consultants. Other benefits include tuition assistance, adoption assistance, paid parental leave, disability benefits, life insurance, 12 paid holidays, and permissive time off.

GE Vernova Inc. or its affiliates (collectively or individually, “GE Vernova”) sponsor certain employee benefit plans or programs GE Vernova reserves the right to terminate, amend, suspend, replace, or modify its benefit plans and programs at any time and for any reason, in its sole discretion. No individual has a vested right to any benefit under a GE Vernova welfare benefit plan or program. This document does not create a contract of employment with any individual.

Similar Jobs

38 Minutes Ago
Easy Apply
Remote
United States
Easy Apply
69K-95K Annually
Mid level
69K-95K Annually
Mid level
Artificial Intelligence • Fintech • Hardware • Information Technology • Sales • Software • Transportation
Provide high-level administrative support to sales executives including calendar, travel, expense reporting, and confidential communication. Prepare reports and sales materials using PowerPoint, Excel, and CRM. Plan and execute internal and external sales events and offsites. Liaise with clients and internal teams, multitask under pressure, and travel a few times a year.
Top Skills: CoupaCRMExcelGoogle SuiteNavanPowerPointSlack
39 Minutes Ago
Remote or Hybrid
Austin, TX, USA
45K-85K Annually
Junior
45K-85K Annually
Junior
Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Handles inbound calls and warm leads, consults customers on insurance needs, recommends appropriate coverage, and converts prospects into policyholders. The role includes paid Property & Casualty licensing and sales training, customer communication, lead conversion, and schedule adherence in a remote environment. Employees must work designated weekday and weekend shifts and maintain a professional home workspace with wired high-speed internet.
Top Skills: PcProperty & Casualty Insurance LicensingWired High-Speed Internet
An Hour Ago
Remote or Hybrid
Junior
Junior
Artificial Intelligence • Cloud • Information Technology • Security • Software • Cybersecurity • Data Privacy
Owns the full sales cycle for new LATAM customers, including outbound prospecting, qualification, technical evaluations, negotiations, and closing. The role also manages and expands existing accounts, supports retention, and develops relationships with DevSecOps, development, and digital transformation stakeholders. Candidates need Spanish or Portuguese fluency, technical SaaS or security sales experience, familiarity with the software development lifecycle, outbound pipeline generation skills, and experience managing varied sales cycles in a remote, collaborative environment.
Top Skills: Ai-Native Developer Security PlatformDevsecopsSaaS

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account