Tract Capital Management Logo

Tract Capital Management

Senior Reliability Engineer

Reposted One Month Ago
Be an Early Applicant
In-Office
Austin, TX, USA
120K-150K Annually
Senior level
In-Office
Austin, TX, USA
120K-150K Annually
Senior level
The Senior Reliability Engineer leads availability modeling and reliability analysis for power solutions, ensuring designs meet targeted availability and resiliency. Responsibilities include identifying risks, defining performance targets, and implementing continuous improvement initiatives in a hybrid environment.
The summary above was generated by AI

Tract Capital adopts a unique approach to digital infrastructure investment. Leveraging experience and strategic insights honed over three decades of creating successful companies in the space, we excel in nurturing and advancing leading-edge digital infrastructure enterprises. Our team of specialized experts are united by a singular purpose: to support the growth of digital infrastructure. Tract Capital goes beyond simple investment by acting as a strategic partner and catalyst for innovation within the sector. We ensure our engagements not only generate strong financial results but also develop essential digital infrastructure to meet growing demands. Tract Capital has introduced multiple digital infrastructure strategies including horizontal powered land, vertical development, and product development.


The Senior Availability / Reliability Engineer leads availability modeling, reliability analysis, and mitigation planning for our behind-the-meter (BTM) power solutions and site-specific conditions. This role partners with engineering, construction, commissioning, and operations to identify risks early, define mitigations, and ensure designs and operating models meet target availability and resiliency.

 

How will you make a difference?

The successful candidate will have experience and practical expertise in the following:

  • Own availability and reliability analysis for BTM power solutions across a variety of technologies such as gas reciprocating engines, turbines, fuel cells, batteries, and site deployments (fault trees, reliability block diagrams, Monte Carlo or scenario modeling as appropriate).
  • Define availability targets and performance assumptions; align with customer requirements and Fleet’s uptime objectives.
  • Identify single points of failure and operational risks; recommend design, controls, procedural, or spares mitigations.
  • Partner with engineering teams to validate redundancy strategies, maintainability, and test/maintenance windows that preserve service availability.
  • Support commissioning readiness by defining test scenarios and success criteria that validate reliability assumptions.
  • Develop Quality Control KPI definitions and reporting for reliability performance (forced outage rate, MTTR, maintenance compliance) and drive continuous improvement.
  • Run Reliasoft or IEEE Goldbook calculations to demonstrate facility uptime based on selection of generation and distribution equipment.
  • Lead root cause analysis and corrective action tracking for reliability-impacting events; ensure lessons learned feed back into standards and roadmaps.
  • Collaborate with vendors and operations teams on maintenance strategies, spares/critical parts planning, and reliability-centered maintenance principles.


Basic Qualifications

  • Bachelor’s degree in Engineering (Electrical, Mechanical, Industrial, or similar).
  • 7+ years in reliability engineering, availability analysis, quality processes and/or asset performance engineering in mission-critical or industrial environments.


Required Qualifications

  • Extensive hands-on experience leading availability and reliability analysis for multi-technology BTM power systems (e.g., gas engines, turbines, fuel cells, batteries)
  • Proficiency with reliability modeling methods and tools (fault trees, reliability block diagrams, Monte Carlo or similar, Reliasoft/IEEE Goldbook) to set and validate availability targets
  • Demonstrated ability to identify single points of failure and operational risks and implement design, controls, procedural, or spares mitigations that measurably improve uptime
  • Experience defining and reporting reliability and quality control KPIs (e.g., forced outage rate, MTTR, maintenance compliance) and driving continuous improvement initiatives
  • Track record leading root cause analysis and corrective actions for reliability-impacting events and collaborating with operations and vendors on maintenance strategies and reliability-centered maintenance


Preferred Qualifications

  • Experience with generation assets and integration into critical electrical systems.
  • Experience with CMMS data, failure coding, and maintenance program optimization.
  • Familiarity with safety and operating discipline (MOP/SOP/EOP, change management, incident response).
  • Experience with ISO and Quality metrics for generation assets.
  • Experience communicating technical risk to executives and customers.


Required Traits

  • Integrity and Ethical Standards: Build trust, ensure fairness, and foster long-term, transparent relationships with suppliers. 
  • Effective Communication: The ability to clearly convey expectations and requirements to suppliers and negotiation parties, while understanding their needs and concerns. Comfortable delivering written and verbal presentations to internal leadership teams. 
  • Emotional Intelligence (EQ): Ability to understand the emotions, cultural nuances, and motivations of others, while effectively managing one's own emotions during high-pressure negotiations. 
  • Strategic Thinking: Recognize how supplier relationships and negotiations align with the broader organizational goals, while aiming for outcomes that benefit both parties. 
  • Critical Thinking Skills: Finding innovative solutions and being flexible in addressing unexpected challenges. 
  • Analytical Ability: Make data-driven decisions, assess cost structures, and identify potential risks, ensuring informed and strategic outcomes.  
  • Influence and Persuasion: Able to effectively advocate for their position, build consensus, and secure favorable agreements without compromising relationships. 
  • Operational Paranoia: Anticipate risks, identify vulnerabilities, and proactively implement mechanisms to prevent and minimize disruptions and safeguard safety, security, availability, and scale.  
  • Relationship Management: Cultivate trust, collaboration, and long-term partnerships, while building a broad network that provides valuable benchmarking, industry insights, and alternative sourcing options.  

 

Location and Travel

  • Work location is flexible to Seattle, WA, Denver, CO, or Alexandria, VA. Hybrid: 3-days in office.
  • Regular travel, as needed, to offices as well as to meet with Vendors.

 

Expected Salary Range

  • $120,000 - $150,000 Salary + Discretionary Bonus


#LI-Hybrid

Tract Capital Employment

Tract Capital employees enjoy competitive compensation and comprehensive benefits, including 100% employer-covered medical, dental, and vision insurance, a 401K program, standard paid holidays, and unlimited PTO.

NOTE: This job description is not intended to be all-inclusive. Employees may perform other related duties to meet the organization's ongoing needs.

Tract Capital is proud to be an Equal Opportunity Employer. Qualified applicants are considered for employment regardless of age, race, color, religion, sex, national origin, sexual orientation, gender identity, disability, or veteran status. If you need assistance applying for any of our open positions, please contact us at [email protected].

Similar Jobs

4 Days Ago
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills: AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
5 Days Ago
Easy Apply
Remote or Hybrid
USA
Easy Apply
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills: AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
11 Days Ago
Easy Apply
Remote or Hybrid
USA
Easy Apply
225K-265K Annually
Senior level
225K-265K Annually
Senior level
Fintech • Information Technology • Software • Financial Services
Own the durability, recoverability, performance, and security of a production PostgreSQL/RDS fleet supporting a live trading platform. Lead replication, failover, backup and restore, disaster-recovery drills, data lifecycle management, access control, encryption, and database observability. Investigate engine-level performance issues including WAL contention, replica lag, bloat, and locking. Build infrastructure and AI-assisted operational tooling while documenting runbooks and reliability decisions.
Top Skills: AlloydbAmazon AuroraAmazon RdsAWSBashBigQueryElkGrafanaKafkaKubernetesLinuxPostgresPrometheusPythonSQLTerraform

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account