Tech Holding Logo

Tech Holding

Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Lead Site Reliability Engineer responsible for establishing performance, capacity, and reliability baselines; defining SLOs/SLIs and error budgets; instrumenting full request paths; identifying bottlenecks; leading load/stress/resilience testing; building capacity and cost models; creating runbooks and production-readiness criteria; and partnering with engineering and leadership to drive remediation and scalable architecture.
The summary above was generated by AI

About us:

Working at Tech Holding isn't just a job, it's an opportunity to be a part of something bigger. We are a full-service consulting firm that was founded on the premise of delivering predictable outcomes and high-quality solutions to our clients.  Our founders and team members have industry experience and have held senior positions in a wide variety of companies – from emerging startups to large Fortune 50 firms – and we have taken our combined experiences and developed a unique approach that is supported by the principles of deep expertise, integrity, transparency, and dependability.

The Role:

We are looking for a hands-on Lead Site Reliability Engineer for a project based assignment to establish and continuously improve the performance, reliability, and scalability of our platform.This role will define what the platform can reliably sustain today, identify where constraints will emerge, and ensure the organization is prepared to scale before demand arrives.
This is not a traditional DevOps role or an advisory architecture position. You will work directly across application services, infrastructure, databases, networking, caching, queues, external dependencies, and operational processes to identify bottlenecks, validate system limits, and lead remediation.
You will partner closely with engineering, product, and leadership to provide clear, evidence-based answers around capacity, performance, reliability, risk, and the cost of scaling.

Key Responsibilities:

  • Establish performance, throughput, latency, and capacity baselines for critical customer and platform workflows
  • Define and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds
  • Instrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services
  • Identify system bottlenecks and lead cross-functional remediation efforts with engineering teams
  • Build capacity models that show what the platform can sustain, where constraints will emerge, and what additional scale will cost
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments
  • Develop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events
  • Drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements
  • Partner with Test Automation and Scalability Engineering to establish automated performance testing, regression coverage, and production release gates
  • Own technical readiness assessments for major pilots, partnerships, and production launches
  • Create operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures
  • Lead performance and reliability investigations during incidents and ensure lessons are incorporated into future engineering work
  • Make infrastructure cost, performance, and reliability tradeoffs visible to engineering and executive leadership
  • Recommend capacity and reliability investments before they become production constraints

Required Skills:

  • Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline
  • Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements
  • Deep understanding of observability, performance analysis, capacity planning, and reliability engineering
  • Strong hands-on experience with cloud infrastructure and production distributed systems
  • Deep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes
  • Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics
  • Hands-on experience performing load, stress, soak, scalability, and resilience testing
  • Ability to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements
  • Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios
  • Strong incident management and root-cause analysis experience
  • Ability to translate technical performance and reliability risks into clear business implications for senior leadership
  • Strong judgment around when systems genuinely require optimization versus when additional complexity is premature

Nice to have:

  • Experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms
  • Experience creating capacity-cost models and forecasting infrastructure requirements
  • Experience building performance and reliability gates into CI/CD pipelines
  • Experience preparing platforms for significant increases in traffic associated with enterprise customers or strategic partnerships
  • Experience leading reliability or performance initiatives that span multiple engineering teams
What Success Looks Like
Within your first several months, you will have:
  • Established measurable throughput, latency, and capacity baselines for critical platform journeys
  • Defined initial SLOs, error budgets, dashboards, alerts, and performance thresholds
  • Identified the platform's most significant scalability and reliability constraints and created an actionable remediation roadmap
  • Validated representative high-scale scenarios through load, soak, stress, failure, and recovery testing
  • Developed a capacity and cost model showing how the platform can support significant increases in demand
  • Established production-readiness criteria and clear go/no-go evidence for major launches and partnerships
  • Created repeatable scale-up, incident, rollback, and dependency-failure runbooks
  • Given leadership a clear, evidence-based understanding of the platform's current capacity envelope and future scaling requirements
Employment type:
  • Contract

*Applicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time

 
#LI-Remote

Tech Holding is proud to be an Equal Opportunity Employer and is committed to fostering a diverse and inclusive workplace. We welcome applicants from all backgrounds and experiences, and we consider qualified applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected characteristic. If you require accommodation in the application process, please contact our HR 

Similar Jobs

3 Days Ago
Remote
USA
Senior level
Senior level
Gaming • Software
The Site Reliability Engineer will manage infrastructure stability and scalability, lead cloud migrations, and optimize performance across systems while mentoring team members.
Top Skills: AnsibleAWSAzureBashChefCloudFormationDatadogDockerElk StackGCPGoGrafanaKubernetesPrometheusPuppetPythonTerraformUnix/Linux
3 Days Ago
Remote
United States of America
185K-227K Annually
Senior level
185K-227K Annually
Senior level
Other
The Senior Site Reliability Engineer at Juul Labs ensures operational stability and performance of hybrid cloud infrastructure, leads automation, and handles critical incidents.
Top Skills: AWSBashCloudFormationGCPNutanixPowershellPythonTerraform
32 Seconds Ago
Remote or Hybrid
United States
132K-166K Annually
Senior level
132K-166K Annually
Senior level
Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
The Senior Data Scientist partners with Customer Success to perform deep-dive analyses, diagnose metric shifts, build predictive models (churn/upsell), apply causal inference and experimentation, measure business impact, and develop AI/LLM solutions. Responsibilities include self-serving across the data stack for extraction, pipelining, and dashboards, defining KPIs, and delivering actionable insights to reduce churn, drive revenue, and improve operational efficiency.
Top Skills: Ai/LlmAirflowDbtPythonSigmaSQLTableau

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account