Improve AWS production infrastructure reliability, observability, performance, and operational maturity. Build Terraform infrastructure, enhance CI/CD, automate operational work, manage incident response and on-call operations, lead postmortems, improve application resilience, support capacity planning and database reliability, and collaborate on security hardening and compliance. Mentor engineers and promote reliability practices across the organization.
Senior Site Reliability Engineer
Improve the reliability, performance, and operational maturity of a platform that supports the future of educational fundraising.
CONTRACT-TO-HIRE REMOTE - UNITED STATES SEATTLE / WEST COAST PREFERRED
About the role
Our client is looking for a hands-on Senior Site Reliability Engineer to improve the reliability, performance, and operational maturity of our platform. With our migration to AWS complete, this role will focus on strengthening our production environment: improving observability, automating infrastructure and operational work, enhancing incident response, and partnering with product engineers to build resilient systems. This is a high-impact role for someone who understands both infrastructure and application development. You will work across the stack, contribute code when appropriate, and help ensure our systems remain secure, scalable, and dependable as we grow.
What you'll do
• Operate, maintain, and improve our production infrastructure in AWS.
• Build and maintain infrastructure as code using Terraform.
• Improve monitoring, alerting, dashboards, and service-level indicators using New Relic or comparable observability platforms.
• Reduce alert noise and build systems that identify problems before customers are affected.
• Participate in the 24/7 on-call rotation and help coordinate the response to production incidents.
• Lead blameless postmortems and ensure corrective actions result in durable improvements.
• Partner with product engineers to diagnose performance and reliability issues throughout the application stack.
• Improve application resilience through appropriate use of timeouts, retries, queuing, backpressure, and idempotency.
• Improve CI/CD pipelines and deployment practices using platforms such as GitHub Actions, GitLab CI, or CircleCI.
• Automate repetitive operational work and reduce engineering toil.
• Create and maintain runbooks, system diagrams, troubleshooting guides, and production documentation.
• Support capacity planning, performance testing, database reliability, and production-readiness reviews.
• Collaborate with Security and Engineering teams on infrastructure hardening, access controls, logging, and compliance-related operational practices.
• Mentor engineers and promote effective reliability practices across the Engineering organization
What we're looking for
• 10+ years of overall software engineering, infrastructure, or systems experience, including at least 5 years in an SRE, Platform Engineering, DevOps, or production operations role.
• Previous professional software development experience and the ability to read, debug, and contribute to application code. • Strong, hands-on experience operating production workloads in AWS. • Experience building and maintaining infrastructure with Terraform or a similar infrastructure-as-code tool.
• Strong observability skills using New Relic, Datadog, or another modern monitoring platform. • Experience with incident response, on-call operations, postmortems, and production troubleshooting.
• Experience building or maintaining CI/CD pipelines.
• Working knowledge of networking, Linux, distributed systems, and relational databases.
• Strong judgment when balancing immediate operational needs with long-term maintainability.
• Clear communication skills and the ability to collaborate effectively across engineering disciplines.
• A track record of using automation to improve reliability and create leverage for other engineers.
Bonus points
• Experience with Ruby or Ruby on Rails.
• Strong PostgreSQL administration or performance-tuning experience.
• Experience operating enterprise SaaS products at scale.
• Familiarity with SLOs, SLIs, error budgets, capacity modeling, and load testing.
• Experience with payments, fintech, or other highly regulated systems.
• Experience supporting SOC 2 or similar security and compliance programs. Role details
• Contract-to-hire.
• Remote within the United States.
• Seattle-area or West Coast candidates are preferred to support occasional in-person collaboration, but exceptional candidates elsewhere should also be considered.
• Participation in a shared on-call rotation is required.
Similar Jobs
Digital Media • Information Technology • News + Entertainment
Develop and manage local advertising clients and agencies to achieve sales goals. Create market research, advertising proposals, forecasts, reports, and sales documentation. Prospect for new customers, coordinate advertising schedules and client requirements with internal teams, monitor account activity and collections, and maintain accurate customer records. The role requires independent judgment, punctual attendance, and flexibility to work nights, weekends, variable schedules, and overtime.
Top Skills:
Advertising TechnologyDigital AdvertisingMultiscreen Video AdvertisingTv Advertising
Digital Media • Information Technology • News + Entertainment
Designs and implements complex enterprise network solutions, supports customer pilots and sales, provides Tier IV escalation support, troubleshoots network issues, develops capacity models, evaluates hardware and software, maintains technical documentation and procedures, and conducts security audits. Leads and mentors network engineers, presents technical information, participates in an on-call rotation, and supports customer professional services projects. The role requires independent judgment, occasional travel, and variable night or weekend work.
Top Skills:
AaaAclsAerohiveBgpCiscoCisco Catalyst Sd-WanCisco CceCisco FirepowerCisco MerakiCisco PrimeCradlepointDhcpDmvpnEigrpFortianalyzerFortimanagerFortinet FortigateGlbpGreHipaaHsrpIpsecIwanJuniperMistNetscoutOspfPci DssPolicy RoutingPriQosRadiusRipSd-WanSipSnmpSocSolarwinds OrionStpTacacs+VlansVoipVrrpVtpWhatsup GoldWireless NetworkingWireshark
Cloud • Information Technology • Security • Software • Cybersecurity
Conduct advanced threat research across endpoint and cloud environments, dissecting and replicating adversary techniques to improve detection. Engineer attack automation, test code, and proof-of-concept detections; collaborate with detection engineers, analysts, and threat hunters; document findings; influence product strategy and vendor telemetry; and mentor junior researchers.
Top Skills:
Ai/MlAWSCC++Endpoint Detection And Response (Edr)GitGitGoLinuxmacOSMicrosoft 365OktaPowershellPythonRustWindows
What you need to know about the Austin Tech Scene
Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.
Key Facts About Austin Tech
- Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
- Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
- Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
- Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center


