Owns the reliability, security, performance, and availability of AWS-hosted healthcare infrastructure. Builds infrastructure as code, CI/CD automation, monitoring, observability, backup and disaster recovery capabilities. Leads incident response, optimizes cloud resources, implements security controls, supports customer onboarding and migrations, and ensures compliance with healthcare privacy and software lifecycle requirements. Participates in on-call rotations and provides technical guidance to engineering and support teams.
Position Overview
The Cloud Site Reliability Engineer will own the reliability, performance, and security of Cadwell's cloud-hosted infrastructure supporting healthcare facilities that depend on Cadwell's software suite for neurodiagnostic care. Working within the Enterprise Service Delivery team, this role builds and maintains infrastructure as code, deployment automation, monitoring and observability, and disaster recovery capabilities for Cadwell's AWS environments, and leads incident response for hosted customer systems. The Cloud Site Reliability Engineer partners closely with software engineering, enterprise support, and project delivery teams to ensure hosted deployments meet the uptime, data protection, and regulatory expectations of clinical environments.
The Cloud Site Reliability Engineer will own the reliability, performance, and security of Cadwell's cloud-hosted infrastructure supporting healthcare facilities that depend on Cadwell's software suite for neurodiagnostic care. Working within the Enterprise Service Delivery team, this role builds and maintains infrastructure as code, deployment automation, monitoring and observability, and disaster recovery capabilities for Cadwell's AWS environments, and leads incident response for hosted customer systems. The Cloud Site Reliability Engineer partners closely with software engineering, enterprise support, and project delivery teams to ensure hosted deployments meet the uptime, data protection, and regulatory expectations of clinical environments.
Key Responsibilities
- Build, maintain, and continuously improve Cadwell's AWS cloud infrastructure supporting hosted customer environments, applying infrastructure-as-code practices using Terraform, YAML, and JSON/Jinja.
- Automate build, test, and deployment pipelines for cloud-hosted Cadwell applications to reduce manual effort and eliminate configuration drift across customer environments.
- Build and maintain log ingestion, monitoring, alerting, and observability systems that provide early warning of degradation in hosted clinical environments.
- Lead incident response for cloud-hosted environments, serving as the escalation point for availability and performance events and conducting post-mortems with corrective actions tracked to closure.
- Design, implement, and routinely test backup and disaster recovery strategies for hosted customer data to meet the recovery time and recovery point objectives committed to customers.
- Optimize cloud compute, storage, and lifecycle policies to balance performance, clinical data retention requirements, and cost across the hosted footprint.
- Implement and maintain cybersecurity best practices across cloud infrastructure, including identity and access management, network segmentation, encryption, patching, and vulnerability remediation.
- Partner with software engineering to improve application reliability, scalability, and performance, contributing to architecture decisions for new and migrating hosted deployments.
- Partner with Cadwell's enterprise support and project delivery teams on hosted environment onboarding, migrations, and upgrades, providing cloud infrastructure expertise throughout the customer lifecycle.
- Ensure cloud infrastructure and processes comply with applicable healthcare data privacy and security requirements (e.g., HIPAA, GDPR) and with Cadwell's quality system, including IEC 62304 software lifecycle processes.
- Document infrastructure architecture, runbooks, and escalation procedures so that support and on-call staff can operate hosted environments consistently.
- Other activities as directed, assigned, or requested.
Knowledge, Skills, and Abilities
- Expert-level knowledge of core AWS services and architectural patterns, including compute, object storage, networking (VPC, security groups, load balancers), managed databases, and container orchestration with ECS.
- Demonstrated command of infrastructure as code and deployment automation, including Terraform, YAML and JSON/Jinja templating, CI/CD pipeline design, and scripting in Bash, Python, or JavaScript/TypeScript.
- Practical knowledge of cloud security, including identity and access management, least-privilege design, secret management, encryption in transit and at rest, network segmentation, and vulnerability remediation.
- Hands-on experience building monitoring, logging, and observability systems, and designing alerting that surfaces real problems without creating noise.
- Sound judgment in backup, disaster recovery, and business continuity planning, including replication, restore validation, and recovery objectives appropriate to clinical data.
- Solid grounding in managed database administration and cloud cost management, including backups, patching, performance tuning, right-sizing, storage lifecycle policies, and capacity planning.
- Familiarity with healthcare data privacy and security requirements such as HIPAA and GDPR, and with software lifecycle standards such as IEC 62304.
- Ability to explain complex technical concepts clearly to technical and non-technical audiences, lead incident response under pressure, and mentor engineers and support staff on cloud practices.
- Ability to work both independently and as an effective team member, exhibiting Cadwell's core values of integrity, initiative, accountability, commitment to excellence, and service to others.
Education and Experience
- Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent work experience.
- 8+ years of experience in site reliability engineering, cloud infrastructure, or DevOps, including recent hands-on experience operating production AWS environments.
- Demonstrated experience operating production infrastructure under formal on-call and incident management practices.
- Experience working in a regulated environment (medical device, healthcare, or financial services) preferred.
- 2+ years of support within a healthcare environment preferred.
- AWS certification such as Solutions Architect or SysOps Administrator preferred.
- Familiarity with virtualization technologies such as VMware, Hyper-V, or Citrix is preferred.
- Familiarity with HL7 integration tooling such as Mirth Connect is preferred.
Physical Requirements and Working Conditions
- Positions working with Cadwell equipment generally may require some reaching, bending, stooping, squatting, crawling, kneeling, pushing, pulling, and occasional lifting and carrying up to 50 pounds, finger dexterity, repetitive motions, standing, walking, sitting, hearing, visual acuity, color vision, and 2-way written/verbal communication. More specific details may be provided as needed or requested.
- Extensive use of a computer will be required.
- Reliable high-speed internet required.
- Participation in an on-call rotation for hosted environment escalations, including occasional after-hours and weekend response, is required.
- Infrastructure maintenance is frequently scheduled outside of customer clinical hours; flexibility to work evenings or weekends for planned changes is required.
- Travel up to 10% for company meetings.
Cadwell Industries, Inc. is an Equal Opportunity Employer, and as such affirms the right of every person to participate in all aspects of employment without regard to race, religion, color, national origin, citizenship, sex, sexual orientation, gender identity, age, veteran status, disability, genetic information, or any other protected characteristic. If you are interested in applying for employment and need special assistance or an accommodation to apply for a posted position, contact our Human Resources department at [email protected].
Salary Range
$120,000—$130,000 USD
Similar Jobs
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Operate and improve large-scale Kubernetes and GPU clusters across public and private clouds. Build automation, observability, capacity-management, and reliability systems; define SLOs and SLIs; support production launches; lead incident triage and root-cause analysis; conduct blameless postmortems; and participate in on-call support for AI workloads.
Top Skills:
AirflowAnsibleArgo WorkflowsAWSAws Step FunctionsCadenceChefCudaElk StackGoGoogle Cloud PlatformGrafanaKubernetesKubevirtLightstepLinuxAzureNcclNvidia Dgx CloudNvidia DynamoOpentelemetryOracle Cloud InfrastructurePrometheusPuppetPythonPyTorchSglangSplunkTcp/IpTemporalTensorrt-LlmTerraformVllm
Information Technology • Security • Cybersecurity
Own production operations and highly available cloud infrastructure for regulated government environments. Lead incident response, root-cause analysis, disaster recovery testing, compliance operationalization, audit readiness, vulnerability management, and continuous monitoring. Build automation, secure CI/CD pipelines, infrastructure-as-code, observability, and compliance tooling across Kubernetes, Linux, containers, and cloud platforms. Partner with security, compliance, and engineering teams to improve reliability, deployment safety, and regulatory sustainment.
Top Skills:
Aws GovcloudBashCi/CdDod Il4Dod Il5FedrampGitopsGoGrafanaKubernetesLinuxNist 800-53PrometheusPythonStigTerraformUnixZero Trust
Automotive
Leads SRE engineering leaders and engineers while defining enterprise observability, reliability, and platform strategy across GCP, on-premise, manufacturing, distribution, and campus environments. Oversees vendor-agnostic tooling, OpenTelemetry integrations, CI/CD observability, SRE maturity models, and Agentic AI initiatives. Drives adoption of SRE practices, develops technical roadmaps, partners with senior leadership and operational teams, and maintains hands-on architectural and technical credibility.
Top Skills:
Agentic AiAWSAzureCi/CdDatadogDynatraceGCPNew RelicOpentelemetryOtel Genai Semantic ConventionsSource Control PlatformsSplunkTerraform
What you need to know about the Austin Tech Scene
Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.
Key Facts About Austin Tech
- Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
- Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
- Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
- Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center


