San Francisco Compute Company Logo

San Francisco Compute Company

HPC/ GPU Cluster Architect

Reposted One Month Ago
Be an Early Applicant
Remote or Hybrid
Hiring Remotely in San Francisco, CA
220K-300K Annually
Senior level
Remote or Hybrid
Hiring Remotely in San Francisco, CA
220K-300K Annually
Senior level
As an HPC/GPU Cluster Architect, you will design, architect, and scale GPU clusters, ensuring performance and reliability while automating fleet operations and mentoring junior engineers.
The summary above was generated by AI
About San Francisco Compute

San Francisco Compute exists to give ambitious teams large scale supercomputers without taking on a catastrophic balance sheet risk. Most GPU clouds force you to sign long term contracts with no way out. If you underbuy, you miss the window. If you overbuy, you’ll blow up with no way out. It’s a desperate gamble that San Francisco Compute does not force you to play.

Originally, SFC was an AI lab that bought too big of a GPU cluster and was forced to sublease it or go under. Today, we build data centers, supercomputers, and then lease those machines on contracts that let our customers sublease. More than anyone else in the industry, we deeply felt the pain and fear of being forced to bet the farm with your neocloud vendor. To make it possible to sublease, we invented the first compute market and became the first one to offer physical settlement. That means we operate GPU clusters like a neocloud and derisk our customers like a market.

Customers buying GPU clusters don’t want to work through a network of brokers. They want to transact with the folks that have the keys to the cluster and can own the SLA.

To keep the industry safe from unnecessary risk, SFC needs to scale fast. To do that, we serve financially motivated parties as a technical partner to build, operate, & lease supercomputers on their behalf. Basically, we help folks who want an economic return build GPU clusters on their balance sheet, but operated in full by us. That allows us to retain the technical control needed to operate a cloud, but scale fast like a market. SFC’s ability to offer subleasing can significantly increase the levered returns for cluster owners. Our hybrid model out-performs brokered contracts or other “compute markets”, who have to go through a network of counterparties to solve problems. It also lets us design custom solutions for our customers, like clusters deployed in regions next to your dataset or physical eval set or a large CPU fleet deployed colocated with your GPU cluster.

Our team includes senior & technical leadership from places like Lambda, Crusoe, Digital Ocean, AWS, and Hut8. Our CTO is the cofounder and former CEO of Voltage Park. In prior roles, our team has deployed 8GW of datacenter capacity & hundreds of thousands of GPUs. We’ve been described as having “the highest talent density in the space.” If you are an honorable, gritty person with eyes wide open & good epistemics, who wants to see AI go right, we’d love to work with you. There has never been a more critical moment in history and it is up to us to shape it on behalf of those who come after us.

About the Role

GPU clusters are some of the most performant computers on the planet. Even smaller clusters by today’s standards would have ranked in the TOP500 five years ago. Our infrastructure team is responsible for architecting and deploying new clusters around the world and keeping them running smoothly. You’ll participate in on-call rotation, deploy new environments, fix issues when they arise, and lean into automation to enable deployments at scale. We’re a small but ambitious team so you’ll be an early contributor helping to shape culture, mentor junior engineers, and learn from our customers.

About You
  • You will have 5+ years of experience with hands-on designing, architecting and scaling at least one HPC or GPU compute cluster in production (ideally >1,000 GPUs, but not required)

  • You deeply understand server hardware fundamentals, including GPUs, NICs, PCIe, memory, thermals, and power

  • You’re comfortable debugging performance and reliability issues across hardware, OS, drivers, and networking layers; full-stack.

  • The idea of automating fleet operations (provisioning, monitoring, remediation) excites you — you embrace infrastructure-as-code

  • You appreciate, value, and generate strong operational documentation and runbooks

  • You have the ability and willingness to mentor junior engineers and contribute to team culture

  • You’re open to domestic travel when required

Some Nice to Haves
  • Familiarity with data center operations including power, cooling, and colo/vendor engagements

  • Strong Linux systems administration experience, including kernel drivers, RDMA stack tuning, and performance analysis

  • Experience with schedulers and orchestration systems such as Slurm and Kubernetes

  • Exposure to virtualization technologies (KVM, QEMU, libvirt)

  • Experience utilizing telemetry pipelines for predictive hardware failure detection

  • Experience troubleshooting high-speed fabrics such as InfiniBand and/or RoCEv2 Ethernet

BenefitsGenerous equity grant

Team members are offered a competitive salary along with equity in the company

Visa Sponsorships

Yes, we sponsor visas and work permits

Retirement matching

We match 401(k) plans up to 4%

Medical, dental & vision

We offer competitive medical, dental, vision insurance for employees and dependents and cover 100% of premiums

Time off

We offer unlimited paid time off as well as 10+ observed holidays

Parental leave

We offer biological, adoptive, and foster parents paid time off to spend quality time with family

Daily lunch

We cover lunch daily for employees

Unlimited office book budget

You can buy as many books for the office as you want

The San Francisco Compute Company is committed to maintaining a workplace free from discrimination and harassment.

We make employment decisions based on business needs, job requirements, and individual qualifications, without regard to race, color, religion, belief, national origin, social or ethical origin, age, physical, mental, or sensory disability, sexual orientation, gender identity or expression, marital status, civil union or domestic partnership status, past or present military service, HIV status, family medical history or genetic information, family or parental status including pregnancy, or any other status protected by law.

We welcome the opportunity to consider qualified applicants with prior arrest or conviction records. Our commitment to diversity includes hiring talented individuals regardless of their criminal history, in accordance with local, state, and federal laws, including San Francisco’s Fair Chance Ordinance and California’s ban-the-box laws.

Similar Jobs

2 Hours Ago
Remote
United States
136K-198K Annually
Senior level
136K-198K Annually
Senior level
Artificial Intelligence • Information Technology • Professional Services • Software • Analytics • Generative AI • Big Data Analytics
Leads analytics strategy and technical delivery for clients, specializing in Adobe Analytics and Customer Journey Analytics. Conducts measurement workshops, defines KPI and implementation strategies, oversees tagging and data quality, guides experimentation, and reviews dashboards and insights. Advises senior stakeholders, manages and mentors consultants, supports proposals and account growth, and contributes to Adobe Experience Platform initiatives involving Target, RT-CDP, and AJO.
Top Skills: Adobe AnalyticsAdobe Customer Journey AnalyticsAdobe Experience PlatformAdobe Journey Optimizer (Ajo)Adobe TargetCcpaData Layer ArchitectureGdprReal-Time Customer Data Platform (Rt-Cdp)
3 Hours Ago
Remote or Hybrid
United States
93K-115K Annually
Senior level
93K-115K Annually
Senior level
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads the design, governance, implementation, and continuous improvement of marketing review processes. Coordinates Marketing, Legal, Compliance, technology, vendors, and audit stakeholders; oversees workflow tools, controls, documentation, reporting, training, access management, and issue resolution. Improves transparency, audit readiness, user experience, regulatory adherence, and speed to market for marketing content approvals.
Top Skills: Adobe WorkfrontAi ToolsMicrosoft 365Microsoft TeamsRedoakSharepointWorkflow Platforms
4 Hours Ago
Remote or Hybrid
172K-301K Annually
Senior level
172K-301K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Architects and operates production-grade agentic AI systems for enterprise identity security. Responsibilities include designing multi-agent orchestration, tool use, planning, memory, retrieval-augmented generation, model integration, evaluation, safety guardrails, observability, and human-in-the-loop controls. The role provides technical leadership through architecture ownership, code reviews, mentoring, and establishing scalable production AI practices.
Top Skills: Anthropic SdkC++Distributed SystemsGoGoogle Ai SdkHybrid SearchJavaMlopsOpenai SdkPythonRagSemantic SearchVector Stores

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account