Graphcore Logo

Graphcore

Principal System Debug Engineer

Posted 3 Days Ago
Be an Early Applicant
Hybrid
Austin, TX, USA
Expert/Leader
Hybrid
Austin, TX, USA
Expert/Leader
Lead system-level debug and validation for Arm-based server blades and rack platforms. Drive post-silicon bring-up, cross-functional root-cause analysis, debug methodologies, tooling and automation, program metrics, and mentor engineers to ensure POR quality and timely issue resolution.
The summary above was generated by AI
Lead System Debug Engineer – Server & Rack Validation (Principal Level and Above)Position Overview

We are seeking a senior technical leader (Principal Engineer level and above) to lead the bring-up, enablement, and hardware debug of server compute systems and rack-level platforms based on Arm® server architecture.

The successful candidate will be a key member of the System Validation organization, responsible for driving system-level debug activities and facilitating rapid resolution of complex hardware, firmware, and software issues. This role requires close collaboration with engineering teams across the organization to identify root causes, implement corrective actions, and ensure successful program execution.

The ideal candidate will be deeply involved in challenging system debug efforts while developing and executing scalable debug strategies that maximize throughput and ensure Product of Record (POR) quality. In addition, this individual will establish and drive debug methodologies, improve processes, and help create a culture of technical excellence across the organization.

We are looking for a disciplined, dynamic, and highly motivated leader who can thrive in a global environment while fostering strong cross-functional collaboration.

As a Lead System Debug Engineer within Server and Rack Validation, you will drive balanced, scalable, and automated debug solutions that optimize engineering efficiency and product quality. This highly visible role provides the opportunity to innovate and improve debugging capabilities while delivering industry-leading server technologies to market.

Your technical leadership, validation expertise, and problem-solving skills will play a critical role in product development, issue root cause analysis, and resolution. Success in this role requires close collaboration with System Validation, System Architecture, Silicon Engineering, Rack Firmware, and other cross-functional teams.

Primary Responsibilities
  • Develop and drive a Debug Center of Excellence, including scalable debug and triage methodologies, processes, and playbooks for server blade and rack-level issues spanning hardware, firmware, and software integration.

  • Debug issues discovered during server rack bring-up, post-silicon validation, and production phases.

  • Lead complex debug efforts involving silicon, server systems, firmware, and software to determine root causes and drive effective resolutions.

  • Ensure issues are resolved with high quality and within program timelines.

  • Manage and track technical issues, risks, and priorities to remove blockers and achieve key program milestones.

  • Develop and publish debug program metrics and indicators to identify roadblocks and improve overall debug efficiency.

  • Communicate program status, risks, and opportunities to customers, stakeholders, and executive leadership.

  • Drive technical innovation across triage and debug workflows through tool development, scripting, methodology enhancements, and cross-functional engineering initiatives.

  • Mentor engineers and promote best practices in system validation and debug methodologies.

Required Qualifications
  • Strong analytical and problem-solving skills with exceptional attention to detail.

  • Extensive experience in validation and debug roles involving operating systems, firmware, silicon, and hardware issues.

  • Deep understanding of industry-standard server interconnects and software stacks, including PCIe and CXL.

  • Strong knowledge of Arm® CPU or x86 architectures, SoC design, memory subsystems, RAS (Reliability, Availability, and Serviceability), and power management.

  • Extensive experience with system architecture, technical debugging, and validation strategies.

  • Strong understanding of platform-level and system-level debug methodologies, including Operating Systems, Device Drivers, and BIOS interactions.

  • Excellent communication, collaboration, and cross-functional leadership skills.

  • Highly organized and detail-oriented, with the ability to manage multiple priorities and deliver results under tight deadlines.

  • Experience leading technical programs and coordinating cross-functional engineering efforts.

  • Thorough understanding of data center technologies and associated software stacks.

  • Self-motivated with the ability to independently drive tasks from problem identification through resolution.

Preferred Qualifications
  • Master's degree or Ph.D. in Electrical Engineering, Computer Engineering, Computer Science, or a related technical field.

  • Experience with large-scale server platforms, rack-level systems, and hyperscale data center environments.

  • Expertise in automation, scripting, and debug tool development.

  • Experience establishing and scaling debug processes across multiple product generations and engineering organizations.


USA Benefits
In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.


Graphcore Austin, Texas, USA Office

Graphcore Austin Office Office

Austin, TX, United States

Similar Jobs at Graphcore

7 Hours Ago
Hybrid
Austin, TX, USA
Senior level
Senior level
Artificial Intelligence • Semiconductor
As a Principal Hardware Diagnostics Engineer, you will design automated diagnostics software for hardware health monitoring, collaborating across teams to enhance AI infrastructure efficiency.
Top Skills: C#C++LinuxPython
Yesterday
Hybrid
Austin, TX, USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Semiconductor
Define and lead end-to-end security architecture for a large-scale AI inference service platform. Own threat modelling, tenant isolation, platform and hardware security (secure boot, attestation, firmware integrity), network and API protections, secrets and key management, monitoring and incident response, and customer assurance. Partner across engineering, operations, compliance, and customers to set requirements, validate implementations, support audits and pentests, and provide technical leadership and mentorship.
Top Skills: Api SecurityCertificate Lifecycle ManagementCi/Cd Deployment PipelinesConfidential ComputingEncryptionFirmware IntegrityHardware Platform SecurityIncident ResponseKey ManagementLogging And MonitoringMeasured BootNetwork SegmentationPrivileged Access ManagementRemote AttestationSecrets ManagementSecure BootSecure EnclavesSecure Firmware DevelopmentService Control PlanesSupply Chain SecurityTenant IsolationTrusted Execution EnvironmentsZero Trust
2 Days Ago
Hybrid
Austin, TX, USA
Senior level
Senior level
Artificial Intelligence • Semiconductor
Assemble and rework electronic prototypes, perform basic functional testing and troubleshooting of PCBAs, support engineers with test setups, maintain lab equipment and workspace, and manage inventory and procurement of consumables.
Top Skills: Bench Power SupplyInventory Tracking ApplicationsIpc SolderingMultimeterOscilloscopePcb Assembly (Pcba)Soldering (Surface-Mount)SpreadsheetsThrough-Hole Soldering

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account