Sr HPC Hardware Engineer

Posted · Add Comment
Career Techniques Inc
Published
August 27, 2026
Location
Dallas, TX - Hybrid - 3 days/week in-office
Category
 
Job Type

Description

RESPONSIBILITIES

  • Design, configure, and manage a high-performance compute fleet comprising large-scale GPU (NVIDIA V100/A100/H200/GB200) and CPU nodes across the firm's infrastructure.
  • Own the full firmware and BIOS lifecycle across the HPC/AI fleet — from establishing baselines and validation through rollout, compliance, and ongoing maintenance.
  • Lead troubleshooting of hardware components including CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs; drive component replacement and configuration remediation.
  • Automate health checks, onboarding workflows, and recurring hardware issue remediation to accelerate safe deployment and reduce recovery time.
  • Validate and operationalize next-generation AI platforms (e.g., NVL72 / Grace Blackwell) from day one, ensuring stability, performance readiness, and production fitness.
  • Collaborate with vendors on firmware and hardware issues, providing clear reproduction cases, diagnostic logs, and business impact to drive timely resolution.
  • Perform hardware performance analysis, tuning, and capacity planning to ensure reliable scale-out of the compute environment.
  • Define and implement security hardening best practices for hardware infrastructure, maintaining platform integrity across the fleet.
  • Leverage Infrastructure as Code (IaC) methodologies and scripting to drive efficient, repeatable, and scalable infrastructure management.
  • Mentor junior engineers, act as a subject matter expert for infrastructure-related escalations, and champion a culture of continuous improvement across the team.

REQUIREMENTS

  • Bachelor’s degree in Electrical Engineering, Computer Engineering, or a related field, or equivalent hands-on experience.
  • 8+ years of experience managing large-scale HPC or AI compute infrastructure in a production environment.
  • Deep knowledge of server hardware architecture, including processors, memory, storage, networking, power systems, and thermal management.
  • Hands-on experience with bare-metal provisioning, firmware and BIOS lifecycle management, and hardware automation tools such as Ansible, Puppet, or Chef.
  • Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management and diagnostics.
  • Demonstrated ability to troubleshoot and resolve complex hardware issues across GPU and CPU nodes, including NVIDIA-SMI and GPU diagnostics.
  • Experience with hardware monitoring platforms, performance tuning, and capacity planning at scale.
  • Familiarity with Linux-based environments and scripting proficiency in Python, Bash, or PowerShell for infrastructure automation.
  • Experience with OpenStack (particularly Ironic) or equivalent cloud/bare-metal provisioning platforms is strongly preferred.
  • Strong cross-functional communication skills and proven ability to collaborate effectively with software, networking, and vendor teams.
  • Prior technical leadership experience, including mentoring engineers and driving team-wide best practices.
  • Max. file size: 100 MB.
  • Please complete the math question to prove you are human.

Related Jobs

Senior Technical Program Manager - HPC   Dallas, TX - Hybrid - 3 days/week in-office new
August 27, 2026
Machine Learning Research Engineer   New York, NY - Hybrid - 4 days/week in-office
August 12, 2026
Machine Learning Performance Engineer (Inference)   New York, NY - Hybrid - 4 days/week in-office
August 12, 2026
Software Engineer, Machine Lifecycle   New York, NY - Hybrid - 4 days/week in-office
August 12, 2026
HPC Operations Engineer   New York, NY - Hybrid - 4 days/week in-office
August 12, 2026