Career Techniques Inc
Description
RESPONSIBILITIES
- Design, configure, and manage a high-performance compute fleet comprising large-scale GPU (NVIDIA V100/A100/H200/GB200) and CPU nodes across the firm's infrastructure.
- Own the full firmware and BIOS lifecycle across the HPC/AI fleet — from establishing baselines and validation through rollout, compliance, and ongoing maintenance.
- Lead troubleshooting of hardware components including CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs; drive component replacement and configuration remediation.
- Automate health checks, onboarding workflows, and recurring hardware issue remediation to accelerate safe deployment and reduce recovery time.
- Validate and operationalize next-generation AI platforms (e.g., NVL72 / Grace Blackwell) from day one, ensuring stability, performance readiness, and production fitness.
- Collaborate with vendors on firmware and hardware issues, providing clear reproduction cases, diagnostic logs, and business impact to drive timely resolution.
- Perform hardware performance analysis, tuning, and capacity planning to ensure reliable scale-out of the compute environment.
- Define and implement security hardening best practices for hardware infrastructure, maintaining platform integrity across the fleet.
- Leverage Infrastructure as Code (IaC) methodologies and scripting to drive efficient, repeatable, and scalable infrastructure management.
- Mentor junior engineers, act as a subject matter expert for infrastructure-related escalations, and champion a culture of continuous improvement across the team.
REQUIREMENTS
- Bachelor’s degree in Electrical Engineering, Computer Engineering, or a related field, or equivalent hands-on experience.
- 8+ years of experience managing large-scale HPC or AI compute infrastructure in a production environment.
- Deep knowledge of server hardware architecture, including processors, memory, storage, networking, power systems, and thermal management.
- Hands-on experience with bare-metal provisioning, firmware and BIOS lifecycle management, and hardware automation tools such as Ansible, Puppet, or Chef.
- Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management and diagnostics.
- Demonstrated ability to troubleshoot and resolve complex hardware issues across GPU and CPU nodes, including NVIDIA-SMI and GPU diagnostics.
- Experience with hardware monitoring platforms, performance tuning, and capacity planning at scale.
- Familiarity with Linux-based environments and scripting proficiency in Python, Bash, or PowerShell for infrastructure automation.
- Experience with OpenStack (particularly Ironic) or equivalent cloud/bare-metal provisioning platforms is strongly preferred.
- Strong cross-functional communication skills and proven ability to collaborate effectively with software, networking, and vendor teams.
- Prior technical leadership experience, including mentoring engineers and driving team-wide best practices.
