GPU Systems Engineer

Posted · Add Comment
Career Techniques Inc
Published
August 12, 2026
Location
New York, NY - Hybrid - 4 days/week in-office
Category
 
Job Type

Description

As part of R&D, you will join the engineers responsible for the compute, storage, operating systems, and automation behind that work at serious scale: hundreds of petabytes of storage and large CPU and GPU clusters spanning thousands of nodes.

The role is broad by design. One week you might be shaping the architecture of a new AI cluster, the next profiling a training job that will not scale, the next writing automation that keeps the whole fleet healthy with minimal human intervention.

 

Responsibilities:

  • Design, deploy, and scale distributed GPU clusters, from hardware selection and network topology through to production operation.
  • Track down performance bottlenecks across the full stack: compute, storage, network, and the seams between them.
  • Partner with researchers to profile and benchmark GPU workloads, then turn the findings into measurable speedups.
  • Build the automation that lets a small team operate thousands of nodes: provisioning, monitoring, diagnostics, and self-healing.
  • Own infrastructure projects end to end, from scope and design through implementation and long-term support.
  • Qualify new generations of hardware and software, and work directly with vendors to root-cause complex issues.

Qualifications:

  • 5+ years engineering large-scale Linux systems in HPC, AI, or distributed-infrastructure environments.
  • Deep Linux fundamentals: installation, performance tuning, and debugging, down to the kernel when the problem calls for it.
  • Hands-on troubleshooting of distributed GPU workloads, with a strong mental model of GPU performance.
  • Working experience with GPUDirect RDMA. You understand how data moves between GPUs and the network, and what to check when it does not.
  • Solid Python for automation and tooling, plus CUDA or C/C++ experience. You can read, profile, and debug GPU code, not just operate the clusters it runs on.
  • Familiarity with configuration management tools such as Salt, Ansible, Puppet, or Chef.
  • Comfort diagnosing problems that cross hardware, OS, and network boundaries rather than stopping at one layer.
  • Clear communication. You will work daily with researchers, engineers, and vendors.

 

Comp: 200-300K + Bonus

  • Max. file size: 100 MB.
  • Please complete the math question to prove you are human.

Related Jobs

Machine Learning Research Engineer   New York, NY - Hybrid - 4 days/week in-office new
August 12, 2026
Machine Leaning Performance Engineer (Inference)   New York, NY - Hybrid - 4 days/week in-office new
August 12, 2026
Software Engineer, Machine Lifecycle   New York, NY - Hybrid - 4 days/week in-office new
August 12, 2026
HPC Operations Engineer   New York, NY - Hybrid - 4 days/week in-office new
August 12, 2026
Software Engineer, GPU Fleet   New York, NY - Hybrid - 4 days/week in-office new
August 12, 2026