HPC Operations Engineer

Posted · Add Comment
Career Techniques Inc
Published
August 12, 2026
Location
New York, NY - Hybrid - 4 days/week in-office
Category
 
Job Type

Description

This is an operations role, not a platform engineering one. It is about the daily health of the Research compute fleet: you will be the first line of support for HPC users and the primary owner of day-to-day operations across scheduling, compute, storage, and access. The work is transactional by nature, with tickets, triage, provisioning, and maintenance done well, every day.

You will keep an eye on system health, queues, node status, and service availability; work job failures, scheduler errors, and resource constraints as they come in; and drive every issue to resolution or a clean, well-documented escalation.

Responsibilities:

  • Provide first-line support for HPC users across scheduling, compute, storage, and access issues.
  • Troubleshoot job failures, scheduler errors, and resource constraints, driving each issue to resolution or a clean handoff.
  • Triage infrastructure incidents: gather diagnostics, apply known fixes, and escalate to subject-matter experts when a problem extends beyond defined ownership.
  • Monitor fleet health (queues, node status, storage, and service availability) and act on what you see before users have to report it.
  • Carry out established operational procedures for maintenance, patching, and configuration updates across the Research fleet.
  • Provision new machines into the Research fleet (OS installation, configuration, validation, and handoff into service), and handle reinstalls and decommissions as routine work.
  • Write and maintain runbooks, knowledge-base articles, and user guides so the next occurrence of a problem is faster to fix than the first.
  • Spot recurring issues and propose practical refinements, such as better workflows or automation candidates the HPC team can pick up, so the same ticket stops coming back.

Qualifications:

  • A bachelor's degree in computer science, engineering, or a related field, or equivalent practical experience.
  • 2+ years supporting Linux-based production environments.
  • Solid Linux administration fundamentals (RHEL-family and/or Ubuntu).
  • Methodical troubleshooting: you work a problem step by step, know what you have ruled out, and recognize when it is time to escalate.
  • Experience working directly with users in a technical support or operations role.
  • Strong written communication: clear tickets, clear runbooks, clear handoffs.
  • The discipline to follow established processes with genuine attention to detail.

 

Nice to Have:

  • Enough Bash or Python to script away routine operational tasks.
  • Working knowledge of batch schedulers such as Slurm, HTCondor, or LSF.
  • A good grasp of the plumbing behind networked computing: NFS, automounter, LDAP.
  • Hands-on experience provisioning Linux machines: network boot (PXE), unattended installs (kickstart), or configuration management such as Ansible.
  • Prior exposure to HPC or other large-scale compute environments.
  • Familiarity with monitoring and observability stacks such as Prometheus and Grafana.

 

Comp: $175-225K + Bonus

  • Max. file size: 100 MB.
  • Please complete the math question to prove you are human.

Related Jobs

Machine Learning Research Engineer   New York, NY - Hybrid - 4 days/week in-office new
August 12, 2026
Machine Leaning Performance Engineer (Inference)   New York, NY - Hybrid - 4 days/week in-office new
August 12, 2026
Software Engineer, Machine Lifecycle   New York, NY - Hybrid - 4 days/week in-office new
August 12, 2026
GPU Systems Engineer   New York, NY - Hybrid - 4 days/week in-office new
August 12, 2026
Software Engineer, GPU Fleet   New York, NY - Hybrid - 4 days/week in-office new
August 12, 2026