Career Techniques Inc
Description
In this role you will build and own the tooling behind day-to-day GPU operations: management, monitoring, metrics collection, maintenance, and network configuration. When something misbehaves anywhere between an application and the kernel, you will help track it down.
Responsibilities:
- Build and maintain the tooling that automates GPU fleet operations: management, monitoring, metrics collection, maintenance, and network configuration.
- Troubleshoot software and hardware issues across the fleet, from application and network problems down to the operating system and kernel.
- Work with engineering teams across the firm to tune workloads and processes so they use GPUs more efficiently.
- Analyze GPU job statistics to surface trends, inefficiencies, and opportunities for improvement.
Qualifications:
- BS or MS in computer science or a related field.
- 2+ years of relevant experience, including Python development and hands-on GPU management.
- An automation-first mindset. When a workflow is manual, slow, or error-prone, you reach for code.
- Experience deploying, troubleshooting, and tuning a range of GPU hardware.
- Strong computer science fundamentals and sound software design instincts.
- Solid working knowledge of Linux/UNIX and comfort with open-source software.
- Sharp debugging. You get to the bottom of problems quickly and methodically.
- Familiarity with configuration management and monitoring technologies.
- Organized, adaptable, and collaborative. You can juggle several tasks with careful attention to detail, work independently or with the team, and pick up new skills fast.
Comp: 200-300K + Bonus
