Career Techniques Inc
Description
You will focus on expanding our ML research platform to benchmark, rapidly prototype, and stress-test both software and hardware layers across our entire distributed ML stack. By leveraging AI agents and auto-research capabilities, you will push our systems to their limits, identify bottlenecks, and create a frictionless environment to test novel machine learning models on realistic, large-scale data.
Responsibilities:
- Platform Validation & Infrastructure Benchmarking:
- Serve as the primary feedback loop for the entire ML stack.
- Actively run complex models through our full ML pipeline to comprehensively test both the training and inference environments.
- Validate the central infrastructure in practice, seeing exactly how new research ideas fare and identifying system bottlenecks before broader rollout to research teams.
- Streamline Rapid Prototyping for ML Research:
- Build high-level abstractions that allow users to bypass setup friction.
- Integrate our core ML tooling directly with our underlying simulation and data frameworks, providing a unified entry point to access our full tech stack.
- Enable rapid iteration on real-world data and seamless distributed training via Ray.
- Agentic Workflows for ML Research:
- Leverage AI agents and auto-research workflows to autonomously generate experiments, stress-test our distributed clusters, and provide data-driven, actionable feedback on what infrastructure needs to be optimized or built next.
- Research Platform Feedback & Insights Sharing:
- Act as the critical bridge between infrastructure builders and ML researchers.
- Be the first to exhaustively test new models and push the platform's limits.
- Document and publish empirical findings on system capabilities and hardware performance.
- Take your validated insights to assist engineering teams with platform improvements and advise researchers on how to best leverage the stack.
Qualifications:
- Strong Software Engineering Foundation:
- Deep proficiency in Python and software design principles.
- Ability to build clean, scalable APIs and abstractions that other developers and researchers are enthusiatic about using.
- Applied Machine Learning:
- Hands-on experience with modern frameworks (PyTorch, TensorFlow, etc.)
- Strong practical understanding of how to train, evaluate, and deploy models at scale.
- Distributed Compute:
- Experience scaling ML workloads across GPUs and multi-node clusters using frameworks like Ray, Dask, or PyTorch Distributed.
- AI Agent Workflows:
- Familiarity with LLM tooling, agentic frameworks, and using AI to automate coding, research, or testing tasks.
- System Profiling & Optimization:
- Ability to debug and identify bottlenecks across hardware and software layers (e.g., memory limits, GPU utilization, data pipeline latency).
Comp: $200-300K + Bonus
