Staff Software Engineer, HPC
Zoox · Foster City, CA
Posted 38d ago · first seen by the radar 6d ago · last checked on the employer's board 4d ago
Staff · Hybrid · Full-time · $201k–$315k
In this role, you will:
-
Design and implement core services and abstractions for distributed compute infrastructure supporting hundreds of thousands of concurrent jobs
-
Work with customer teams and other infrastructure teams to build a multiyear software engineering roadmap for the HPC platform
-
Lead multi-quarter, cross team initiatives that drive org-wide improvements
-
Create production-grade APIs, SDKs, and tools that make it easy for engineers across Zoox to run large-scale distributed workloads
-
Design and improve job scheduling algorithms and auto-scaling policies to maximize reliability and resource availability
-
Design multi-region orchestration strategies that optimize for data locality, reliability, and performance
-
Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners across multiple teams
-
Evaluate new technologies and paradigms that improve Zoox's computational and storage capabilities
-
Develop capacity planning tools and forecasting models to support Zoox's growing compute needs
-
Mentor junior engineers, guiding them through their career development
Qualifications
-
Experience designing and operating large-scale distributed systems in production
-
Experience with Ray.io, particularly Ray Core and Ray Data (or equivalent technologies)
-
Experience with Kubernetes, particularly for heterogeneous workloads
-
Experience with cloud infrastructure on AWS or similar providers
-
Track record of shipping and operating reliable, highly available scalable infrastructure
-
Demonstrated ability to prioritize development work and build cross-functional consensus around technical tradeoffs
-
Proficiency with Python
Bonus Qualifications
-
Exposure to machine learning workloads (training, inference, data generation)
-
Experience with Kubernetes or SLURM at scale (>10k+ nodes)
-
Experience with SLURM workload manager and advanced scheduling policies
-
Background in algorithmic optimization or operations research
-
Experience building developer tools and platforms used by large engineering organizations
Listing read directly from Zoox's applicant tracking system. Check frequency varies by source. Listings are removed after successful checks confirm they are no longer present.