Staff Software Engineer, HPC
Verified live
Visa & eligibility
Accepts F-1 students
not published
Supports CPT
not published
Supports OPT
not published
Supports STEM OPT
not published
E-Verify employer
not published
Sponsorship stance
Unknown / not mentioned
The job
Location
Foster City, CA
Work setting
Hybrid
Hours per week
not published
Schedule
not published
Experience required
not published
New graduates accepted
not published
Skills mentioned
Compensation & benefits
Pay range
$201,000–$315,000/yr
Relocation assistance
not published
Health insurance
not published
Applying
Application deadline
not published
Typical response time
not published
Number of applicants
not published
Posting verified
Verified live
Last checked
Sept. 23, 2026, 3:45 a.m.
Applies directly to employer
Yes
Apply on the employer's site →
Employer
Company
Zoox
Industry
Technology / Software
Company size
not published
Contact
not published
Contact email
not published
Source
Zoox
Full description
Zoox is looking for an experienced Staff Software Engineer to build, scale, and operate our custom High-Performance Computing infrastructure. As Zoox scales its autonomous vehicle development, our HPC platform must keep pace with rapidly growing compute, storage, and scheduling demands across the company. You will modernize our HPC platform—built on industry-leading technologies like Ray.io, SLURM, and Kubernetes—with a focus on reliability, scalability, and world-class developer velocity.
These HPC services form the backbone of development workflows across all Zoox software teams, from data engineering to training our AI models in Perception, Planner, Prediction, to Simulation, and more. You will have a direct impact on the productivity and effectiveness of every engineering team at Zoox.
The position comes with a high degree of independence and the opportunity to define Zoox's HPC platform strategy, both technically and organizationally. You will work closely with stakeholders in Autonomy and Software teams to understand their workload requirements and translate them into robust, scalable infrastructure.
In this role, you will:
Design and implement core services and abstractions for distributed compute infrastructure supporting hundreds of thousands of concurrent jobs
Work with customer teams and other infrastructure teams to build a multiyear software engineering roadmap for the HPC platform
Lead multi-quarter, cross team initiatives that drive org-wide improvements
Create production-grade APIs, SDKs, and tools that make it easy for engineers across Zoox to run large-scale distributed workloads
Design and improve job scheduling algorithms and auto-scaling policies to maximize reliability and resource availability
Design multi-region orchestration strategies that optimize for data locality, reliability, and performance
Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners across multiple teams
Evaluate new technologies and paradigms that improve Zoox's computational and storage capabilities
Develop capacity planning tools and forecasting models to support Zoox's growing compute needs
Mentor junior engineers, guiding them through their career development
Qualifications
Experience designing and operating large-scale distributed systems in production
Experience with Ray.io, particularly Ray Core and Ray Data (or equivalent technologies)
Experience with Kubernetes, particularly for heterogeneous workloads
Experience with cloud infrastructure on AWS or similar providers
Track record of shipping and operating reliable, highly available scalable infrastructure
Demonstrated ability to prioritize development work and build cross-functional consensus around technical tradeoffs
Proficiency with Python
Bonus Qualifications
Exposure to machine learning workloads (training, inference, data generation)
Experience with Kubernetes or SLURM at scale (>10k+ nodes)
Experience with SLURM workload manager and advanced scheduling policies
Background in algorithmic optimization or operations research
Experience building developer tools and platforms used by large engineering organizations
About Zoox
Zoox is developing the first ground-up, fully autonomous vehicle fleet and the supporting ecosystem required to bring this technology to market. Sitting at the intersection of robotics, machine learning, and design, Zoox aims to provide the next generation of mobility-as-a-service in urban environments. We’re looking for top talent that shares our passion and wants to be part of a fast-moving and highly execution-oriented team.
Follow us on LinkedIn
A Final Note: You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills.
Eligibility signals are matched from the employer's own wording and shown with the source text so you can check them. A blank field means the posting did not say — not that the answer is no.
This is informational only and is not immigration advice. Confirm your work authorization with your DSO before accepting any role.