HPC/ML Infrastructure Engineer
Spellbrush · San Francisco or Tokyo
About this role
[https://app.ashbyhq.com/api/images/user-content/89c1442c-bd9c-46be-a910-20b49b5d9ffc/7832bc0a-28e2-4339-a413-1e13b747c3b5/hpc-admin-wide.png]
We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world. You’ll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the anime models are training.
YOU MAY BE A GOOD FIT IF:
YOU LOVE ANIME AND THE ANIME AESTHETIC.
This probably one of the only jobs in the world where you will get to combine your love of anime and large-scale GPU systems.
YOU’RE FAMILIAR WITH THE MODERN HPC SOFTWARE LANDSCAPE
Once upon a time, our team could install SLURM on a few bare metal nodes and get away with it. Now the landscape has become unbelievable complex, with SLURM deploys through Slinky on K8s, provisioning through warewulf/MAAS/ansible, filesystems through WEKA/VAST/Ceph, VPN and access through tailscale, and monitoring via the Grafana/Prometheus stack.…
Summary from Spellbrush's official Ashby career feed — read the full description on the original posting ↗
The full job description lives on Spellbrush's official careers page. AI Stack Jobs links you straight to it — we never sit between you and the employer.
More AI/ML roles at Spellbrush
- 2025-09-02 San Francisco or Tokyo
- 2025-09-02 San Francisco or Tokyo
- 2024-02-07 Tokyo
- 2024-02-07 San Francisco