Forward Deployed Engineer - SRE
full-time
mid
Posted 3 weeks ago
Before you apply
Build my evidence-backed draft — free Apply on company site →Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.
About this role
FORWARD DEPLOYED ENGINEER - SRE
LOCATION: NORTH AMERICA REMOTE/SF-HYBRID · FULL-TIME
ABOUT ANDROMEDA
Andromeda is a market and infrastructure platform to buy, sell, and operate compute.
We believe demand for compute will grow exponentially. So fast that a handful of vertically integrated providers won't be able to scale across operations, capital, supply chains, and politics to serve it. The result is a massive wave of fragmentation, with AI factories of every shape and size coming to market to fill this demand. Our job is to enable all of that fragmented compute to flow through one platform, delivering reliable capacity to model builders, research labs, and inference providers when they need it. We believe every spare electron should be made productive for AI and we're building the platform that makes that possible.
We sit at the center of three forces:
- Companies that need reliable, high-performance compute fast
- A fragmented global supply of GPUs across hyperscalers, neoclouds, and independent data centers
- Capital, risk, and operational complexity that most teams are not equipped to manage
When we succeed, trillions of dollars of compute will flow through Andromeda. Builders get capacity when they need it. Providers get a reliable way to monetize, operate, and finance infrastructure at scale. Capital gets an easy way to deploy, hedge, and underwrite.
In five years, Andromeda won't just participate in the AI infrastructure market. We will shape it.
The Role
This is not a generalist SRE role, and it is not a support role. You will embed directly with the teams running large-scale training and inference on our clusters. You are responsible for onboarding them, tuning their jobs, and debugging their failures alongside them, while owning the infrastructure and automation that makes those clusters reliable in the first place.
Forward deployed means you spend real time inside customer environments: reading their training code, sitting in their Slack channels, watching their runs, and shipping fixes that land in our platform. When a multi-hundred-GPU run stalls, you are the person who figures out whether it's the fabric, the driver, the scheduler, or their dataloader, then you make sure it can't happen the same way twice.
We're looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework. Equally important: you can explain what you found to someone else's engineering team without condescension, and turn that conversation into a product improvement.
What You’ll Do
- Serve as the primary technical point of contact for teams running large-scale training and inference workloads. Own onboarding end to end; environment setup, orchestration choice (Slurm, Kubernetes, or direct SSH), storage layout, first successful run at scale. You will continue to stay engaged as their workloads grow.
- Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns, container and driver mismatches. Read their code when you need to. Reproduce, isolate, fix, and write it down.
- Profile and improve distributed training performance on live workloads. Improving MFU, cutting idle GPU time, and reducing time-to-first-successful-run for new deployments.
- Own reliability outcomes for the accounts you're deployed on.
- Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training. Diagnose and resolve fabric-level issues that degrade collective operations.
- Build deep visibility into GPU utilization, memory pressure, interconnect throughput, job performance, and hardware health.
- Turn every repeated deployment problem into automation: cluster provisioning, GPU health checks and burn-in, preflight validation, self-healing, firmware/driver lifecycle management, and reusable reference configurations for common training and serving stacks.
- Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks. Own the customer-facing communication during the incident and the blameless postmortem and systemic fix after it.
- You will see our rough edges before anyone else does. Bring that signal back to influence the roadmap, file the hard bugs, and build the missing pieces yourself when that's the fastest path.
What We’re Looking For
- Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent). You understand GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes from direct experience.
- Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training. You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree top
Similar Jobs
Related searches:
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.