Member of the Technical Staff - Platform

Andromeda · San Francisco, CA
full-time lead Posted 17 hours ago
Apply Now Stand out: build a proof-of-work pitch →

Free GitHub-based preview. Direct apply stays one click away.

Get weekly job alerts like this →

Hiring for this role?

AI Market Demand Pack · $29 one-time

Compare this role's skills with the full AI hiring market. Get ranked demand, salary bands, leading companies, public source URLs, and a decision brief.

See the live sample →

About this role

MEMBER OF THE TECHNICAL STAFF, PLATFORM Location: North America Remote / San Francisco, CA · Full-Time ABOUT ANDROMEDA Compute is the most sought-after resource in the world, yet it still trades like commercial real estate: year-long contracts, manual fulfillment, capacity sitting idle because nobody can move it. We are building the liquidity layer at Andromeda. Andromeda was founded by Nat Friedman and Daniel Gross to give startups the scaled AI infrastructure once reserved for hyperscalers. The first cluster filled almost instantly. The years since went into the platform that makes compute liquid: it deploys into foreign datacenters and turns the hardware it finds into clusters that leading AI labs train on. Today Andromeda operates compute for 80+ customers across 30+ capacity providers, with tens of thousands of GPUs under management and billions of GPU-hours supported, on everything from A100 to GB300. THE ROLE We are looking for engineers to build and operate the control plane that runs our fleet. Your responsibilities will include: - The automated systems that take a cluster from bare machines to customer-ready - Machine lifecycle between tenants: join, wipe, verify, rejoin - Operating Kubernetes and Postgres across the fleet - Contributing to our custom Kubernetes operators - Scaling clusters from tens of nodes to thousands - Participating in on-call rotations REQUIREMENTS - Impressive technical work you can go deep on, with impact in the world. That can take three years or twenty. - 2+ years of on-call experience for critical production services - Deep Kubernetes experience - Strong Linux fundamentals: kernel, cgroups, containers, networking, storage - Experience operating databases, monitoring, CI/CD, and cloud infrastructure at scale - Familiarity with fleet management and capacity planning Prior experience writing operators, or with GPUs, HPC scheduling, or bare-metal hardware is nice to have, but not required. Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Similar Jobs

Related searches:

Remote Jobs Lead Jobs Remote Lead Jobs Lead NLP & Language AILead AI InfrastructureLead Machine Learning AI Jobs in San Francisco NLP & Language AI in San FranciscoAI Infrastructure in San FranciscoMachine Learning in San Francisco cloudllmplatform

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.