TPM Manager, Infrastructure
full-time
lead
Posted 20 hours ago
Before you apply
Build my evidence-backed draft — free Apply on company site →Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.
About this role
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
About the Role
Anthropic’s Infrastructure organization is responsible for the systems that train our models, serve our products, and support our engineering teams. That includes datacenter operations, capacity planning across cloud providers and our own facilities, accelerator cluster management, production serving infrastructure, developer tooling, data pipelines, and networking. It’s a lot of surface area, and the demands on it are growing fast.
We’re hiring a TPM leader to own program management across this whole ecosystem from compute provisioning through to production workloads. We’re bringing up new datacenters, scaling multi-cloud compute across AWS, GCP, and Azure, managing datacenter construction, and building the software infrastructure to keep pace. The team is growing very quickly, and we’re looking for a senior leader with experience at scale to build and scale this team to support Anthropic’s rapid growth. You’ll report to the Head of TPM, partnering closely with various engineering leaders on technical strategy, roadmapping, and aligning TPM support where it is most impactful.
You’ll personally drive 2–3 critical programs while leading your teamin parallel. This is a role where you need to be comfortable doing the work yourself before you can hand it off.
What you'll do:
IC Program Leadership
Own and drive 2–3 of the highest-priority programs across infrastructure while you build the team
Run the actual programs—datacenter bring-up timelines, capacity scaling plans, infrastructure migrations, cross-team reliability efforts, or whatever the most pressing needs are
Build the processes and playbooks as you go—figure out what works by doing it, then codify it for the team
Earn credibility with engineering leads through solid execution, not just strategy
Team Building & Development
Set the standard for what good TPM work looks like in this domain through your own output
Coach and develop TPMs
Transition programs to your team as you hire
Grow the team with excellent TPMs by defining roles, writing JDs, sourcing candidates, and closing hires
Planning & Prioritization
Work with various engineering leads to identify work that would most benefit from TPM support
Make real tradeoffs about what to staff vs. what to skip given limited TPM capacity during the build phase
Maintain portfolio-level visibility across programs—status, risks, dependencies, blockers
Represent the team in planning cycles and leadership reviews
Cross-functional Coordination
Coordinate across Infrastructure and partner teams (Research, Product, Security, Finance, Legal) on programs that span organizational boundaries
Drive alignment on programs that cross the hardware/software line—e.g., capacity plans that feed into training schedules, or efficiency work that spans accelerator kernels and serving systems
Own executive communication on program status, risks, and resource needs
You May Be a Good Fit If You:
Have 10+ years of experience in technical program management, with 7+ years directly managing TPMs and ideally some experience leading larger TPM organizations
Have built a team or function from scratch before—you know the difference between hiring for a defined role vs. figuring out what the roles should be
Have scaled TPM teams to support rapidly-growing, fast-moving company environments
Have worked across physical and software infrastructure—datacenters, networking, hardware ops, distributed systems, cloud platforms, developer tooling. You don’t need to be deep in all of it, but you need to be conversant enough to ask the right questions and spot the real risks.
Have run large-scale compute or infrastructure programs—capacity planning, cluster deployments, datacenter build-outs, cloud migrations, or similar
Can communicate complex programs clearly to senior leadership without losing the important details
Are good at context-switching between doing the work and managing people, and don’t see the IC work as beneath you
Are comfortable making staffing and prioritization decisions without perfect information
Deadline to apply: None, applications will be received on a rolling basis.
The annual compensation range for this role is listed below.
For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.
Annual Salary:
$365,000 — $565,000 USD
Logistics
Minimum education: Bachelor’s degree or an equ
Similar Jobs
Related searches:
Hybrid Jobs
Lead Jobs
Hybrid Lead Jobs
Lead Backend & SystemsLead AI Safety & SecurityLead AI ResearchLead Data EngineeringLead AI Infrastructure
AI Jobs in San Francisco
Backend & Systems in San FranciscoAI Safety & Security in San FranciscoAI Research in San FranciscoData Engineering in San FranciscoAI Infrastructure in San Francisco
distributed-systemsdata-pipelinealignmentinfrastructure
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.