Staff Software Engineer, Production Engineering

Harvey AI · San Francisco, CA · $231k - $340k
full-time lead Posted 15 hours ago
Apply Now Stand out: build a proof-of-work pitch →

Free GitHub-based preview. Direct apply stays one click away.

Get weekly job alerts like this →

Hiring for this role?

AI Market Demand Pack · $29 one-time

Compare this role's skills with the full AI hiring market. Get ranked demand, salary bands, leading companies, public source URLs, and a decision brief.

See the live sample →

About this role

WHY HARVEY At Harvey, we’re transforming how legal and professional services operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work gets done for decades to come. This is a rare chance to help build a generational company at a true inflection point. We have strong product-market fit and world-class investor support. We’re scaling fast and defining a new category in real time. The work is ambitious, the bar is high, and the opportunity for growth — personal, professional, and financial — is unmatched. Our team moves fast, takes ownership, and is deeply committed to the mission — operating with intensity, staying close to our customers, and pushing each other for excellence. We live by three values: Decisiveness, Simplicity, and Job's Not Finished. We act quickly on clear judgment over perfect information, we believe simplicity is what scales, and we're never satisfied with where we are. If you want to do the best work of your career alongside people who share that drive, we'd love to build with you. At Harvey, the future of professional services is being written today — and we’re just getting started. ROLE OVERVIEW Harvey is building the AI platform trusted by the world’s leading law firms and enterprises. Our infrastructure is the foundation that powers every customer interaction, every model inference, and every production workload. We’re looking for a Production Engineer to help build and operate Harvey’s core compute and networking infrastructure, Kubernetes platform, workflow orchestration platform, and production infrastructure foundations. You’ll work on the systems that enable engineering teams to move quickly and operate reliable services at scale. In this role, you’ll improve the reliability, scalability, security, and efficiency of Harvey’s infrastructure platform. You’ll solve complex production challenges across compute fleet management, capacity planning, infrastructure automation, and production operations. You’ll partner closely with Product Engineering, Security, AI Infrastructure, and Platform teams to ensure our infrastructure scales with Harvey’s rapid growth. At Harvey, we value Decisiveness, Simplicity, and the belief that Job’s Not Finished. We move quickly, prioritize clarity, and continuously raise the bar for engineering excellence. WHAT YOU'LL DO INFRASTRUCTURE ENGINEERING & TECHNICAL LEADERSHIP - Design, build, and operate the production infrastructure that powers Harvey’s products and AI workloads. - Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations. - Lead complex, cross-functional technical initiatives that improve reliability, scalability, security, operational efficiency, and infrastructure cost. - Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate product and business requirements into resilient infrastructure solutions. - Establish reusable patterns, tooling, and paved paths that help engineering teams ship and operate production services safely. - Raise the engineering bar through thoughtful design reviews, clear technical documentation, operational rigor, and mentorship. INFRASTRUCTURE FOUNDATION & PRODUCTION OPERATIONS - Build and operate Harvey’s global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance. - Improve compute utilization, performance, and service availability while supporting rapidly growing AI workloads. - Develop capacity models, demand forecasts, and fleet lifecycle automation to help infrastructure scale efficiently with business growth. - Operate and continuously improve Harvey’s Kubernetes platform, including cluster provisioning, upgrades, networking, monitoring, reliability, performance, and operational automation. - Drive infrastructure cost efficiency through capacity management, resource rightsizing, workload optimization, and utilization monitoring. - Build secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls. - Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi. - Improve observability, monitoring, alerting, incident response, and operational readiness across the infrastructure platform. - Participate in the on-call rotation, lead incident response when needed, and turn production learnings into durable engineering improvements. WHAT YOU HAVE - 10+ years of software, infrastructure, site reliability, or production engineering experience. - Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform. - Strong hands-on experience operating Kubernetes in production, including

Similar Jobs

Related searches:

Hybrid Jobs Lead Jobs Hybrid Lead Jobs Lead AI InfrastructureLead Machine LearningLead Backend & SystemsLead NLP & Language AILead AI Agents & RAG AI Jobs in San Francisco AI Infrastructure in San FranciscoMachine Learning in San FranciscoBackend & Systems in San FranciscoNLP & Language AI in San FranciscoAI Agents & RAG in San Francisco agentsdistributed-systemsllmcloud

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.