Inference Systems Performance Architect

SambaNova Systems · San Jose, CA · $245k - $325k
full-time principal Posted 1 week ago

Before you apply

Build my evidence-backed draft — free Apply on company site →

Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.

Get weekly job alerts like this →

About this role

The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale. SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets. About the role As an Architect on the Inference Systems Performance team, you'll own the discipline of end-to-end performance for large-scale LLM inference at SambaNova, from how a request moves through tokenization, prefill, decode, and the fabric between them, to how an entire deployment is sized against customer SLOs. Inference-systems performance is a nascent field; the results of design choices are being discovered daily rather than inherited from a mature craft, and this role exists to bring rigor to that frontier. The work spans two coupled pillars. The first is reproducible workload capture and benchmarking -- building faithful, replayable representations of real and increasingly agentic traffic, so that what we measure reflects production rather than an artifact of a naive load script.  The second is performance modeling and simulation - analytic and simulation models that turn measurement into a "what-if" capability, letting us reason about configurations and hardware that do not exist yet. Together these feed both today's serving optimization and the next generation of system planning. The technical frontier you'll help define is heterogeneous, disaggregated inference - GPU on prefill, the RDU on decode - which explores hard problems across networking, storage, prompt caching, and tail-latency-bound data movement. You will be the go-to person for inference-systems performance across SambaNova, and a resource the entire organization relies on to answer "how fast can this go, and what will it take." Responsibilities Define and drive the technical strategy for inference-systems performance including workload capture, benchmarking, modeling, and simulation, while developing and architecture that enables many potential futures Build the workload-capture and agentic-benchmarking capability - capture representative production traffic and enforce the discipline of interrogating results, spotting artificial contention or misleadingly high cache-hit rates that never occur in real use Own the performance-modeling and simulation practice - models that predict how a configuration change moves the output, informing capacity planning against customer SLOs and next-generation system and hardware planning Attack the end-to-end profiling gap - drive tooling that produces accurate, actionable profiles of a distributed inference pipeline so bottlenecks can be localized across host, accelerator, and fabric Serve as the senior technical voice across model-optimization, systems, hardware, and product, tying together multiple engineering activities and teams, and weighing trade-offs of reliability, scalability, operational cost, and ease of adoption Act as a resource for the entire organization including representing SambaNova's performance story to customers and partners Mentor and multiply by raising the capability of principal and senior engineers, building the systems, tools, and patterns that make everyone more productive Drive the resolution of the most ambiguous, novel challenges that span organizational boundaries or have no established answer in the field yet Required qualifications 12+ years of experience in performance engineering, with a demonstrated record of technical leadership on large-scale, complex systems Deep expertise in end-to-end performance analysis of distributed systems with many moving parts and the ability to localize bottlenecks that others cannot Proven command of realistic workload generation and simulation and of performance modeling, including calibrating models against real, variable workloads Demonstrated ability to enter an unfamiliar domain and apply core performance methods with transferable discipline expertise  Ability to lead cross-functional efforts, mentor senior engineers, and influence organizational direction Experience representing an organizations credibly to customers and partners Track record of independently scoping and delivering high-complexity, high-ambiguity work with significant impact on products or roadmap Preferred qualifications Direct experience with LLM inference serving - continuous batching, prompt/KV cach

Similar Jobs

Related searches:

On-site Jobs Principal Jobs On-site Principal Jobs Principal Machine LearningPrincipal AI InfrastructurePrincipal AI Agents & RAGPrincipal Backend & SystemsPrincipal NLP & Language AIPrincipal Generative AI AI Jobs in San Jose Machine Learning in San JoseAI Infrastructure in San JoseAI Agents & RAG in San JoseBackend & Systems in San JoseNLP & Language AI in San JoseGenerative AI in San Jose llmagentsdistributed-systemsgenerative-aiinference

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.