Platform Engineer (Agent Runtime)
full-time
senior
Posted 8 hours ago
Apply Now
Stand out: build a proof-of-work pitch →
Free GitHub-based preview. Direct apply stays one click away.
Get weekly job alerts like this →Hiring for this role?
AI Market Demand Pack · $29 one-time
Compare this role's skills with the full AI hiring market. Get ranked demand, salary bands, leading companies, public source URLs, and a decision brief.
About this role
About Netskope
Today, there's more data and users outside the enterprise than inside, causing the network perimeter as we know it to dissolve. We realized a new perimeter was needed, one that is built in the cloud and follows and protects data wherever it goes, so we started Netskope to redefine Cloud, Network and Data Security.
Since 2012, we have built the market-leading cloud security company and an award-winning culture powered by hundreds of employees spread across offices in Santa Clara, St. Louis, Bangalore, London, Paris, Melbourne, Taipei, and Tokyo. Our core values are openness, honesty, and transparency, and we purposely developed our open desk layouts and large meeting spaces to support and promote partnerships, collaboration, and teamwork. From catered lunches and office celebrations to employee recognition events and social professional groups such as the Awesome Women of Netskope (AWON), we strive to keep work fun, supportive and interactive. Visit us at Netskope Careers. Please follow us on LinkedIn and Twitter @Netskope .
Every agent on this platform runs inside its own isolated environment, and someone has to make sure that environment actually behaves the way it's supposed to under real load, not just in a clean test run. As a Senior Agent Runtime Engineer , you'll own the health of the systems agents run on: keeping sessions isolated from each other, catching the failure modes that only show up under real concurrency, and making sure a slow or expensive agent gets caught before it becomes everyone's problem. You'll work closely with the engineers building the agents themselves, since you're often the first call when something behaves strangely in a real environment but not in a test one. If you like being the person who understands a system well enough to know exactly why it's misbehaving, this is that job.
Skills and competencies:
Own the health of the agent execution environment day to day — session isolation, resource limits, and catching the failure modes that only appear once real concurrency and real traffic patterns show up, not just in a clean test run.
Investigate and resolve cases where an agent behaves differently in production than it did in testing, working directly with the Agent Engineer or Quality Engineer who built it to figure out whether the problem is the agent's design or the environment it's running in.
Tune cold-start and concurrency settings for the platform's critical-path functions, and review them on a regular cadence as usage patterns shift rather than setting them once and forgetting them.
Understand the platform's session isolation model well enough to reason about its limits — where isolation is guaranteed by the underlying compute layer, and where the platform has to add its own controls because that guarantee doesn't fully hold.
Build and maintain the observability that lets someone answer "is this agent actually working correctly," not just "is it technically up" — tracing, behavioral drift detection, and quality signals sitting alongside the usual metrics and logs.
Track per-agent and per-session cost and efficiency, and flag agents that are burning more tokens, calling more tools, or running longer than the task should reasonably require.
Deploy and manage runtime-layer infrastructure resources through existing CI/CD pipelines as needed — provisioned concurrency settings, runtime-specific IAM roles, observability configurations, etc.
Run and improve the tests that validate isolation actually holds — for example, confirming that one agent's session genuinely can't reach or affect another's, not just assuming it because the platform is supposed to guarantee it.
Must-Have:
At least 5 years in a platform, SRE, or infrastructure engineering role, with real production experience running serverless or containerized workloads on AWS (Lambda, Fargate, Docker, Kubernetes or equivalent) at meaningful scale.
Hands-on experience with AWS observability tooling (CloudWatch, X-Ray, or a comparable distributed tracing stack) — able to go from "something's wrong" to a root cause using traces and logs, not just dashboards.
Real experience deploying and managing infrastructure through CI/CD independently — comfortable owning IaC changes (Terraform, CDK, or similar) end to end rather than handing them off to someone else.
Working understanding of compute isolation concepts (containers, microVMs, or similar sandboxing models) and where their guarantees actually stop, since a lot of this role is reasoning about the edges of what a platform promises versus what it might not fully cover on its own.
Comfort investigating cost and performance problems at a granular level — able to trace an unexpectedly expensive or slow workload back to a specific cause, not just flag that costs went up.
Strong incident response instincts: staying calm and methodical while root-causing a live production issue, and following
Similar Jobs
Related searches:
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.