Software Engineer — Platform
full-time
lead
Posted 5 days ago
Before you apply
Build my evidence-backed draft — free Apply on company site →Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.
About this role
About Snorkel
At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data.
We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!
About the Team
Snorkel's Platform organization owns the infrastructure and services that power everything at Snorkel — the pipelines, evals, access layers, event systems, governance, compute, and agent infrastructure that every product team and customer deployment depends on. We're a small team with a large surface area, and we're in the middle of a foundational architecture shift: moving from a single-database data path to a multi-source, event-driven agent first platform. The decisions being made now will define how our platform scales for years. You will be making them.
About the Role
We're looking for Platform Engineers who combine strong infrastructure chops with real backend engineering depth. You'll build and operate the systems, services, and agent infrastructure that let product teams move fast and reliably — from data access layers and event-driven pipelines to the agent first architecture that will transform our ability to scale. Your work will directly allow us to scale the amount of high quality data we are able to deliver to our customers.
We are looking to grow our team of Platform Engineers, and are hiring at multiple levels.
What You'll Do
Design and build agent infrastructure that allows us to safely and reliably speed up workflows from engineering and operation teams
Design and implement event-driven data flows using event brokers, CDC connectors, schema registries, event routing, and dead letter queues — ensuring events flow reliably and failures are visible and recoverable
Build the systems that track how data moves through the platform (lineage), enforce who can access what (governance and RBAC), and log what happened (auditing), including PII handling, retention policy enforcement, and audit infrastructure for enterprise and regulatory compliance
Set strategy and architecture for build systems, testing frameworks, and CI/CD pipelines, and drive the transition toward robust, automated continuous deployment
Instrument services with OpenTelemetry, define and monitor SLOs (query latency, pipeline success rates, service reliability), and build alerting that catches issues before they become incidents — you will be on-call for the systems you build
Contribute to infrastructure cost visibility and optimization — query cost estimation, workload right-sizing, and routing data to the most cost-effective storage tier for its access pattern
Collaborate with engineers, product managers, and designers to bring consistency and high standards to codebases, infrastructure, and processes
What You'll Bring
5+ years building platform infrastructure, backend services, or data systems in production — you have built and operated pipelines, data access layers, distributed services, or ETL/ELT systems at scale
Strong proficiency in Python, and experience designing REST APIs for internal services and developers
Strong background in distributed systems and cloud platforms (AWS preferred) — hands-on experience with services like S3, RDS, EKS, EventBridge, and IAM, and comfort working in a Terraform-managed environment
Familiarity with data orchestration tools (Prefect, Airflow, or Dagster) and transformation frameworks (dbt)
Understanding of data governance concepts — RBAC, PII handling, audit logging, data lineage
Track record of leading complex engineering initiatives, influencing stakeholders, and delivering measurable impact
Ability to work in a fast-paced environment with strong technical communication skills
Fluency with modern developer tooling and a willingness to adopt new tools quickly — the team evaluates and integrates new tooling regularly to improve velocity and reliability
Nice to Have
Experience building shared libraries or SDKs consumed by multiple teams — versioning, backwards compatibility, migration support
Experience with event-driven architectures — CDC, event buses, schema registries, at-least-once delivery semantics
Experience with OpenTelemetry, ClickHouse, or similar observability infrastructure
Prior work in regulated environments (SOC 2, FedRAMP, HIPAA) where compliance requirements shaped system design
Exp
Similar Jobs
Related searches:
On-site Jobs
Lead Jobs
On-site Lead Jobs
Lead Data EngineeringLead Generative AILead AI InfrastructureLead Backend & Systems
AI Jobs in New York
Data Engineering in New YorkGenerative AI in New YorkAI Infrastructure in New YorkBackend & Systems in New York
data-pipelineapi-designgenerative-aiclouddistributed-systems
Get jobs like this delivered weekly
Free AI jobs newsletter. No spam.