Data Engineer (Agentic Search)

Nebius · Israel
full-time senior Posted 10 hours ago

Before you apply

Build my evidence-backed draft — free Apply on company site →

Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.

Get weekly job alerts like this →

About this role

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Product In a rapidly evolving world, trust in AI depends on AI agents being grounded in fresh, verified real-world data. Search is the foundation that makes this possible. We are building an agent-native search platform designed specifically for AI systems rather than human users. Our product provides programmatic, low-latency, and observable search APIs that AI agents use to retrieve, filter, and reason over real-world information at scale. Behind every search request is a rich stream of signals — query patterns, retrieval decisions, crawling outcomes, ranking quality, usage, and revenue events. Turning that stream into a trustworthy, queryable data platform is what makes the product improvable, the business measurable, and the models trainable.   The Role We are looking for a  Data Engineer to help build and scale the data platform behind our search quality, ML pipelines, product analytics, and business operations. In this role, you will contribute across the full data lifecycle: ingesting data from production systems, designing and evolving our data warehouse, building batch and streaming pipelines, and making high-quality datasets available to researchers, engineers, analysts, and product teams across the company. The platform spans tens of terabytes and ingests data from tens of proprietary and third-party sources — including our search engine and its components, CRM, billing, identity, and product analytics across multi-region production environments. Around 100 internal users rely on it daily. You will work closely with engineers and stakeholders across the company, contribute to architectural and modeling decisions, and help improve the reliability, usability, and scalability of the data platform as it grows. In this position, your responsibility will be to: Contribute to the design, development, and operation of Tavily's data platform — from real-time ingestion through data warehouse medallion layers to consumer-facing datasets and dashboards. Build and maintain reliable batch and streaming pipelines that ingest data from production services and external systems. Design and evolve scalable, analytics-ready data models in the data warehouse. Work closely with engineers across the company to ensure data produced by production systems is reliable, well-structured, and usable downstream. Improve observability across the data platform, including data quality checks, freshness monitoring, lineage, schema evolution, and cost controls. Partner with researchers, engineers, analysts, finance, and product managers to deliver trustworthy datasets for product, search quality, ML, and GTM analytics. Contribute to defining the objects, entities, and relationships that represent Tavily's search domain — including agent inputs, URLs, chunks, agent sessions, crawls, and the connections between them — and translate them into clean, queryable data models. Improve engineering practices around testing, documentation, deployment, and incident response. Investigate and resolve production data issues, including broken pipelines, corrupted datasets, schema changes, and large-scale backfills. Contribute to technical standards and best practices for data engineering across the company. Help maintain high standards of data quality, integrity, security, and governance across environments. You may be a good fit if you: Have 5+ years of Data Engineering experience, with strong experience designing and implementing scalable, analytics-ready data models and cloud data warehouses such as Snowflake or BigQuery. Have hands-on experience with Snowflake, or a comparable cloud data warehouse, and a strong understanding of modern data warehouse architecture, preferably including medallion-style modeling. Have deep knowledge of databases, including schema design, query optimization, and familiarity with NoSQL use cases. Have strong experience with modern data orchestration and transformation frameworks such as Airflow and dbt. Understand cloud data services on AWS or GCP and have experience with streaming platforms such as Kafka or Pub/Sub. Have hands-on experience with Spark, MapReduce, or similar distributed processing systems, and understand whe

Similar Jobs

Related searches:

On-site Jobs Senior Jobs On-site Senior Jobs Senior AI InfrastructureSenior AI Agents & RAGSenior Fintech & Payments AISenior Data Engineering searchdata-pipelineagentscloudpaymentsdata-engineering

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.