Data/Infrastructure Advocate Engineer - EMEA Remote

Hugging Face · Paris, France

full-time mid Posted 1 month ago

Apply Now Stand out: build a proof-of-work pitch →

Free GitHub-based preview. Direct apply stays one click away.

Get weekly job alerts like this →

Hiring for this role?

AI Market Demand Pack · $29 one-time

Compare this role's skills with the full AI hiring market. Get ranked demand, salary bands, leading companies, public source URLs, and a decision brief.

See the live sample →

data-pipeline infrastructure

About this role

<p></p><p>At Hugging Face, we're on a journey to democratize good AI. We are building the fastest growing platform for AI builders with over 11 million users who collectively shared over 4 million models, 1 million datasets & 1.5 million Gradio apps. Our open-source libraries have more than 700,000 stars on Github.</p><h3>About the Role</h3><p>As our first Data/Infrastructure Advocate Engineer, you'll bridge the gap between cutting-edge data infrastructure and the global community of data engineers, researchers, and developers. You'll champion Xet storage on the Hugging Face Hub, helping users efficiently store, version, and collaborate on large-scale datasets. This role is for someone who thrives at the intersection of technical depth (storage, Parquet, deduplication) and community advocacy, helping define the future of open data workflows.</p><p>You'll collaborate with teams like Datasets, Hub, and Infrastructure to shape how developers interact with data on our platform, and inspire a community to build better, faster, and more scalable data pipelines.</p><h3>Your main missions</h3><ul><li>Grow and nurture the open-source data/infra community: launch initiatives, collaborate with data-focused groups, and organize events or challenges. Engage with communities like Apache Parquet, Open Table Formats, and data engineering forums to promote best practices and Hugging Face tools.</li><li>Promote the Hugging Face Hub as the go-to platform for data storage, versioning, and collaboration, curating and showcasing datasets, benchmarks, and tools like Xet.</li><li>Highlight use cases like efficient large-dataset updates, Parquet editing, and deduplication to demonstrate the Hub's value for data workflows.</li><li>Create demos, benchmarks, and tools (for example Colab notebooks) that illustrate best practices for data storage and versioning, and experiment with Xet, Parquet, and other formats.</li><li>Produce high-quality tutorials, blog posts, and videos that make complex topics accessible.</li><li>Share insights on storage optimization, dataset versioning, and deduplication to empower developers.</li><li>Actively participate in online communities (Discord, GitHub, forums) to highlight contributions, answer questions, and foster collaboration.</li><li>Make sure datasets and tools released on the Hub are well-documented, with clear examples, benchmarks, and use cases.</li></ul><h3>About You</h3><p>You're already an active voice in the data and ML community. You build in public, you publish, and people follow your work on LinkedIn and X.</p><p>You're a hands-on builder who loves experimenting with data tools, storage optimization, and dataset versioning. You can take a complex topic like deduplication, compression, or Parquet editing and make it click for other developers through writing, demos, or talks. You're passionate about open source and knowledge sharing, and you thrive in fast-moving environments.</p><h3>What you'll need</h3><ul><li>3+ years in developer relations or developer advocacy, ideally for data engineering, infrastructure, or ML tools and platforms</li><li>An established public presence as a technical voice, with a track record of regularly publishing data/infra/ML content and a demonstrable, engaged audience on LinkedIn and X (Twitter)</li><li>A portfolio of developer-facing content you can point to: tutorials, blog posts, videos, demos, benchmarks, or conference talks</li><li>Hands-on experience building and engaging open-source or developer communities (Discord, GitHub, forums)</li><li>Strong Python skills</li><li>Hands-on experience with data libraries such as pandas, pyarrow, and huggingface/datasets</li><li>Practical experience with storage systems and formats: Parquet, Open Table Formats, and S3</li><li>Working knowledge of dataset versioning, deduplication, and compression</li><li>Ability to explain complex technical topics clearly through writing, demos, or talks</li><li>Fluent written and spoken English</li></ul><h3>Nice to have</h3><ul><li>Experience with the Hugging Face Hub and datasets ecosystem, or with Xet</li><li>Open-source maintainer or contributor experience</li><li>Familiarity with large-scale data pipelines and data engineering workflows</li><li>Experience producing notebooks (for example Colab) for tutorials and benchmarks</li></ul><h3>A note on fit</h3><p>If you're interested in joining us but don't tick every box above, we still encourage you to apply. We're building a diverse team whose skills, experiences, and backgrounds complement one another, and we're happy to consider where you might make the biggest impact.</p><h3>One more thing</h3><p>At Hugging Face we believe great AI shouldn't require a massive cluster, we build for everyone, especially the GPU-poor. And because we read every application, here's a small sign that you read this one too: start your answer to the first application question with the words “<strong>GPU-poor and proud 🤗</strong>”. No trick,