Infrastructure and MLOps Engineer

Graphcore · Gdańsk, Pomeranian Voivodeship, Poland
full-time mid Posted 4 days ago
Apply Now Stand out: build a proof-of-work pitch →

Free GitHub-based preview. Direct apply stays one click away.

Get weekly job alerts like this →

Hiring for this role?

About this role

About Graphcore  At Graphcore, we’re building the future of AI compute. We’re a team of semiconductor, software and AI experts, with deep experience in creating the complete AI compute stack - from silicon and software to infrastructure at datacenter scale. As part of the SoftBank Group, backed by significant long-term investment, we are delivering key technology into the fast-growing SoftBank AI ecosystem.To meet the vast and exciting AI opportunity, Graphcore is expanding its teams around the world.We are bringing together the brightest minds to solve the toughest problems, in a place where everyone has the opportunity to make an impact on the company, our products and the future of artificial intelligence . Job Summary  Join our dynamic Software Infrastructure team and take a pivotal role in scaling and managing our infrastructure. You will develop essential tools and services that empower our broader software team. Your contributions will enhance the build, test, deployment, and productisation processes of our Machine Learning Software components. Work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems The Team The Software Infrastructure team provides critical platforms and services for software development teams across the business. Our responsibilities include managing the CI platform and services, build engineering, component integration, and packaging and release systems. We operate in squads, fostering a culture of service ownership and empowerment for our engineers. We focus on long-term engineering solutions and strive to eliminate toil wherever possible.   Responsibilities and Duties   Develop, own, and maintain tools and services to support AI research and engineering teams   Deploy and maintain services with Kubernetes and Docker   Manage our Cloud Infrastructure using tools such as Terraform   Candidate Profile    Essential:   Knowledge of Python   Familiarity with cloud services (e.g. AWS)   Experience managing or developing in Linux environments   Understanding of CI/CD principles   Experience using Kubernetes (k8s)   Experience of one of the following: maintaining machine learning applications. deploying ML orchestration tools (e.g. NV Ray, KFP, SkyPilot). managing ML accelerator hardware (e.g. DCGM). Desirable   Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu)   Experience with GitHub Actions   Experience with modern observability tooling (e.g. Prometheus)   Experience with Grafana   Knowledge of Go/Java/C++ (or similar language) Benefits In addition to a competitive salary, Graphcore offers annual leave policy, medical and dental health plans, a gym card, and employee pension (matched up to 4%). We review our benefits on a yearly basis to ensure we offer a valuable and rewarding benefits programme to our employees. We welcome people of different backgrounds and experiences; we’re committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments

Similar Jobs

Related searches:

On-site Jobs Mid-Level Jobs On-site Mid-Level Jobs Mid-Level AI InfrastructureMid-Level Backend & Systems distributed-systemscloudmlopsinfrastructure

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.