ML Researcher - Image / Video Diffusion

Krea · San Francisco, CA
full-time mid Posted 1 day ago

Before you apply

Build my evidence-backed draft — free Apply on company site →

Paste your relevant resume section or 2–4 true bullets. See supported requirements and honest gaps. No account and no application sent.

Get weekly job alerts like this →

About this role

ABOUT KREA At Krea, we are building next-generation AI creative tools. We're dedicated to making AI intuitive and controllable for creatives - our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats - text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium. We recently took this a step forward with the launch of Krea 2 https://www.krea.ai/blog/krea-2-image-model, our first foundation model, built completely from scratch for aesthetic diversity and stylistic control. We've raised over $83M and are backed by world-class investors such as a16z, Bain Capital, and Abstract. We work full-time and in-person at our waterfront office in San Francisco. We care about creativity: our team includes musicians, designers, visual artists, and engineers.   We're looking for an experienced Researcher with engineering skills who can work on large-scale image and video models training experiments, with experience training image models at scale. OUR CULTURE - We work full-time and in-person at our North Beach office in San Francisco. - We believe that demonstrated interest in the creative space is key: our team includes musicians, designers, visual artists and more. - Fast iteration and execution speed. Bias towards action, agency, and independence. WHAT YOU'LL DO - Train diffusion models for image and video generation on large GPU clusters. - Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication. - Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP. - Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design. - Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU, NVLink, Infiniband (IB) components as well as monitoring numerical errors and NCCL issues. - Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance of our models. WHAT WE'RE LOOKING FOR - Proven track record in working with image or video models at scale (publications or open-source contributions a plus). - Strong proficiency in PyTorch and understanding of its inner workings. - Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs. - Experience in profiling and debugging large distributed training. Being comfortable with analyzing traces to identify bottlenecks and look for improvements. - Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8. - Solid understanding of diffusion model training pipeline across pretraining, midtraining, preference optimization, and reinforcement learning. - Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research. - Being comfortable working in a goal-oriented research environment. - Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data. - Comfortable working with underspecified goals. We expect every technical member to take an ambiguous research goal and break it down into concrete requirements, plans, experiment plan, and execution items. - Good research taste — bias towards simplicity and methods that scale well with compute, data, and minimal human supervision. - Ability to iterate rapidly, and propose creative research directions. - Be comfortable getting your hands dirty with data and designing custom data pipelines to improve data quality. WHAT WE OFFER - Team: Work alongside a world-class team building the future of AI creative tooling - Impact: Significant scope and company-wide impact - Competitive compensation: generous salary & equity packages - Health & wellness: 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverage - Time off: Flexible PTO policy - Financial planning: 401k with a 4% company-sponsored match - Meals in the office: breakfast, lunch, dinner - you name it, we'll cover it - Transit: Ubers covered to & from the office - Sponsorship: We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3). - And more! Please note the above benefits & perks are for full-time employees

Similar Jobs

Related searches:

On-site Jobs Mid-Level Jobs On-site Mid-Level Jobs Mid-Level Computer VisionMid-Level AI InfrastructureMid-Level Robotics & AutonomyMid-Level Backend & SystemsMid-Level AI ResearchMid-Level Generative AIMid-Level Machine Learning AI Jobs in San Francisco Computer Vision in San FranciscoAI Infrastructure in San FranciscoRobotics & Autonomy in San FranciscoBackend & Systems in San FranciscoAI Research in San FranciscoGenerative AI in San FranciscoMachine Learning in San Francisco generative-aigpuroboticspytorchdistributed-systemsdiffusion-modelsreinforcement-learningpre-training

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.