QA Engineer

Nebius · Taiwan
full-time senior Posted 1 day ago
Apply Now Stand out: build a proof-of-work pitch →

Free GitHub-based preview. Direct apply stays one click away.

Get weekly job alerts like this →

Hiring for this role?

AI Market Demand Pack · $29 one-time

Compare this role's skills with the full AI hiring market. Get ranked demand, salary bands, leading companies, public source URLs, and a decision brief.

See the live sample →

About this role

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. About the Role We are looking for a technically strong, hands-on QA Engineer to join our hardware team on-site at ODM factories in Taiwan. This is not a checklist job - we're looking for someone who enjoys digging deep into technical issues, investigating root causes, and taking ownership of complex hardware problems.You'll be the key person ensuring the quality of our servers and racks before they ship, but more importantly, you'll play a critical role in debugging failures, analyzing test data, and working closely with RnD, logistics, and factory teams to continuously improve the process and the product.This is a deeply technical role that blends hardware validation, manufacturing QA, and problem-solving - perfect for someone who understands how servers are built and tested, and wants to make sure every unit that leaves the factory is production-grade. What You'll Own Technical Investigation & Debugging Investigate complex problems (e.g., high GPU failure rate, power-related test failures), gather logs, run diagnostics, and escalate with context to RnD when needed. Drive root cause analysis across factory teams and internal engineering groups. Document findings and help define preventive actions for recurring problems. Act as the first line of technical escalation for hardware issues discovered during factory QA or internal testing.Engineering Support Participate in new platform bring-up sessions together with the visiting RnD teams during on-site trips to ODM labs. Provide technical support, coordination, and hands-on assistance during the bring-up process. Help ensure early-stage hardware behaves as expected, and escalate integration or platform issues to the relevant teams.On-Site Product QA Perform visual inspections of completed products (servers, racks) before packaging. Define and maintain QA checklists and inspection procedures tailored to different product lines. Verify inventory records at the factory against internal system data (part numbers, serials, configurations). Oversee the product packaging process for compliance with defined standards. Supervise pickup operations: ensure outbound trucks meet shipment conditions and schedules.Failure Rate Monitoring & Analytics Collect failure data from vendor-side burn-in and our own test systems. Analyze failure trends and estimate spare part needs for future datacenter deployments. Use dashboards and structured reporting to communicate insights with QA, engineering, and supply chain teams.Feedback Loop & Quality Improvement Gather and process feedback from datacenters on each delivered batch of equipment:     * Report on packaging issues, impact sensor triggers, shipping anomalies.     * Assess rack-level build quality: cabling, bracket alignment, labeling.     * Log systemic hardware issues (design flaws, infant mortality, recurring failures). Forward the feedback to the teams: logistics, ODM partners, hardware RnD, QA.Test Infrastructure & Validation Assist with deployment and maintenance of test infrastructure on-site. Ensure Nebius post-manufacturing hardware validation tests run smoothly (uptime, monitoring, coordination with support team). Coordinate real-time issue escalation and basic triage with factory and internal teams.Local Insight & Communication Communicate relevant local risks and context (e.g., typhoons, holidays, factory-specific constraints) to our global logistics and hardware teams. Maintain productive relationships with factory staff, logistics providers, and internal stakeholders. Working Conditions & Tools During production peaks, issues may arise that require fast, hands-on debugging and resolution on-site. Flexibility is expected: you may need to stay late to investigate failures in freshly built batches or arrive early to verify and unblock outbound truck shipments. Rapid response and clear communication with engineering and factory teams are critical during these high-pressure periods. Occasional international travel may be expected to Nebius headquarters in Amsterdam or to datacenters in Europe and the US. Daily work tools involve:   * Managing workflows and escalation via Jira   * Writing and maintaining technical documentation in C

Similar Jobs

Related searches:

On-site Jobs Senior Jobs On-site Senior Jobs Senior AI Infrastructure cloudinfrastructure

Get jobs like this delivered weekly

Free AI jobs newsletter. No spam.