All jobs
Save

L3 Support Engineer

Nebius
Contract type
Ongoing
Work mode
On-site · On-site (see listing for address)
Experience
3 years

Job description

Key details

  • Lead root cause analysis beyond L2 depth (GPU failures, firmware issues, Linux-level faults, HW/SW interactions)
  • Detect recurring patterns across sites and convert findings into durable fixes
  • Own technical workstreams during high-severity incidents
  • Build evidence packs and drive escalations with ODM and R&D
  • Push for firmware, component, and platform-level resolutions
  • Track outcomes and ensure knowledge flows back to operations
  • Support validation and rollout of firmware updates (risk assessment, staging, rollback planning)
  • Help operationalize platform standards across datacenters
  • Create scalable runbooks, troubleshooting guides, and error catalogs
  • Turn investigations into playbooks that elevate L1/L2 teams
  • Travel to datacenters for complex troubleshooting, new platform readiness, or incident containment
  • Company mission
  • Nebius is leading a new era in cloud infrastructure for the global AI economy, building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment

Primary stack

Core technologies

Python (Programming Language)

Benefits

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture
  • Opportunity to work on impactful AI projects
  • International environment and talented teams

Requirements & details

  • Strong hands-on experience with datacenter servers and deep Linux troubleshooting
  • Ability to diagnose across hardware, BIOS/BMC firmware, and Linux (logs, drivers, storage basics, performance triage)
  • Structured incident response experience and clear communication under pressure
  • Experience driving evidence-based escalations with vendors/R&D
  • Fluent English (written and spoken)
  • Bonus: Strong familiarity with GPU server platforms and tooling (for example: nvidia-smi, dcgmi, Linux logs correlation)
  • Bonus: Experience with ipmitool and Redfish workflows, firmware lifecycle, and staged rollouts
  • Bonus: Scripting skills (bash and basic Python) for log collection, triage automation, and simple reliability analysis
  • Bonus: Exposure to OCP-based platforms and ODM manufacturing ecosystems
  • Bonus: Experience supporting enterprise bare metal customers under contractual SLAs
  • Work location: Béthune, France
  • Python, Linux, bash, nvidia-smi, dcgmi, ipmitool, Redfish
  • Python (Programming Language)

Apply