N
L3 Support Engineer
Nebius
Contract type
Ongoing
Work mode
On-site · On-site (see listing for address)
Experience
3 years
Job description
Key details
- Lead root cause analysis beyond L2 depth (GPU failures, firmware issues, Linux-level faults, HW/SW interactions)
- Detect recurring patterns across sites and convert findings into durable fixes
- Own technical workstreams during high-severity incidents
- Build evidence packs and drive escalations with ODM and R&D
- Push for firmware, component, and platform-level resolutions
- Track outcomes and ensure knowledge flows back to operations
- Support validation and rollout of firmware updates (risk assessment, staging, rollback planning)
- Help operationalize platform standards across datacenters
- Create scalable runbooks, troubleshooting guides, and error catalogs
- Turn investigations into playbooks that elevate L1/L2 teams
- Travel to datacenters for complex troubleshooting, new platform readiness, or incident containment
- Company mission
- Nebius is leading a new era in cloud infrastructure for the global AI economy, building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment
Primary stack
Core technologies
Python (Programming Language)
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
Requirements & details
- Strong hands-on experience with datacenter servers and deep Linux troubleshooting
- Ability to diagnose across hardware, BIOS/BMC firmware, and Linux (logs, drivers, storage basics, performance triage)
- Structured incident response experience and clear communication under pressure
- Experience driving evidence-based escalations with vendors/R&D
- Fluent English (written and spoken)
- Bonus: Strong familiarity with GPU server platforms and tooling (for example: nvidia-smi, dcgmi, Linux logs correlation)
- Bonus: Experience with ipmitool and Redfish workflows, firmware lifecycle, and staged rollouts
- Bonus: Scripting skills (bash and basic Python) for log collection, triage automation, and simple reliability analysis
- Bonus: Exposure to OCP-based platforms and ODM manufacturing ecosystems
- Bonus: Experience supporting enterprise bare metal customers under contractual SLAs
- Work location: Béthune, France
- Python, Linux, bash, nvidia-smi, dcgmi, ipmitool, Redfish
- Python (Programming Language)
