Staff Systems Software Engineer- Server
Indexed description
We are looking for a highly motivated and creative Staff System Software Engineer who is experienced and passionate about diagnostics for NVIDIA next-generation GPU products. You will lead and contribute to the design, implementation, and integration of these software-based validations into the sophisticated manufacturing flow for NVIDIA GPU products.
What You’ll Be Doing
- Implement and enhance GPU diagnostics covering power, thermal, memory, PCIe, NVLink, and system‑level checks on boards, servers, and racks.
- Develop stress and validation tests that exercise GPU subsystems and platform components; add clear pass/fail criteria, telemetry, and error codes suitable for automation and failure attribution.
- Contribute to the integration of diagnostics into L6/L10/L11 factory flows and datacenter workflows, working with senior engineers to define coverage, runtime, and sequencing.
- Execute and monitor automated regression tests and pipelines on top of orchestration systems.
- Debug issues in cooperation with hardware, firmware, and other teams (Ops/TE/AE, etc.); root‑cause problems that span HW/FW/SW boundaries.
- Improve logging and reporting for diagnostics (e.g., structured logs, JSON, exit codes) to make triage simpler.
- Analyze factory and field data (yields, error trends, intermittency) to identify gaps in diagnostics or weak debug signals and propose targeted test or logging improvements.
- Help write and maintain technical documentation: test specs, user guides, SOPs, and troubleshooting guides for internal teams and external partners (ODMs/OEMs/CSPs).
- BS or MS in Computer Science, Electrical/Computer Engineering, or a related field, with 8+ years of relevant experience in system software, diagnostics, or platform validation.
- Strong programming skills in C/C++ and Python, plus familiarity with shell scripting for lab tools and automation.
- Experience developing low‑level system software on Linux (e.g., board bring‑up tools, diagnostics, or drivers), working close to firmware and hardware registers.
- Solid understanding of server platforms:
- x86 and/or ARM server architecture (CPU, memory, PCIe topology, firmware/boot flow).
- Basic GPU architecture concepts and how GPU accelerators are integrated into servers.
- Ability to read and interpret HW and system logs and understand system/block diagrams and schematics in order to connect failures back to the underlying architecture.
- Demonstrated debugging and problem‑solving skills, including use of tools like gdb, perf, tracing frameworks, and vendor debug utilities.
- Demonstrated technical leadership with a strong commitment to driving projects, investigations, and critical bugs to closure across cross-functional teams.
- Strong communication skills and a collaborative mindset; comfortable working with globally distributed teams across HW, FW, SW, operations, and manufacturing.
- Hands‑on experience with diagnostic tools in factory or datacenter environments.
- Direct involvement in defining or improving diagnostic flows or RMA qualification flows for complex hardware products.
- Proven track record writing clear, high‑coverage test plans and test cases for complex HW/SW systems.
- Experience using AI‑assisted tools (for example, for log triage, code assistance, data analysis, or test generation) to accelerate diagnostics development and debugging.
, , JR2014668
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search