- Location
- Santa Clara, CA,US, US · Austin, TX,US, US · Westford, MA,US, US · Durham, NC,US, US · Redmond, WA,US, US
- Type
- Full-time
- Department
- Engineering
- Seniority
- Senior
- Source
- Eightfold
Description
Architect Failure Attribution Frameworks: Build a scalable "flight recorder" for EDA jobs that captures high-fidelity state across the CPU, GPU, and Fabric at the moment of failure. Build automated diagnostics that correlate GPU XID errors, PCIe bus failures, and CUDA memory exceptions. Connect these errors with system-level events such as OOM kills or NUMA-related hangs. Distributed Logging & Tracing: Implement low-overhead tracing mechanisms (using tracing tools or custom agents) that provide access to job execution across multi-node Slurm or Kubernetes clusters. Root Cause Automation: Develop heuristics and models based on machine learning to classify failures as "Hardware Fault," "Software Bug," or "Environment Issue." This reduces the Mean Time to Identify (MTTI) for R&D teams. Resiliency Engineering: Work closely with hardware and infrastructure teams to define "signals of impending failure," enabling proactive job migration or check-pointing before a crash occurs. Experience building automated RCA (Root Cause Analysis) pipelines for HPC or cloud-scale environments. Experience with tools doing non-intrusive monitoring of application health and syscall-level failure patterns. Experience with checkpoint/restore technologies (like CRIU) and their application in long-running EDA flows.