Incident

NVMe Failure in a ZFS Mirror

Incident work after an NVMe failure without data loss, from preserving state to analyzing write load.

What happened#

An NVMe device failed in a ZFS mirror. No data was lost. The device later reappeared, but its return was not treated as sufficient evidence that the incident was over.

Immediate response#

The first step was to preserve the state. Backups and ongoing operations were then checked. This order kept the observed situation available before further analysis or changes could alter it.

Investigation#

Once the data, backups, and operation had been checked, the write load was analyzed. The goal was not to stop at the visible recovery of one component, but to rebuild confidence in the system.

What changed afterwards#

The analysis led to targeted corrections. No specific hardware cause, SMART values, pool names, or internal sizes are established by the available facts, so none are stated here.

Lesson#

A device returning does not automatically close an incident. The component may be back while confidence in the state of the wider system still needs to be recovered.