Article
The disk came back. The incident wasn't over.
Why the return of a component is not the same as the return of confidence in a system.
Visible recovery is not confidence#
An NVMe device failed in a ZFS mirror. No data was lost. The device later reappeared. That was a positive signal, but it was not sufficient reason to declare the incident over.
The return of one component initially answers only a narrow question: is it visible again right now? It does not establish that the state of the wider system is understood, that recovery measures are sound, or that the conditions surrounding the failure have been investigated far enough.
Recovery of a component is not the same as recovery of confidence.
Confidence does not come from one reassuring signal. It comes from several verified statements about data, recoverability, and operation. While one of those statements remains open, the assessment of the incident remains open as well. The visible symptom may disappear while the underlying technical uncertainty remains.
Preserve the state first#
The first step was to preserve the state. This matters because every subsequent action can change what is observable. Correcting the situation immediately may remove exactly the information needed to form a dependable account later.
Backups were checked next. No data loss in the mirror is an important finding, but it does not replace verification of independent recovery. An incident concerns more than the data that is currently visible. It also concerns the ability to return to a known state if conditions deteriorate.
Ongoing operations were checked as well. A system can appear available while still operating in a condition that deserves attention. Operational verification therefore belongs beside data and backup checks, not behind a premature all-clear.
Analyze after securing the situation#
Only after the state, backups, and operation had been checked was the write load analyzed. That order separates two jobs. First, uncertainty is prevented from turning into an uncontrolled state. Then the load affecting the system can be examined.
The write-load analysis led to targeted corrections. The documented record does not support a more specific causal story. There is no established hardware cause here, no published SMART data, and no basis for invented pool names or commands.
That limitation is part of responsible incident documentation. A plausible story about the cause would be easier to tell, but it would imply confidence that the known sequence does not provide.
When the incident has actually moved forward#
An incident is not complete simply because its most visible symptom has disappeared. The technical state needs to be preserved, recovery needs to be checked, and ongoing operation needs to be understood. Analysis can then justify targeted changes.
In this case, the NVMe device returning was an event within the investigation, not its conclusion. Progress came from restored confidence: no data loss, backups checked, operation checked, write load analyzed, and targeted corrections derived from that analysis.
The practical lesson is concise. When a component returns, the first question is not whether the all-clear can be given. It is what uncertainty remains.
That question changes the standard for recovery. The presence of the component is not enough. What matters is how many dependable statements about the system can be established again. Those statements turn visible recovery back into justified confidence.