SMART said PASSED. The drive had 477 dead sectors.
Every hard drive carries a self-assessment: one bit that says healthy or failing. It is the field most monitoring setups check, because it is the one the drive volunteers. We spent two weeks measuring how far that bit can sit from the raw counters underneath it. On one drive the answer was 477 retired sectors and a pending count swinging up to 96, with the bit reading PASSED throughout.
What we set out to test
Glassmkr exists on a claim: that we can tell you a drive is going before it takes your data with it. That is an easy claim to make and an uncomfortable one to check, because checking it honestly means finding out whether your own product is mostly theatre.
So we built a fleet to find out. Thirteen boxes of real spinning rust, most of them years into service. Three of them we designated as expendable and abused deliberately. The rest we left alone as a control group, because a detector that fires on everything is not a detector.
The catch
One drive, a 4 TB HGST at just over 38,000 power-on hours, was carrying 477 reallocated sectors. A reallocated sector is one the drive has given up on and silently swapped for a spare from a small reserve pool. Four hundred and seventy-seven of them is not a rough patch. It is a drive working through its reserve.
Glassmkr raised a critical alert on it. What makes that worth writing about is what the drive itself was saying at the same moment:
serial V6G97HRR reallocated 477 health PASSED PASSED. Not in one unlucky reading: in all 140 readings we logged for that drive. Not once, across the entire campaign, did its own self-assessment admit anything was wrong. As of publication it is still saying PASSED, still at 477, some eleven days after we first looked.
The other twelve campaign boxes produced no drive alert at all, and that negative result matters as much as the positive one. Widen the lens to the whole fleet we monitor, 21 hosts and 87 drives at the time of writing, and this is still the only drive raising a SMART failure. One true positive, everything else quiet. A detector that finds the bad drive and stays silent about the other 86 is the only kind worth running.
Why the health bit is like this
This is not a bug in the drive, and the drive is not lying. The overall-health verdict is a pass or fail against normalised attribute values, not against the raw counts an operator reads. Every SMART attribute carries both. The raw value is the physical count: 477 sectors. The normalised value is a vendor-scaled score, typically starting near 100 and falling as the attribute worsens, and the drive reports FAILED only when that score drops to the vendor's threshold.
Those thresholds are set where the manufacturer expects imminent, warranty-relevant death, and on a 4 TB drive the spare sector pool is large enough that 477 retired sectors barely moves the score. So the raw count sat at a number any operator would act on immediately, while the normalised score stayed comfortably in passing territory. Both facts are true at once, and the verdict is behaving exactly as specified.
The bit is doing its job. Its job is simply not the job people use it for. It answers "has this drive crossed the manufacturer's replace-under-warranty line", and it gets read as "is this drive fine". The mistake is not that the firmware got it wrong. The mistake is trusting a one-bit verdict instead of reading the raw attributes underneath it.
The same drive passed a full surface test
Here is the part that surprised us. We ran a complete destructive write-and-verify pass over that drive: write every sector, read it all back, compare. If 477 sectors are dead, a full surface test should notice.
It came back clean. Zero errors.
That is correct behaviour, and it is the whole problem. Reallocation is the drive hiding the damage: the bad sector is retired, a spare is mapped in its place, and every read and write after that is served by good media. From the operating system's point of view nothing happened. The host-visible IO path is clean precisely because the drive is spending its reserve to keep it clean.
So two independent checks both came back clean, and neither was wrong. The health verdict truthfully reported that no normalised attribute had crossed its threshold. The surface test truthfully reported that every sector it touched read back correctly. Both answered the question they were asked. Neither was asked "how much of this drive has already been retired", and the only place that answer exists is the raw attribute counter, visible only to something watching it over time.
The part that should end the argument: pending sectors
If the 477 does not move you, this should. Pending sectors, SMART attribute 197, are sectors the drive has found suspect and has not yet been able to retire. A non-zero pending count is the one reading that even operators who trust the health verdict treat as replace-this-now: it means there is data the drive currently cannot read reliably.
Across our 140 readings of that drive, all three series at once:

When we started watching, the count was already at 96 and took about two days to drain to zero. Later in the window it rose from zero three more times, peaking at 16, then 32, then 16, all inside twenty hours. In total 89 of the 140 readings had pending sectors above zero. At no point during any of that did the overall-health verdict stop saying PASSED. This is the hard version of the finding, and it does not depend on anyone agreeing that 477 is alarming: the drive was repeatedly unable to read its own sectors, and its self-assessment reported healthy throughout.
There is a second lesson buried in the same data, and it is why we do not alert on pending. The counter is supposed to move: the drive finds a suspect sector, retries it later, and either clears it or retires it. An alert keyed on pending would have fired and self-cleared several times in a little over a week on this one drive, and nobody keeps reading an alert that does that. So we anchor the critical on reallocated sectors, which only ever go up, and on the drive's own failing verdict. We considered pending as a trigger and the data talked us out of it, while that same data made the case above.
We tried to kill four healthy drives and could not
The other half of the campaign was an attempt to manufacture a failure, so we would have more than one bad drive to learn from. We took drives with roughly 59,000 power-on hours, which is about six and a half years of continuous service, and hammered them.
Two days of continuous random writes produced nothing. Eight of the nine drives under stress finished with exactly zero reallocated and zero pending sectors, jobs completing normally throughout. Then full write-and-verify passes over four of them, several hours each, every sector written and read back: zero errors, zero new defects, all four still reporting clean afterwards.
We could not break them, and the failure of that experiment is itself a finding. Reallocations track latent media defects, not how hard you have been writing. Write volume is not the driver, which means "this drive has had a heavy workload" is not a risk signal, and the only drive that produced new defects was the one that was already marginal.
What we are not claiming
One failing drive is a demonstration, not a dataset, and we would rather say so than dress it up.
We found exactly one bad drive, and we found it already at 477. We did not watch it climb from zero, so we cannot tell you what number is worth worrying about. Is a drive at 3 reallocated sectors in trouble, or is it fine for another four years? Our campaign cannot answer that, and any threshold we published off a single drive would be an invention with a number attached to it.
That has a concrete consequence for the product. We considered splitting the alert so that a low, stable reallocated count became a quieter advisory rather than a critical. We are not shipping that yet, because we cannot calibrate where "low" ends, and the two ways of being wrong are not equally bad. Telling you to investigate a drive that is actually dying costs you data. Paging you about a drive that turns out to be stable costs you a ticket. Until we can draw that line from real failures rather than from one anecdote, reallocated sectors stay critical.
What to take from this
If you run bare metal, the useful conclusions are cheap to act on:
- Do not read the overall-health verdict as a health check. It is a normalised-threshold test with a warranty question behind it. A drive can be deep into its spare sector pool, and intermittently unable to read its own sectors, and still report PASSED. Ours did, in every reading.
- Read the raw attributes, and read them over time. This is the one rule that needs no threshold and no failure statistics, so it is the one we can hand you with a straight face: a single reading tells you almost nothing, and the same counters read weekly tell you nearly everything. Trajectory is the signal. The verdict is not.
- Alert on the counter that latches, and only that one. Reallocated sectors only ever go up: once a drive retires a sector it does not un-retire it, and ours sat at exactly 477 in all 140 readings. That makes it safe to page on. Pending sectors are the opposite, oscillating by design, and ours swung between 0 and 96 while the actual damage never changed, so read them as evidence rather than wiring them to a pager. Offline uncorrectable sectors, attribute 198, behave more like pending than like reallocated: they can clear too, so do not assume they latch.
- A clean surface test does not clear a drive. Reallocation exists to make the damage invisible to exactly that test.
- A raw count needs interpreting in both directions. The drive above is the alarming version: 477 retired sectors under a PASSED verdict. Another drive in the same fleet is the mirror image, sitting at 138 interface CRC errors with zero reallocated sectors and, yes, also a PASSED verdict. Attribute 199 counts errors on the cable and backplane between controller and drive, not on the media, so that second drive is fine and replacing it would fix nothing; the thing to reseat is the cable. Same verdict on both drives, opposite realities, and the verdict distinguished neither. Only the raw attributes did.
Glassmkr does this watching for you, which is why the drive above raised a critical alert on the raw counts while its own verdict, correctly and unhelpfully, still read PASSED. But none of the findings here depend on our product. They are properties of the drives you already own, and the counters are there whether or not anything is reading them.