On 2026-09-26 21:37, Andy Smith wrote:
Hi Gary,
On Sat, Sep 26, 2026 at 08:14:02PM -0400, Gary Dale wrote:
The flaw in the logic is that the drive is obviously working as it is part
of a RAID array.
I have tried to explain multiple times that this is not an accurate
description. An accurate description is that parts of your drive are in
active RAID arrays. That isn't the same as "all of my drive is in a RAID
array".
Linux can only see I/O errors when it encounters them. If one part of a
drive is bad, but it is never read from or written to, Linux will never
know. So it is perfectly possible to have an unreadable part of a drive
that you can't put into an MD array, but other parts of the drive are in
working MD arrays.
What isn't working it getting one partition to rejoin the
array it used to belong to.
As far as I was aware this was because you saw in your logs that Linux
encountered an I/O error when you tried to do that.
Anyway, a successful write to the partition had no impact.
So to elaborate, are you saying that you looked at the logs to find out
which sector could not be read, confirmed it could not be read with
hdparm (or dd) and then wrote over it with hdparm, as suggested? Or do
you mean that you wrote over every sector of the partition in some other
way?
The reason why I keep saying about confirming if things can be read
with other tools is precisely to rule out hardware problems (one or more
unreadable sectors).
I know you did manage to do a SMART long self-test that came out fine,
but if the kernel is reporting an I/O error I still think it would be
good to confirm whether that sector is readable generally or not.
There is something about the partition that mdadm doesn't like as evidenced
by the error message:
#mdadm --manage /dev/md0 --add /dev/sda1
mdadm: add new device failed for /dev/sda1 as 7: Invalid argument
mdadm: Cannot read superblock on /dev/sda1
And when this happens are there any relevant logs in the journal?
If there's an I/O error does it mention a sector number? Is it the same
sector number every time? Can you read that sector number with other
tools?
Either it is in the disk I am trying to add or it is in the array
information.
It may be a good time to ask on the linux-raid mailing list.
THanks,
Andy
Thanks Andy, but from what I've read, it's simpler to use e2fsck to
check for bad blocks. It turns out that the issue had nothing to do with
bad blocks at all. When I tried e2fsck, which I wouldn't expect to see a
valid ext file system, it pointed to the real problem:
# e2fsck /dev/sda1 -L /root/badblocks.text
e2fsck 1.47.2 (1-Jan-2025)
ext2fs_open2: Bad magic number in super-block
e2fsck: Superblock invalid, trying backup blocks...
e2fsck: Bad magic number in super-block while trying to open /dev/sda1
The superblock could not be read or does not describe a valid ext2/ext3/ext4
filesystem. If the device is valid and it really contains an ext2/ext3/ext4
filesystem (and not swap or ufs or something else), then the superblock
is corrupt, and you might try running e2fsck with an alternate superblock:
e2fsck -b 8193 <device>
or
e2fsck -b 32768 <device>
So I ran mkfs -t ext4 /dev/sda1. When it finished, I was able to add the
partition back in. The array is currently rebuilding.