Extremely specific disk write errors

to put it very briefly, i spent close to a month trying to get 2 of my new drives to behave, they just contain steam games so nothing critical

no the disks are not bad, turns out both just had bad sata cables, i also checked on a test bench and both drives are fine hardware wise

yet weirdly still, every now n then on the odd reboot, one of the two panics and goes into read only mode

you can still see the files, a reboot or shutdown and come back later fixes it, but it is annoying that it keeps happening, and its not sata port dependant

both have btrfs, in fact my entire system does, and the rest of the other 17 disks in my systems have no issues, its only the two ssds i have mounted in the back of my pc

does anyone know why this is happening if all the hardware is verified as good? even checked the psu

dmesg isnt very happy about whats happening either, and tries to reset the sata link when this happens, repeatedly, where its suddenly unresponsive

2386.358895] sd 6:0:0:0: [sdc] tag#0 FAILED Result: hostbyte=DID_BAD_TARGET driverbyte=DRIVER_OK cmd_age=0s
[ 2386.358897] sd 6:0:0:0: [sdc] tag#0 CDB: Read(10) 28 00 00 0e 19 c0 00 00 20 00
[ 2386.358898] I/O error, dev sdc, sector 924096 op 0x0:(READ) flags 0x1000 phys_seg 4 prio class 2
[ 2386.358899] BTRFS error (device sdc1 state EA): bdev /dev/sdc1 errs: wr 14, rd 2931, flush 0, corrupt 0, ge
n 0
[ 2433.645390] scsi_io_completion_action: 815 callbacks suppressed
[ 2433.645396] sd 6:0:0:0: [sdc] tag#0 FAILED Result: hostbyte=DID_BAD_TARGET driverbyte=DRIVER_OK cmd_age=0s
[ 2433.645399] sd 6:0:0:0: [sdc] tag#0 CDB: ATA command pass through(16) 85 06 20 00 00 00 00 00 00 00 00 00 0
0 00 e5 00

and then that spams for a billion lines

So with 17 disks there are the motherboard and (some) extensions cards involved? And you switched cables and ports (cards/board)? Are the disks from the same manufacturer? Did the manufacturer maybe issue a firmware update that can be applied?

its 19 disks total, and while yes i do have an extension card, these disks arent on it, sata port switching did nothing for it aswell

both are a fanxiang S101Q if it helps any, but poking on their website i dont see any firmware downloads

Aliexpress…i would not expect to much from it.

i got mine from amazon, and it certainly wasnt an ali express price

Are these the only drives that are fanxiang S101Q?

This means the drive is unreachable for some reason (perhaps obvious). That might be a cable issue, or plugs, or an issue with the drive itself.

That’s 2,931 read errors, and 14 write errors.

What’s the output of:

sudo smartctl --all /dev/sdc

nope, i have a third but it has no issues

=== START OF INFORMATION SECTION ===
Device Model:     Fanxiang S101Q 1TB
Serial Number:    MX_00000000000010011
LU WWN Device Id: 5 000000 000000011
Firmware Version: SN17003
User Capacity:    1,024,209,543,168 bytes [1.02 TB]
Sector Size:      512 bytes logical/physical
Rotation Rate:    Solid State Device
Form Factor:      2.5 inches
TRIM Command:     Available, deterministic, zeroed
Device is:        Not in smartctl database 7.5/5706
ATA Version is:   ACS-4 (minor revision not indicated)
SATA Version is:  SATA 3.2, 6.0 Gb/s (current: 6.0 Gb/s)
Local Time is:    Thu May 28 15:35:02 2026 CDT
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status:  (0x00) Offline data collection activity
                                        was never started.
                                        Auto Offline Data Collection: Disabled.
Self-test execution status:      (   0) The previous self-test routine completed
                                        without error or no self-test has ever 
                                        been run.
Total time to complete Offline 
data collection:                (   33) seconds.
Offline data collection
capabilities:                    (0x7b) SMART execute Offline immediate.
                                        Auto Offline data collection on/off support.
                                        Suspend Offline collection upon new
                                        command.
                                        Offline surface scan supported.
                                        Self-test supported.
                                        Conveyance Self-test supported.
                                        Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
                                        power-saving mode.
                                        Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
                                        General Purpose Logging supported.
Short self-test routine 
recommended polling time:        (   2) minutes.
Extended self-test routine
recommended polling time:        (  85) minutes.
Conveyance self-test routine
recommended polling time:        (   2) minutes.
SCT capabilities:              (0x0031) SCT Status supported.
                                        SCT Feature Control supported.
                                        SCT Data Table supported.

SMART Attributes Data Structure revision number: 20
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
  5 Reallocated_Sector_Ct   0x0013   100   100   050    Pre-fail  Always       -       0
  9 Power_On_Hours          0x0012   100   100   000    Old_age   Always       -       3950
 12 Power_Cycle_Count       0x0012   100   100   000    Old_age   Always       -       779
167 Unknown_Attribute       0x0022   100   100   000    Old_age   Always       -       0
168 Unknown_Attribute       0x0012   100   100   000    Old_age   Always       -       0
169 Unknown_Attribute       0x0013   100   100   010    Pre-fail  Always       -       196614
171 Unknown_Attribute       0x0032   000   000   000    Old_age   Always       -       0
172 Unknown_Attribute       0x0032   000   000   000    Old_age   Always       -       0
173 Unknown_Attribute       0x0012   200   200   000    Old_age   Always       -       8590393347
175 Program_Fail_Count_Chip 0x0022   100   100   010    Old_age   Always       -       0
177 Wear_Leveling_Count     0x0012   100   100   000    Old_age   Always       -       9280464
180 Unused_Rsvd_Blk_Cnt_Tot 0x0033   100   100   000    Pre-fail  Always       -       117
187 Reported_Uncorrect      0x0032   100   000   000    Old_age   Always       -       0
192 Power-Off_Retract_Count 0x0012   100   100   000    Old_age   Always       -       163
194 Temperature_Celsius     0x0022   042   042   000    Old_age   Always       -       42 (Min/Max 30/53)
196 Reallocated_Event_Count 0x0012   100   100   000    Old_age   Always       -       0
199 UDMA_CRC_Error_Count    0x0012   100   100   000    Old_age   Always       -       0
206 Unknown_SSD_Attribute   0x0032   200   200   000    Old_age   Always       -       2
207 Unknown_SSD_Attribute   0x0032   200   200   000    Old_age   Always       -       7
208 Unknown_SSD_Attribute   0x0032   200   200   000    Old_age   Always       -       3
209 Unknown_SSD_Attribute   0x0032   200   200   000    Old_age   Always       -       1
210 Unknown_Attribute       0x0032   200   200   000    Old_age   Always       -       19
211 Unknown_Attribute       0x0032   200   200   000    Old_age   Always       -       7
231 Unknown_SSD_Attribute   0x0023   100   100   005    Pre-fail  Always       -       0
241 Total_LBAs_Written      0x0032   100   100   000    Old_age   Always       -       332
242 Total_LBAs_Read         0x0032   100   100   000    Old_age   Always       -       759
245 Unknown_Attribute       0x0032   100   100   000    Old_age   Always       -       524296

SMART Error Log Version: 1
No Errors Logged

SMART Self-test log structure revision number 1
No self-tests have been logged.  [To run self-tests, use: smartctl -t]

SMART Selective self-test log data structure revision number 1
 SPAN  MIN_LBA  MAX_LBA  CURRENT_TEST_STATUS
    1        0        0  Not_testing
    2        0        0  Not_testing
    3        0        0  Not_testing
    4        0        0  Not_testing
    5        0        0  Not_testing
Selective self-test flags (0x0):
  After scanning selected spans, do NOT read-scan remainder of disk.
If Selective self-test is pending on power-up, resume after 0 minute delay.

The above only provides legacy SMART information - try 'smartctl -x' for more

That is an way too many Unknown Attributes for a SSD.

what am i supposed to do about it?

Device Model:     Fanxiang S101Q 1TB
Serial Number:    MX_00000000000010011
LU WWN Device Id: 5 000000 000000011

Fanxiang are low budget SSD devices and the Serial Number seems to be fake. This very much looks like bad hardware to start with.

What you could try is to turn off SATA link power management (lpm) and see if that helps.

Add this to the kernel commandline:

ahci.mobile_lpm_policy=1

that appears to be for mobile chipsets, not sure if it matters

but i can just turn it off in the bios

Maybe those two are bad!

they still work, it was just aggresive sata power management since they start up quite slow and the max delay the system expects is like 10ms

they read, they write, and its a real bad time to get replacements

So the solution was ahci.mobile_lpm_policy=1
?

it was a bios option is how i turned it off is all, i didnt try the launch option, might work for a laptop user tho

Should work for you in this application also.