Analyze btrfs uncorrectable errors

Hi,

Booted up my Kinoite station today and quickly ran into issues with my main drive running btrfs. Programs start to break reporting that some config files can not be written to. Seems like my drive is remounted as read-only because of some issues.

journalctl:

Sep 08 20:56:24 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:56:24 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:56:24 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:56:24 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:56:24 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:56:24 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:56:24 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:56:24 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:56:38 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:56:38 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:56:38 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:56:38 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:56:38 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:56:38 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:57:05 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:57:05 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:57:05 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:57:05 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:57:05 bude-desktop kernel: BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:57:18 bude-desktop kernel: BTRFS: error (device nvme0n1p3 state A) in btrfs_add_link:6937: errno=-5 IO failure
Sep 08 20:57:18 bude-desktop kernel: BTRFS info (device nvme0n1p3 state EA): forced readonly
Sep 08 20:57:18 bude-desktop kernel: BTRFS: error (device nvme0n1p3 state EA) in btrfs_create_new_inode:6873: errno=-5 IO failure
Sep 08 20:57:22 bude-desktop kernel: btrfs_validate_extent_buffer: 76 callbacks suppressed
Sep 08 20:57:22 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:57:22 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:57:22 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:57:22 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:57:54 bude-desktop flatpak[4169]: [Parent 2, IPC I/O Parent] WARNING: process 623 exited on signal 15: file checkouts/gecko/ipc/chromium/src/chrome/common/process_watcher_posix_sigchld.cc:161
Sep 08 20:58:08 bude-desktop auditd[1039]: Audit daemon is suspending logging due to previously mentioned write error
Sep 08 20:58:21 bude-desktop kernel: btrfs_validate_extent_buffer: 14 callbacks suppressed
Sep 08 20:58:21 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:59:47 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:59:47 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:59:47 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 20:59:47 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 20:59:47 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:00:02 bude-desktop kernel: btrfs_validate_extent_buffer: 8 callbacks suppressed
Sep 08 21:00:02 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 21:00:02 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:00:13 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 21:00:13 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:00:27 bude-desktop kdialog[7246]: This plugin does not support propagateSizeHints()
Sep 08 21:00:28 bude-desktop kdialog[7259]: This plugin does not support propagateSizeHints()
Sep 08 21:00:28 bude-desktop kdialog[7262]: This plugin does not support propagateSizeHints()
Sep 08 21:00:28 bude-desktop kdialog[7263]: This plugin does not support propagateSizeHints()
Sep 08 21:00:28 bude-desktop kdialog[7265]: This plugin does not support propagateSizeHints()
Sep 08 21:00:29 bude-desktop wireplumber[2139]: wp-state: <WpState:0x55ac648d1df0> could not save stream-properties: Failed to create file “/var/home/bude/.local/state/wireplumber/stream-properties.U55IV3”: Read-only file system
Sep 08 21:02:15 bude-desktop systemd[2009]: app-org.kde.kwrite@130093f1672047d0861d6940be4467c2.service: Consumed 1.023s CPU time over 6min 21.585s wall clock time, 187.6M memory peak.
Sep 08 21:02:20 bude-desktop firefox-bin[4169]: App Icon is not available when using Portal Notifications
Sep 08 21:02:21 bude-desktop xdg-desktop-portal[7599]: (xdg-desktop-portal-validate-icon:2): GLib-WARNING **: 19:02:21.001: getpwuid_r(): failed due to unknown user id (1000)
Sep 08 21:02:21 bude-desktop xdg-desktop-portal[7599]: (xdg-desktop-portal-validate-icon:2): GLib-WARNING **: 19:02:21.001: Could not find home directory: $HOME is not set, and user database could not be read.
Sep 08 21:02:23 bude-desktop systemd[2009]: app-org.kde.konsole@8d44cf5514a6407783c385c17f4b4164.service: Consumed 648ms CPU time over 23min 34.979s wall clock time, 202.9M memory peak.
Sep 08 21:02:23 bude-desktop systemd[2009]: app-org.kde.konsole-3272.scope: Consumed 8.899s CPU time over 23min 34.780s wall clock time, 45.6M memory peak.
Sep 08 21:02:46 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:02:47 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 21:02:47 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:02:50 bude-desktop wireplumber[2139]: wp-state: <WpState:0x55ac648d1df0> could not save stream-properties: Failed to create file “/var/home/bude/.local/state/wireplumber/stream-properties.8LISV3”: Read-only file system
Sep 08 21:02:55 bude-desktop systemd[2009]: Started app-systemsettings@8ce06a5a47f44619ae002a5d46fbd608.service - System Settings - System Settings.
Sep 08 21:02:56 bude-desktop wireplumber[2139]: wp-state: <WpState:0x55ac648d1df0> could not save stream-properties: Failed to create file “/var/home/bude/.local/state/wireplumber/stream-properties.TSHJV3”: Read-only file system
Sep 08 21:02:56 bude-desktop systemsettings[8044]: Failed to connect to Bolt manager DBus interface: 
Sep 08 21:02:57 bude-desktop systemsettings[8044]: Failed to register with host portal QDBusError("org.freedesktop.portal.Error.Failed", "Could not register app ID: Connection already associated with an application ID")
Sep 08 21:02:58 bude-desktop systemd[2009]: app-systemsettings@8ce06a5a47f44619ae002a5d46fbd608.service: Consumed 1.106s CPU time over 3.361s wall clock time, 138.5M memory peak.
Sep 08 21:03:06 bude-desktop kernel: btrfs_validate_extent_buffer: 46 callbacks suppressed
Sep 08 21:03:06 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 21:03:06 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:03:06 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 21:03:06 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:03:06 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 21:03:06 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:03:20 bude-desktop systemd[2009]: Started app-org.kde.dolphin@d1aa4f926406420bbc3f741f20e1cfbb.service - Dolphin - File Manager.
Sep 08 21:03:20 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 21:03:20 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:03:20 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 21:03:20 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:03:20 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
Sep 08 21:03:20 bude-desktop kernel: BTRFS error (device nvme0n1p3 state EA): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
Sep 08 21:03:21 bude-desktop wireplumber[2139]: wp-state: <WpState:0x55ac648d1df0> could not save stream-properties: Failed to create file “/var/home/bude/.local/state/wireplumber/stream-properties.EGLIV3”: Read-only file system
Sep 08 21:03:21 bude-desktop kdialog[8223]: This plugin does not support propagateSizeHints()
Sep 08 21:03:21 bude-desktop kdialog[8238]: This plugin does not support propagateSizeHints()
Sep 08 21:03:21 bude-desktop kdialog[8241]: This plugin does not support propagateSizeHints()
Sep 08 21:03:21 bude-desktop kdialog[8243]: This plugin does not support propagateSizeHints()
Sep 08 21:03:21 bude-desktop kdialog[8242]: This plugin does not support propagateSizeHints()
Sep 08 21:03:22 bude-desktop wireplumber[2139]: wp-state: <WpState:0x55ac648d1df0> could not save stream-properties: Failed to create file “/var/home/bude/.local/state/wireplumber/stream-properties.6E7TV3”: Read-only file system

btrfs check (–force; ran via live image with the same result):

Filebin | 51gt3vhrk6z41y8i (14MB, I can upload somewhere else but pastebin and github gist both would not work)

btrfs scrub:

UUID:             4ba34f82-4666-4347-93fd-599d39e9e1d9
Scrub started:    Tue Sep  8 20:39:22 2026
Status:           finished
Duration:         0:05:58
Total to scrub:   618.98GiB
Rate:             1.73GiB/s
Error summary:    verify=8
Corrected:      0
Uncorrectable:  8
Unverified:     0

dmesg:

[   77.107107] BTRFS info (device nvme0n1p3): scrub: started on devid 1
[   77.194811] BTRFS warning (device nvme0n1p3): scrub: tree block 173948928 mirror 1 has bad generation, has 63386 want 63489
[   77.195980] BTRFS warning (device nvme0n1p3): scrub: tree block 173948928 mirror 1 has bad generation, has 63386 want 63489
[   77.196237] BTRFS warning (device nvme0n1p3): scrub: tree block 173948928 mirror 1 has bad generation, has 63386 want 63489
[   77.196361] BTRFS warning (device nvme0n1p3): scrub: tree block 173948928 mirror 1 has bad generation, has 63386 want 63489
[   77.196364] BTRFS error (device nvme0n1p3): scrub: unable to fixup (regular) error at logical 173932544 on dev /dev/nvme0n1p3 physical 182321152
[   77.196371] BTRFS warning (device nvme0n1p3): scrub: generation error at logical 173932544 on dev /dev/nvme0n1p3, physical 182321152: metadata leaf (level 0) in tree 7
[   77.196374] BTRFS error (device nvme0n1p3): scrub: unable to fixup (regular) error at logical 173932544 on dev /dev/nvme0n1p3 physical 182321152
[   77.196378] BTRFS warning (device nvme0n1p3): scrub: generation error at logical 173932544 on dev /dev/nvme0n1p3, physical 182321152: metadata leaf (level 0) in tree 7
[   77.196381] BTRFS error (device nvme0n1p3): scrub: unable to fixup (regular) error at logical 173932544 on dev /dev/nvme0n1p3 physical 182321152
[   77.196385] BTRFS warning (device nvme0n1p3): scrub: generation error at logical 173932544 on dev /dev/nvme0n1p3, physical 182321152: metadata leaf (level 0) in tree 7
[   77.196388] BTRFS error (device nvme0n1p3): scrub: unable to fixup (regular) error at logical 173932544 on dev /dev/nvme0n1p3 physical 182321152
[   77.196391] BTRFS warning (device nvme0n1p3): scrub: generation error at logical 173932544 on dev /dev/nvme0n1p3, physical 182321152: metadata leaf (level 0) in tree 7
[   77.196394] BTRFS error (device nvme0n1p3): bdev /dev/nvme0n1p3 errs: wr 0, rd 0, flush 0, corrupt 2, gen 5
[   77.888084] BTRFS warning (device nvme0n1p3): scrub: tree block 173948928 mirror 2 has bad generation, has 63386 want 63489
[   77.889176] BTRFS warning (device nvme0n1p3): scrub: tree block 173948928 mirror 2 has bad generation, has 63386 want 63489
[   77.889433] BTRFS warning (device nvme0n1p3): scrub: tree block 173948928 mirror 2 has bad generation, has 63386 want 63489
[   77.889684] BTRFS warning (device nvme0n1p3): scrub: tree block 173948928 mirror 2 has bad generation, has 63386 want 63489
[   77.889688] BTRFS error (device nvme0n1p3): scrub: unable to fixup (regular) error at logical 173932544 on dev /dev/nvme0n1p3 physical 1256062976
[   77.889694] BTRFS warning (device nvme0n1p3): scrub: generation error at logical 173932544 on dev /dev/nvme0n1p3, physical 1256062976: metadata leaf (level 0) in tree 7
[   77.889698] BTRFS error (device nvme0n1p3): scrub: unable to fixup (regular) error at logical 173932544 on dev /dev/nvme0n1p3 physical 1256062976
[   77.889702] BTRFS warning (device nvme0n1p3): scrub: generation error at logical 173932544 on dev /dev/nvme0n1p3, physical 1256062976: metadata leaf (level 0) in tree 7
[   77.889705] BTRFS error (device nvme0n1p3): scrub: unable to fixup (regular) error at logical 173932544 on dev /dev/nvme0n1p3 physical 1256062976
[   77.889708] BTRFS warning (device nvme0n1p3): scrub: generation error at logical 173932544 on dev /dev/nvme0n1p3, physical 1256062976: metadata leaf (level 0) in tree 7
[   77.889711] BTRFS error (device nvme0n1p3): scrub: unable to fixup (regular) error at logical 173932544 on dev /dev/nvme0n1p3 physical 1256062976
[   77.889714] BTRFS warning (device nvme0n1p3): scrub: generation error at logical 173932544 on dev /dev/nvme0n1p3, physical 1256062976: metadata leaf (level 0) in tree 7
[   77.889717] BTRFS error (device nvme0n1p3): bdev /dev/nvme0n1p3 errs: wr 0, rd 0, flush 0, corrupt 2, gen 6
[   81.904947] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[   81.906694] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[   82.024805] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[   82.026672] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[   82.027642] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[   82.030183] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[   88.354692] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[   88.356191] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[   89.457652] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[   89.458952] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  130.714741] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[  130.715609] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  130.885315] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[  130.885980] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  131.017530] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[  131.017837] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  131.127023] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[  131.128719] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  131.244445] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[  131.245293] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  135.715738] btrfs_validate_extent_buffer: 150 callbacks suppressed
[  135.715741] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[  135.715877] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  135.802637] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[  135.805471] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  135.838462] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[  135.839846] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  135.910646] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[  135.912990] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  136.020509] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 1 wanted 63489 found 63386
[  136.021248] BTRFS error (device nvme0n1p3): parent transid verify failed on logical 173948928 mirror 2 wanted 63489 found 63386
[  435.270459] BTRFS info (device nvme0n1p3): scrub: finished on devid 1 with status: 0

I am willing to try destructive options like wiping the drive and reinstalling everything. But obviously avoiding this is less hassle and more importantly I want to make sure whether my hardware is failing or not.

I think that’s the big question. Usually, filesystems don’t just randomly corrupt themselves except for bugs (and BTRFS should be sufficiently bug-free) and hardware issues. So ruling out any hardware issues would be my first step.

Have you had to force power off the system?
You may have failing hardware. It’s worth checking how well the nvme is seated.

What does sudo smartctl -x /dev/nvme0n1 report?

Any good idea on how? Do you think reseating and or placing the NVMe into a different slot is a good first approach? What do I have to change so it is still correctly mounted in that case?

I do have to do force power off occasionally. For some new game releases and kernel/mesa combination my system tends to run into hard freezes. I am talking about every few months maybe, and last time has been a few months ago. So nothing recent.

This is the output:

smartctl 7.5 2025-04-30 r5714 [x86_64-linux-6.19.10-300.fc44.x86_64] (local build)
Copyright (C) 2002-25, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Number:                       XPG GAMMIX S11 Pro
Serial Number:                      2L31291S1DHC
Firmware Version:                   32B3T8ED
PCI Vendor/Subsystem ID:            0x1cc1
IEEE OUI Identifier:                0x707c18
Controller ID:                      1
NVMe Version:                       1.3
Number of Namespaces:               1
Namespace 1 Size/Capacity:          1,024,209,543,168 [1.02 TB]
Namespace 1 Utilization:            667,782,488,064 [667 GB]
Namespace 1 Formatted LBA Size:     512
Local Time is:                      Tue Sep  8 20:24:48 2026 UTC
Firmware Updates (0x14):            2 Slots, no Reset required
Optional Admin Commands (0x0017):   Security Format Frmw_DL Self_Test
Optional NVM Commands (0x005f):     Comp Wr_Unc DS_Mngmt Wr_Zero Sav/Sel_Feat Timestmp
Log Page Attributes (0x0e):         Cmd_Eff_Lg Ext_Get_Lg Telmtry_Lg
Maximum Data Transfer Size:         64 Pages
Warning  Comp. Temp. Threshold:     75 Celsius
Critical Comp. Temp. Threshold:     80 Celsius

Supported Power States
St Op     Max   Active     Idle   RL RT WL WT  Ent_Lat  Ex_Lat
 0 +     9.00W       -        -    0  0  0  0        0       0
 1 +     4.60W       -        -    1  1  1  1        0       0
 2 +     3.80W       -        -    2  2  2  2        0       0
 3 -   0.0450W       -        -    3  3  3  3     2000    2000
 4 -   0.0040W       -        -    4  4  4  4    15000   15000

Supported LBA Sizes (NSID 0x1)
Id Fmt  Data  Metadt  Rel_Perf
 0 +     512       0         0

=== START OF SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

SMART/Health Information (NVMe Log 0x02, NSID 0xffffffff)
Critical Warning:                   0x00
Temperature:                        36 Celsius
Available Spare:                    100%
Available Spare Threshold:          10%
Percentage Used:                    18%
Data Units Read:                    186,050,618 [95.2 TB]
Data Units Written:                 133,991,289 [68.6 TB]
Host Read Commands:                 1,329,925,680
Host Write Commands:                1,241,262,139
Controller Busy Time:               29,061
Power Cycles:                       2,335
Power On Hours:                     12,709
Unsafe Shutdowns:                   18
Media and Data Integrity Errors:    0
Error Information Log Entries:      1
Warning  Comp. Temperature Time:    0
Critical Comp. Temperature Time:    0
Thermal Temp. 1 Transition Count:   351
Thermal Temp. 1 Total Time:         1654

Error Information (NVMe Log 0x01, 16 of 256 entries)
No Errors Logged

Self-test Log (NVMe Log 0x06, NSID 0xffffffff)
Self-test status: No self-test in progress
No Self-tests Logged

You should run smartctl -t long on that device and then look at the results again. Your current output shows no tests ever run. The self tests probably would show if there were hardware errors. Running smartctl -h will show the many ways that smartctl may be used.

It also shows no specific errors logged so it may not be the drive itself and the filesystem errors might be a result of the forced power-off events.

The repeated freezes could easily be memory issues as well so running memtest86+ might be a wise choice as well.

Unfortunately, I have never had to debug any NVMe issues, so I wouldn’t know where to start.

It probably won’t hurt.

Nomally, Fedora identifies its partitions via UUIDs and subvolume names (for BTRFS), which are stable across different NVMe slots. You shouldn’t need any changes.

However, with the errors your filesystem currently has, I would dispute that it mounts correctly in the slot that it currently is in. :wink: You would need to fix those in some way before putting it in another m.2 would give you any useful information.

While a hard reboot a few months ago is probably not the reason for the corruption you are seeing today, it might be a good idea to look into enabling the Magic SysRq key, for a more controlled reboot/shutdown.

I have mine set to 176, a combination of 128, 64, and 32:

$ cat /etc/sysctl.d/90-sysrq.conf
kernel.sysrq = 176

With these three enabled, I can hold down Alt+PrtScr and then press S (sync filesystems), U (remount r/o), and either B (reboot) or O (poweroff) in slow succession (give it a few seconds between keys so it can do what it needs to do).

Note that anybody can press these keys, so if you often leave your system unattended in public places, anyone can then force poweroff your system with a quick Alt+PrtScr + O (though, to be fair, they could achieve the same thing by just holding the power button).

I did the short (a few seconds) and the long (a few minutes) test. It still looks the same to me :confused:


smartctl 7.5 2025-04-30 r5714 [x86_64-linux-6.19.10-300.fc44.x86_64] (local build)
Copyright (C) 2002-25, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Number:                       XPG GAMMIX S11 Pro
Serial Number:                      2L31291S1DHC
Firmware Version:                   32B3T8ED
PCI Vendor/Subsystem ID:            0x1cc1
IEEE OUI Identifier:                0x707c18
Controller ID:                      1
NVMe Version:                       1.3
Number of Namespaces:               1
Namespace 1 Size/Capacity:          1,024,209,543,168 [1.02 TB]
Namespace 1 Utilization:            667,782,488,064 [667 GB]
Namespace 1 Formatted LBA Size:     512
Local Time is:                      Wed Sep  9 05:11:57 2026 UTC
Firmware Updates (0x14):            2 Slots, no Reset required
Optional Admin Commands (0x0017):   Security Format Frmw_DL Self_Test
Optional NVM Commands (0x005f):     Comp Wr_Unc DS_Mngmt Wr_Zero Sav/Sel_Feat Timestmp
Log Page Attributes (0x0e):         Cmd_Eff_Lg Ext_Get_Lg Telmtry_Lg
Maximum Data Transfer Size:         64 Pages
Warning  Comp. Temp. Threshold:     75 Celsius
Critical Comp. Temp. Threshold:     80 Celsius

Supported Power States
St Op     Max   Active     Idle   RL RT WL WT  Ent_Lat  Ex_Lat
 0 +     9.00W       -        -    0  0  0  0        0       0
 1 +     4.60W       -        -    1  1  1  1        0       0
 2 +     3.80W       -        -    2  2  2  2        0       0
 3 -   0.0450W       -        -    3  3  3  3     2000    2000
 4 -   0.0040W       -        -    4  4  4  4    15000   15000

Supported LBA Sizes (NSID 0x1)
Id Fmt  Data  Metadt  Rel_Perf
 0 +     512       0         0

=== START OF SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

SMART/Health Information (NVMe Log 0x02, NSID 0xffffffff)
Critical Warning:                   0x00
Temperature:                        34 Celsius
Available Spare:                    100%
Available Spare Threshold:          10%
Percentage Used:                    18%
Data Units Read:                    186,051,190 [95.2 TB]
Data Units Written:                 133,991,289 [68.6 TB]
Host Read Commands:                 1,329,934,661
Host Write Commands:                1,241,262,141
Controller Busy Time:               29,061
Power Cycles:                       2,336
Power On Hours:                     12,709
Unsafe Shutdowns:                   18
Media and Data Integrity Errors:    0
Error Information Log Entries:      1
Warning  Comp. Temperature Time:    0
Critical Comp. Temperature Time:    0
Thermal Temp. 1 Transition Count:   351
Thermal Temp. 1 Total Time:         1654

Error Information (NVMe Log 0x01, 16 of 256 entries)
No Errors Logged

Self-test Log (NVMe Log 0x06, NSID 0xffffffff)
Self-test status: No self-test in progress
Num  Test_Description  Status                       Power_on_Hours  Failing_LBA  NSID Seg SCT Code
 0   Extended          Completed without error               12709            -     -   -   -    -
 1   Short             Completed without error               12709            -     -   -   -    -
 2   Short             Completed without error               12709            -     -   -   -    -

Well it does actually “boot correctly”, it is shortly after that everything falls apart :joy:. I get your point.

I am pretty sure I actually set up the Magic SysRq key at some point, but my system was not responding to it in those gpu freeze cases.

I also ran memtest for 80 min (for now).

Does that mean my ram is for sure broken?
I will try to reseat both ram and the nvme to see if this fixes anything.

Edit:
It actually crashed on test 8 …

It might not be bad RAM. What stands out to me is the frequency, 3000 MT/s. The 5800X should handle 3200MT/s fine if your RAM supports it. I would repeat the test with standard JEDEC timing (2133MT/s for DDR4, IIRC). If that completes without errors, I would re-seat the RAM and then increase speeds slowly while verifying every new frequency/timing. Once errors start to occur, you might have reached the end of your memory controller and/or RAM.

For example, I have a 9900X and DDR5 6000, but the combination of the CPU, RAM and mainboard is only stable for me at 5900MT/s.

From that output there are no warning signs.

I think it likely the forced power off has lead to the corruption.

When i have seen my system freeze i have been able to ssh in.

Do you happen to have a second system you can you to ssh into your gaming system?
If so next time the system hangs ssh in and tell it to reboot.

First of all, I would mount the drive read-only (-o ro) and backup all important data.

It could be that even if the last hard-reboot was months ago - btrfs does not scan or verify every block during normal day-to-day usage. Only a btrfs scrub would reveal the corruption

It is also possible that you RAM is defective (or overclocked) since the RAM test is failing - maybe a bit flip introduced the corruption.

As for repairing your metadata corruption, I can’t really help, I don’t use btrfs. However, I recommend not to blindly use btrfs check --repair. Executing a repair on a broken metadata tree can permanently delete unreferenced files. First, check if the filesystem structure can be recovered: sudo btrfs check /dev/nvme0n1p3

Or bad connector, damage due to mishandling (static electricity), bad system board. If there are multiple ram modules see if test failures follow the module when places are swapped. A marginal RAM module may pass tests using lower clock speeds.

I reseated both ram and nvme. I also flipped the ram modules (same slots).

Still fails memtest (even reproducing the freeze).

Then I removed the the XMP profile and it ran fine.

Enabled again and I triggered the freeze again.

Disabled profile and again did not run into the issue. I don’t really have the time to do more in-depth testing right now, though.

Let me try to summarize:
My filesystem is corrupted and there is really no way around wiping it and reinstalling fedora.
Should I stay on btrfs or go with ext4? Would there be a difference in how I could have handled the issue if I was on ext4?

Now obviously something triggered the corruption.
Either one of the freezes/crashes/power downs or the flaky ram.

For the RAM I will disable the overclock and run a longer mem stress test. If successful I will leave it like that or tinker with the overclock until I have something stable.
If that fails I might consider buying a new kit (but I really don’t want to. Could be CPU/mainboard as well, right :smiling_face_with_tear:?)

For the other issue, I will make sure that my magic sysreq key is properly set up and I will setup my laptop and/or phone so I can reboot via ssh from system from there.

Any other suggestion for where to go from here?
Thanks a lot for all the replies :heart:

If you are overclocking RAM, and it is causing Memtest to fail, it can definitely a cause of disk corruption over time. It would also explain the random freeze occurrences.

Indeed. I had a problem like this in the past (RAM wasn’t reliable on the XMP profile and had to be downclocked a bit from that) and the problem was the CPU - the same RAM was fine with a newer CPU on the same board.

I don’t know if I would classify using the predefined XMP profile as overclocking, but I suppose that is what it is.
Random freezes 100% where the GPU (gfx ring timeout), but maybe the actual culprit was ram messing up whatever was uploaded to the GPU.

The system is running like that for 2 years and has been running with a different CPU for 3 more years before.
I do occasionally reinstall fedora, so maybe hardware degradation making the flakiness worse finally broke something in a way that I have to deal with instead of just restarting.

5 years is enough time for components to age out of spec, corrosion to affect connectors, etc. Some problems can be cured with a new power supply. It is common for large enterprises to replace systems having problems at 5+ years old.