BTRFS issues, check stopped by kernel OOM killer

Recently my Fedora 43 refused to boot, saying it couldn’t find the root drive. Eventually after a few reboots it booted normally, but since then has been sluggish at times, occasionally locking up for 10 seconds or so when loading a program like Firefox. When a check is run on the BTRFS filesysem, it gets to “[3/8] checking extents”, then the btrfs check process consumes a massive amount of memory and gets stopped by the kernel OOM killer. The filesystem is 1TB and I have 24GB of RAM.
The SSD the filesystem is on doesn’t seem to be failing, read and write speeds are normal and a SMART test doesn’t show anything.

It may help to have the output from inxi -ezxx so we have a more complete picture of your system hardware and configuration. Some drives have firmware updates, so worth checking fwupdmgr get-updates. In cases like this it is important have good backups of critical data in case recovery attempts fail. I like to make an image of the problem drive so I can get back to the same starting point after a recovery attempt fails.

Thank you for your reply. Here is the output from inxi -ezxx:

System:
Kernel: 6.18.42-200.fc43.x86_64 arch: x86_64 bits: 64 compiler: gcc
v: 15.3.1
Desktop: KDE Plasma v: 6.7.3 tk: Qt v: N/A wm: kwin_wayland dm: SDDM
Distro: Fedora Linux 43 (KDE Plasma Desktop Edition)
Machine:
Type: Desktop System: Gigabyte product: X570 UD v: N/A
serial:
Mobo: Gigabyte model: X570 UD serial: Firmware: UEFI
vendor: American Megatrends LLC. v: F40d date: 09/02/2024
Battery:
Device-1: hid-20:91:df:4d:95:80-battery model: Magic Trackpad serial: N/A
charge: N/A status: discharging
CPU:
Info: 6-core model: AMD Ryzen 5 5500 bits: 64 type: MT MCP arch: Zen 3
rev: 0 cache: L1: 384 KiB L2: 3 MiB L3: 16 MiB
Speed (MHz): avg: 2391 min/max: 411/4269 boost: enabled cores: 1: 2391
2: 2391 3: 2391 4: 2391 5: 2391 6: 2391 7: 2391 8: 2391 9: 2391 10: 2391
11: 2391 12: 2391 bogomips: 86234
Flags-basic: avx avx2 ht lm nx pae sse sse2 sse3 sse4_1 sse4_2 sse4a
ssse3 svm
Graphics:
Device-1: Advanced Micro Devices [AMD/ATI] Navi 33 [Radeon RX 7600/7600
XT/7600M XT/7600S/7700S / PRO W7600] vendor: Sapphire driver: amdgpu
v: kernel pcie: speed: 16 GT/s lanes: 8 ports: active: DP-3 off: HDMI-A-1
empty: DP-1,DP-2,Writeback-1 bus-ID: 03:00.0 chip-ID: 1002:7480
Display: wayland server: X.org v: 1.21.1.24 with: Xwayland v: 24.1.13
compositor: kwin_wayland driver: X: loaded: modesetting
alternate: fbdev,vesa dri: radeonsi gpu: amdgpu display-ID: 0
Monitor-1: DP-3 model: BenQ EW3270U res: 3840x2160 hz: 60 dpi: 140
diag: 801mm (31.5")
Monitor-2: HDMI-A-1 model: LG (GoldStar) FULL HD res: 1920x1080 dpi: 102
diag: 551mm (21.7")
API: EGL v: 1.5 platforms: device: 0 drv: radeonsi device: 1 drv: swrast
gbm: drv: kms_swrast surfaceless: drv: radeonsi wayland: drv: radeonsi x11:
drv: radeonsi
API: OpenGL v: 4.6 compat-v: 4.5 vendor: amd mesa v: 25.3.6 glx-v: 1.4
direct-render: yes renderer: AMD Radeon RX 7600 (radeonsi navi33 LLVM
21.1.8 DRM 3.64 6.18.42-200.fc43.x86_64) device-ID: 1002:7480
display-ID: :0.0
API: Vulkan v: 1.4.341 surfaces: N/A device: 0 type: discrete-gpu
driver: mesa radv device-ID: 1002:7480 device: 1 type: cpu
driver: mesa llvmpipe device-ID: 10005:0000
Info: Tools: api: clinfo, eglinfo, glxinfo, vulkaninfo
de: kscreen-console,kscreen-doctor gpu: corectrl, nvidia-settings,
nvidia-smi, radeontop wl: wayland-info x11: xdriinfo, xdpyinfo, xprop,
xrandr
Audio:
Device-1: Advanced Micro Devices [AMD/ATI] Navi 31 HDMI/DP Audio
driver: snd_hda_intel v: kernel pcie: speed: 16 GT/s lanes: 8
bus-ID: 03:00.1 chip-ID: 1002:ab30
Device-2: Advanced Micro Devices [AMD/ATI] Renoir/Cezanne HDMI/DP Audio
driver: snd_hda_intel v: kernel pcie: speed: 8 GT/s lanes: 16
bus-ID: 0c:00.1 chip-ID: 1002:1637
Device-3: Advanced Micro Devices [AMD] Ryzen HD Audio vendor: Gigabyte
driver: snd_hda_intel v: kernel pcie: speed: 8 GT/s lanes: 16
bus-ID: 0c:00.6 chip-ID: 1022:15e3
API: ALSA v: k6.18.42-200.fc43.x86_64 status: kernel-api
Server-1: PipeWire v: 1.4.11 status: active with: 1: pipewire-pulse
status: active 2: wireplumber status: active 3: pipewire-alsa type: plugin
4: pw-jack type: plugin
Network:
Device-1: Realtek RTL8111/8168/8211/8411 PCI Express Gigabit Ethernet
vendor: Gigabyte driver: r8169 v: kernel pcie: speed: 2.5 GT/s lanes: 1
port: e000 bus-ID: 07:00.0 chip-ID: 10ec:8168
IF: enp7s0 state: up speed: 1000 Mbps duplex: full mac:
Device-2: Microsoft Xbox 360 Wireless Adapter driver: N/A type: USB
rev: 2.0 speed: 12 Mb/s lanes: 1 bus-ID: 5-1:2 chip-ID: 045e:0719
IF-ID-1: tailscale0 state: unknown speed: -1 duplex: full mac: N/A
Bluetooth:
Device-1: Cambridge Silicon Radio Bluetooth Dongle (HCI mode) driver: btusb
v: 0.8 type: USB rev: 2.0 speed: 12 Mb/s lanes: 1 bus-ID: 3-4:2
chip-ID: 0a12:0001
Report: btmgmt ID: hci0 rfk-id: 3 state: up address: bt-v: 4.0
lmp-v: 6
RAID:
Device-1: lager type: zfs status: ONLINE level: mirror-0 raw: size: 7.27 TiB
free: 1.35 TiB zfs-fs: size: 7.14 TiB free: 1.22 TiB
Components: Online: 1: sdc 2: sdd
Drives:
Local Storage: total: 18.08 TiB used: 7.54 TiB (41.7%)
ID-1: /dev/nvme0n1 vendor: Samsung model: SSD 970 EVO Plus 1TB
size: 931.51 GiB speed: 31.6 Gb/s lanes: 4 serial: temp: 41.9 C
ID-2: /dev/nvme1n1 vendor: Samsung model: SSD 990 EVO Plus 2TB
size: 1.82 TiB speed: 126 Gb/s lanes: 4 serial: temp: 37.9 C
ID-3: /dev/sda vendor: Samsung model: SSD 860 EVO 500GB size: 465.76 GiB
speed: 6.0 Gb/s serial:
ID-4: /dev/sdb vendor: Samsung model: SSD 850 EVO 250GB size: 232.89 GiB
speed: 6.0 Gb/s serial:
ID-5: /dev/sdc vendor: Seagate model: ST8000VN004-3CP101 size: 7.28 TiB
speed: 6.0 Gb/s serial:
ID-6: /dev/sdd vendor: Seagate model: ST8000VN004-3CP101 size: 7.28 TiB
speed: 6.0 Gb/s serial:
ID-7: /dev/sdf vendor: Silicon Power model: SPCC M.2 SSD size: 119.24 GiB
type: USB rev: 3.1 spd: 5 Gb/s lanes: 1 serial:
Partition:
ID-1: / size: 1.81 TiB used: 1.63 TiB (90.0%) fs: btrfs dev: /dev/nvme1n1p3
ID-2: /boot size: 3.83 GiB used: 821.2 MiB (20.9%) fs: ext4
dev: /dev/nvme1n1p2
ID-3: /boot/efi size: 598.8 MiB used: 19.3 MiB (3.2%) fs: vfat
dev: /dev/nvme1n1p1
ID-4: /home size: 1.81 TiB used: 1.63 TiB (90.0%) fs: btrfs
dev: /dev/nvme1n1p3
Swap:
ID-1: swap-1 type: file size: 16 GiB used: 0 KiB (0.0%) priority: -2
file: /swap/swapfile
ID-2: swap-2 type: zram size: 8 GiB used: 1.37 GiB (17.1%) priority: 100
dev: /dev/zram0
Sensors:
System Temperatures: cpu: 55.4 C mobo: 30.0 C gpu: amdgpu temp: 60.0 C
mem: 58.0 C
Fan Speeds (rpm): N/A gpu: amdgpu fan: 0
Info:
Memory: total: 24 GiB available: 23.32 GiB used: 7.72 GiB (33.1%)
Processes: 623 Power: uptime: 1h 27m wakeups: 0 Init: systemd v: 258
default: graphical
Packages: pm: rpm pkgs: N/A note: see --rpm pm: flatpak pkgs: 74
Compilers: clang: 21.1.8 alt: 17/18 gcc: 15.3.1 Shell: pk-command-not
running-in: konsole inxi: 3.3.41

Note that the BTRFS drive is actually the 2TB SSD not 1TB as stated earlier. No firmware updates are available for this drive.

This:

Plus

Would encourage me to not trust this drive for anything important. I’d be tempted to reformat it and restore it from a backup.

That kernel is far behind the current 7.1.8

Have you read https://btrfs.readthedocs.io/en/latest/Swapfile.html? I agree with the other responses that suggest a fresh install. Given the current rash of AI generated exploits, keeping current is more important than in the past.

,

Be aware of the fact that f43 is soon to be reaching EOL (near the end of November). If you decide to reinstall in solving this issue it may be prudent to install f44 instead and avoid the need to upgrade later when f43 reaches EOL.

I suggest btrfs scrub start / and check it with watch btrfs scrub status /

I would also like to see some kernel logs, e.g. journalctl -k --since=-10d --no hostname -o short-monotonic > klog.log

btrfs check --mode=lowmem will likely solve the oomkiller problem but it will be quite a lot slower.

I would like to preserve all my settings if I reinstall, would using rsync -a to copy all the hidden folders in the home directory be the best way to do this?

The older kernel is so I can use ZFS on my data HDDs without the openZFS module compatibility being disrupted by kernel updates. Perhaps I should move these to a BSD VM or separate NAS.

A scrub found no errors. I tried running btrfs check --mode=lowmem overnight, while it didn’t freeze it also didn’t complete in ~11 hours.

Here is from the klog.log file: [ 3.539824] kernel: Linux version 6.18.42-200.fc43.x86_64 (mockbuild@455c35aa - Pastebin.com

rsync -a is one reliable way to do it, yes

It’s also possible to replicate from one Btrfs to another Btrfs using read-only snapshots and btrfs send and btrfs receive. I think the differences are nuanced. The btrfs send/receive might be very slightly faster for a full replication. Subsequent replications are definitely faster mainly because no deep traversal is required on either the source or destination file systems to determine what has changed. In effect a sort of change log is very efficiently computed from file system metadata between any two snapshots.

That is a long time. What do you get for btrfs fi us / ? I’m just curious how much storage is allocated.

Scrub verifies data and metadata blocks compared to a stored checksum. It’ll tell us if things have changed since they were written. But it doesn’t tell us if the file system is consistent, only btrfs check can do that. It is slow in this low memory mode.

Here is from the klog.log file: [ 3.539824] kernel: Linux version 6.18.42-200.fc43.x86_64 (mockbuild@455c35aa - Pastebin.com

Is this just for the most recent boot? I don’t see anything in it suggestive of any delays mounting or problems with the file system or the storage device. If I were having this problem I’d like to look at dmesg/kmesg for the time frame the problem is occurring. It might be there are no Btrfs issues but maybe the kernel notices some issue with the storage device.

If this is strictly a latency problem then finding out what part of the kernel it’s occurring in is trickier. I’ve used bcc-tools for this (it’s in the repos) and more info is GitHub - iovisor/bcc: BCC - Tools for BPF-based Linux IO analysis, networking, monitoring, and more · GitHub including the various layers you’d want to look at like biolatency for the device itself (independent of VFS or the file system). You’d likely want to start at a “lower” level like device and work your way to VFS (e.g. filesslower) and then to Btrfs (e.g. btrfsslower).

Here is the output of btrfs fi us /

Overall:
Device size: 1.81TiB
Device allocated: 1.68TiB
Device unallocated: 137.58GiB
Device missing: 0.00B
Device slack: 0.00B
Used: 1.44TiB
Free (estimated): 333.80GiB (min: 265.01GiB)
Free (statfs, df): 333.80GiB
Data ratio: 1.00
Metadata ratio: 2.00
Global reserve: 512.00MiB (used: 0.00B)
Multiple profiles: no

Data,single: Size:1.58TiB, Used:1.39TiB (87.88%)
/dev/nvme0n1p3 1.58TiB

Metadata,DUP: Size:51.00GiB, Used:27.71GiB (54.32%)
/dev/nvme0n1p3 102.00GiB

System,DUP: Size:8.00MiB, Used:224.00KiB (2.73%)
/dev/nvme0n1p3 16.00MiB

Unallocated:
/dev/nvme0n1p3 137.58GiB

I have a large amount of storage used by games so I will delete those and see if btrfs check will work again.

I can’t see anything relevant in dmesg/kmesg, the only thing that shows up after boot is “[ 423.140033] evm: overlay not supported”

btrfs send looks very handy, I think at this point I will use that to take a complete backup of root to an image and either reinstall or format the drive and then copy back to the reformatted drive. I already have a copy of the home folder taken by rsync too.

I’m skeptical this is a kernel or file system problem. Only a btrfs check can tell us for certain but in theory the file system is always consistent.

@hepha3stus has 10x more metadata than my systerm, and I have less RAM, 16G. Normal check takes less than a minute. Lowmem mode was working for 20 minutes at which point I decided I had other things to do. :upside_down_face:

A possible work around for the OOM might be to setup conventional disk based swap if you have a partition to spare for it. Enabling zswap, while disabling zram might improve this type of workload’s performance since it works on a LRU basis to push less used pages to disk and keep most used pages in memory. Whereas swap on zram alone in tight memory situations will lead to OOM.

Since 10x more metadata is substantial, I can’t say if lowmem mode will beat normal mode with a lot of disk based swap available.