Fedora Freezes Within 24 hours of Running - Please help

,

Hello wonderful people,

I am once again asking for some help. I am dealing with a situation that is causing me frustration and confusion, and I have hit the end of my knowledge of how to track down this issue(s).

For a bit of background I use Fedora KDE for my main machine for multiple months, slowly transitioning all my workflows from Windows happily, and that work means that I am not afraid of the terminal and know to read the man pages. However, this issue requires (in my opinion) using logs and then being able to understand these logs and well - I don’t know what logs or how to read what I have looked up. So I am happy to provide logs, please know that I may not know the commands required for the log you want.

This is a saga, so jump to the TLDR if you wish but if you need the history of things I have tried, it will be in the below wall of text.

I have a server that used to run Debian, after an update dialling into (via vnc) the server caused the colours to be inverted and I choose at that point to upgrade the server to Fedora KDE as eventually I would like to run it headless using Fedora Server, as well as Fedora working so well on my main machine.

This worked great, for months. It is primarily used as a NAS using Samba, however the thought is to run other solutions through it. The first thing that I tried to do is run a Davinci Project Server, which worked fine but I started to notice that the server CPU was struggling so decided to upgrade the hardware.

Old Hardware

  • Asus Radeon HD 6770
  • Asrock 990FX Extreme 3
  • 8gb Integral IN3T2GNZBII DDR3 PC3-1333MHz (4 x 2gb sticks)
  • AMD FX8350 CPU
  • Corsair CV650 - 650 Watt 80 Plus Bronze PSU
  • SATA SSD boot drive

New Hardware

  • Asus Radeon HD 6770
  • Asus B450-plus II TUF Gaming
  • 16gb Corsair Vengence DDR4 3000MHz Memory (2 x 8gb sticks)
  • AMD Ryzen 7 1700X Eight-Core Processor
  • Corsair CV650 - 650 Watt 80 Plus Bronze PSU
  • NVME boot Drive

As I was moving the install from a SSD to an NVME I decided to do a fresh install from a USB Live stick so did the following steps:

  1. Created a Live USB for Fedora KDE (44) using Fedora Media Writer
  2. Installed Fedora KDE (44) to the NVME drive
  3. Booted into the installed Fedora
  4. Shut down the machine
  5. Removed the Live USB
  6. Plugged in the old boot SSD
  7. Booted into the new Fedora Install
  8. Mounted the old boot SSD drive
  9. Copied the home folder over from the SSD to the NVME
  10. Shut Down
  11. Connected all Data Drives
  12. Edited fstab to mount those drives on boot
  13. Reboot to test
  14. Edited Samba file and other settings to get file sharing up and running
  15. Went to bed

The next day the computer was frozen, the Desktop Clock Widget was showing a time that was many hours in the past and the super key did not bring up the Application Launcher, so I pressed and held the power button down till it powered down, then as I had some issues with the copying of the home drive the previous night - I assumed I had caused this issue so did the following:

  1. Flipped the power button on the power supply to off
  2. Unplugged all drives, except NVME
  3. Flipped the power button on the power supply to on
  4. Did a fresh install from the Live USB
  5. Put a Desktop Widget, the Digital Clock, show the seconds Always
  6. Changed all settings I could find that would power down/hibernate the system or lock the account
  7. Left the Machine displayed on my second monitor and watched to see if it froze
  8. It did within ~24 hours

As the NVME drive was new - I suspected that first so did the following:

  1. Shut down
  2. Flipped the power button on the power supply to off
  3. Removed the NVME
  4. Flipped the power button on the power supply to on
  5. Booted to the Live USB
  6. Put a Desktop Widget, the Digital Clock, show the seconds Always
  7. Changed all settings I could find that would power down/hibernate the system or lock the account, so that it will remain powered and logged in
  8. Left the Machine displayed on my second monitor and watched to see if it froze
  9. It did within ~24 hours
  10. Repeated steps 1 to 7 with a different USB, it froze within the time window

So now I was starting to suspect that it is a hardware and software issue, but I installed the server originally from a ISO of 43 - luckily I am a pack rat and still have that download so did the following with Fedora KDE 43:

  1. Shut down
  2. Flipped the power button on the power supply to off
  3. Installed the NVME
  4. Flipped the power button on the power supply to on
  5. Did a fresh install from the Live USB
  6. Update and Upgrade so fully up to date
  7. Put a Desktop Widget, the Digital Clock, show the seconds Always
  8. Changed all settings I could find that would power down/hibernate the system or lock the account, so that it will remain powered and logged in
  9. Left the Machine displayed on my second monitor and watched to see if it froze
  10. It did within ~24 hours

Then:

  1. Shut down
  2. Booted to the Live USB
  3. Put a Desktop Widget, the Digital Clock, show the seconds Always
  4. Changed all settings I could find that would power down/hibernate the system or lock the account, so that it will remain powered and logged in
  5. Left the Machine displayed on my second monitor and watched to see if it froze
  6. It did within ~24 hours

Guess who has the old ISO of KDE 42? Yes me, and yes it did the same thing.

So… I downloaded old faithful Debian and did the following:

  1. Shut down
  2. Did a fresh install from the Live USB
  3. Update and Upgrade so fully up to date
  4. Put a Desktop Widget, the Digital Clock, show the seconds Always
  5. Changed all settings I could find that would power down/hibernate the system or lock the account, so that it will remain powered and logged in
  6. Left the Machine displayed on my second monitor and watched to see if it froze
  7. It did not freeze, even after 48 hours

So now I am in a confused puddle and starting to feel crazy. At this point it feels like it is an issue with something that Fedora has that Debian does not, does that mean that Debian will never have this issue or if I keep my Server updated, it will one day start displaying this issue?

I feel like I have to track down this issue and file a bug report somewhere, so I decide to try Fedora Workstation, wondering if it was a KDE issue.

Yes it froze within ~24 hours.

What next, well it is clearly a hardware compatibility issue so - look around at all my old gear at what can I do, replace the motherboard:

  • MSI X370 GAMING PRO CARBON

Freezing still happened within ~24 hours.

Put the new motherboard back and replace the GPU with a NVIDIA 1080 and install Fedora KDE 42, install the NVIDIA driver and leave it.

This is my current hardware:

  • NVIDIA GeForce GTX 1080
  • Asus B450-plus II TUF Gaming
  • 16gb Corsair Vengence DDR4 3000MHz Memory (2 x 8gb sticks)
  • AMD Ryzen 7 1700X Eight-Core Processor
  • Corsair CV650 - 650 Watt 80 Plus Bronze PSU
  • NVME Drive

With KDE Plasma Version: 6.3.4 and Fedora 42

Looking up this freezing issue, I saw that a question that was asked regularly is “Does the caps lock respond?”, not once did I check this so I am currently running it to find out - however about an hour into this it just rebooted on its own.

I do not have other RAM to test with, nor another AM4 CPU, without pulling my main machine apart, so I am hopeful one of you lovely people can help me. As an aside this RAM was in a machine running Fedora for months without the hint of an issue, so I feel safe ruling it out.

Whilst writing up this saga, the machine did freeze and I can confirm that caps lock does not respond. If I try to SSH into the system, I get the following error: “ssh: connect to host xxx.xxx.x.xxx port 22: No route to host”, I can confirm that I could connect via SSH prior to the freeze.

TLDR

I have a computer that through multiple Fedora installations will freeze and lock up, please help me find out why.

Thank you for taking the time to read and I may have missed some troubleshooting I have done, I am honestly exhausted - so please feel free to ask questions.

Mango

Maybe try running an overnight memtest on this.

I used to have the Ryzen 5 2400G (which is the same underlying architecture, I think) and it couldn’t run stably with memory at 3000 MHz.

Downclocking a bit to 2800 MHz fixed it to me. (The RAM itself was fine and had no problems running at the rated speed on other CPUs.)

This will require using the rpmfusion-nonfree repo to install the nvidia 580xx driver for support.

The suggestion by @pg-tips above seems quite reasonable for testing.

You may also need to verify the firmware (bios) on the motherboard is at the latest version. Sometimes newer CPUs may require updated bios to support newer features in those devices.

You can show us the current info for almost all your current hardware with running the command inxi -ezxx (inxi may need to be installed) and posting it here as preformatted text

I’m glad you mentione dthis, wasn’t there something Memory related with this generation? I think I remember Level1Techs having a video about it. Not sure if a BIOS update was the fix.

Yeah, I don’t remember the details, but I remember looking into it enough at the time to decide that it was a generic problem, and not just my bad luck to get a defective unit.

I don’t think subsequent BIOS updates fixed it, but the performance hit wasn’t big enough that I worried much about retesting it, so I might have missed something.

Eventually I swapped the CPU out for a 5700X, and then the RAM worked fine at its rated speed.

Some 1700 series were defective, well known problem. AMD replaced them. Nothing to do with os. Reddit - Please wait for verification

Hi P G, I forgot that I set the memory to 3000 in the BIOS, I have now chosen “auto” - as that is the only other option, which puts the memory to 2133 MHz - running a test to see if it freezes now. If it does, then I will do a memtest, but as said in my post it was in a machine before I upgraded and it ran Fedora fine, but it was not at 3000. :crossed_fingers:

Hi Jeff, I feel I did install the correct driver (but, I might be wrong), and as said the problem persisted with an AMD GPU first, however I can switch back to the AMD GPU if that will help diagnosis.

My BIOS is one version out of date, I am reticent to update the BIOS without confirmation that it will still support this CPU, as I have heard people encountering this issue.

Output of inxi -ezxx is:

System:
  Kernel: 6.19.14-108.fc42.x86_64 arch: x86_64 bits: 64 compiler: gcc v: 15.2.1
  Console: pty pts/1 DM: SDDM Distro: Fedora Linux 42 (KDE Plasma Desktop Edition)
Machine:
  Type: Desktop System: ASUS product: N/A v: N/A serial: N/A
  Mobo: ASUSTeK model: TUF GAMING B450-PLUS II v: Rev X.0x serial: <filter>
    part-nu: SKU Firmware: UEFI vendor: American Megatrends v: 4631 date: 01/14/2025
CPU:
  Info: 8-core model: AMD Ryzen 7 1700X bits: 64 type: MT MCP arch: Zen rev: 1 cache:
    L1: 768 KiB L2: 4 MiB L3: 16 MiB
  Speed (MHz): avg: 3400 min/max: 2200/3400 boost: enabled cores: 1: 3400 2: 3400
    3: 3400 4: 3400 5: 3400 6: 3400 7: 3400 8: 3400 9: 3400 10: 3400 11: 3400 12: 3400
    13: 3400 14: 3400 15: 3400 16: 3400 bogomips: 108593
  Flags-basic: avx avx2 ht lm nx pae sse sse2 sse3 sse4_1 sse4_2 sse4a ssse3
Graphics:
  Device-1: NVIDIA GP104 [GeForce GTX 1080] vendor: Gigabyte driver: nvidia
    v: 580.159.03 arch: Pascal pcie: speed: 2.5 GT/s lanes: 16 ports: active: HDMI-A-1
    empty: DP-1, DP-2, DP-3, DVI-D-1 bus-ID: 06:00.0 chip-ID: 10de:1b80
  Display: unspecified server: Xwayland v: 24.1.6 compositor: kwin_wayland driver:
    gpu: nvidia,nvidia-nvswitch tty: 88x40
  Monitor-1: HDMI-A-1 model: Acer KG271 res: 1920x1080 dpi: 82 diag: 686mm (27")
  API: EGL v: 1.5 platforms: device: 0 drv: nvidia gbm: drv: nvidia surfaceless:
    drv: nvidia inactive: wayland,x11
  API: OpenGL v: 4.6.0 vendor: nvidia v: 580.159.03 note: console (EGL sourced)
    renderer: NVIDIA GeForce GTX 1080/PCIe/SSE2
  API: Vulkan v: 1.4.304 surfaces: N/A device: 0 type: discrete-gpu driver: nvidia
    device-ID: 10de:1b80 device: 1 type: cpu driver: mesa llvmpipe device-ID: 10005:0000
  Info: Tools: api: clinfo, eglinfo, glxinfo, vulkaninfo
    de: kscreen-console,kscreen-doctor gpu: nvidia-settings wl: wayland-info
    x11: xdriinfo, xdpyinfo, xprop, xrandr
Audio:
  Device-1: NVIDIA GP104 High Definition Audio vendor: Gigabyte driver: N/A pcie:
    speed: 8 GT/s lanes: 16 bus-ID: 06:00.1 chip-ID: 10de:10f0
  Device-2: Advanced Micro Devices [AMD] Family 17h HD Audio vendor: ASUSTeK
    driver: N/A pcie: speed: 8 GT/s lanes: 16 bus-ID: 08:00.3 chip-ID: 1022:1457
  API: ALSA v: k6.19.14-108.fc42.x86_64 status: inactive
  Server-1: PipeWire v: 1.4.1 status: n/a (root, process) with: 1: pipewire-pulse
    status: active 2: wireplumber status: active 3: pipewire-alsa type: plugin
    4: pw-jack type: plugin
Network:
  Device-1: Realtek RTL8111/8168/8211/8411 PCI Express Gigabit Ethernet
    vendor: ASUSTeK driver: r8169 v: kernel pcie: speed: 2.5 GT/s lanes: 1 port: f000
    bus-ID: 03:00.0 chip-ID: 10ec:8168
  IF: enp3s0 state: up speed: 1000 Mbps duplex: full mac: <filter>
Drives:
  Local Storage: total: 238.47 GiB used: 9.9 GiB (4.1%)
  ID-1: /dev/nvme0n1 vendor: Fanxiang model: S500Pro 256GB size: 238.47 GiB
    speed: 31.6 Gb/s lanes: 4 serial: <filter> temp: 36.9 C
Partition:
  ID-1: / size: 60 GiB used: 9.24 GiB (15.4%) fs: btrfs dev: /dev/nvme0n1p4
  ID-2: /boot size: 973.4 MiB used: 342.4 MiB (35.2%) fs: ext4 dev: /dev/nvme0n1p2
  ID-3: /boot/efi size: 598.8 MiB used: 19.3 MiB (3.2%) fs: vfat dev: /dev/nvme0n1p1
  ID-4: /home size: 168.89 GiB used: 309.4 MiB (0.2%) fs: btrfs dev: /dev/nvme0n1p5
Swap:
  ID-1: swap-1 type: zram size: 8 GiB used: 0 KiB (0.0%) priority: 100 dev: /dev/zram0
  ID-2: swap-2 type: partition size: 8 GiB used: 0 KiB (0.0%) priority: -1
    dev: /dev/nvme0n1p3
Sensors:
  System Temperatures: cpu: 45.5 C mobo: N/A gpu: nvidia temp: 38 C
  Fan Speeds (rpm): N/A
Info:
  Memory: total: 16 GiB available: 15.5 GiB used: 2.66 GiB (17.2%)
  Processes: 434 Power: uptime: 6m wakeups: 0 Init: systemd v: 257
    target: graphical (5) default: graphical
  Packages: pm: rpm pkgs: N/A note: see --rpm Compilers: gcc: 15.2.1 Shell: Bash
    v: 5.2.37 running-in: pty pts/1 inxi: 3.3.40

Hi Alf, I actually had one that got replaced 1800x - it is in my main machine it caused a Kernel Panic and would not boot. Why would this one work fine on Debian but not Fedora though?

Actually it was not limited to 1700. For instance, Early production batches of the Ryzen 1800X (manufactured before week 25 of 2017) are notorious for a silicon defect that causes Segfaults (SEG) under heavy, multi-threaded workloads. Do you have the product date, from a photo of the lid? No way to find out with software. Search for “ryzen 1800x faulty” on google for more details.

The inxi output seems to indicate the proper driver.
That can be verified with dnf list --installed \*nvidia\* and look for the 580xx in the package names.

Good news = System ran for just over 27 hours
Bad news = It rebooted on its own

Using the command journalctl | grep error I see the following:

Jul 16 19:21:55 SERVER kernel: x86/amd: Previous system reset reason [0x08000800]: an uncorrected error caused a data fabric sync flood event

Which has led me to this page: AMD Ryzen Zen 4: random reboots caused by data fabric sync flood (0x08000800) — investigation and fix · GitHub

But, honestly I am baffled and confused by what those smart people are saying.

As promised, I will run a memtest overnight - but is there any other log(s) that would be helpful?

Looks like I installed the wrong one, however it is displaying on my monitor fine (Linux Magic) lol.

Installed packages
akmod-nvidia.x86_64                        3:580.159.03-1.fc42 rpmfusion-nonfree-nvidia-
kmod-nvidia-6.19.14-108.fc42.x86_64.x86_64 3:580.159.03-1.fc42 @commandline
nvidia-gpu-firmware.noarch                 20250311-1.fc42     781e4eb56ba449a5876af2cc0
nvidia-modprobe.x86_64                     3:580.159.03-1.fc42 rpmfusion-nonfree-nvidia-
nvidia-settings.x86_64                     3:580.159.03-1.fc42 rpmfusion-nonfree-nvidia-
xorg-x11-drv-nvidia.x86_64                 3:580.159.03-1.fc42 rpmfusion-nonfree-nvidia-
xorg-x11-drv-nvidia-cuda-libs.x86_64       3:580.159.03-1.fc42 rpmfusion-nonfree-nvidia-
xorg-x11-drv-nvidia-kmodsrc.x86_64         3:580.159.03-1.fc42 rpmfusion-nonfree-nvidia-
xorg-x11-drv-nvidia-libs.x86_64            3:580.159.03-1.fc42 rpmfusion-nonfree-nvidia-
xorg-x11-drv-nvidia-power.x86_64           3:580.159.03-1.fc42 rpmfusion-nonfree-nvidia-

Will work out how to get the right driver installed, thanks for the prompt - however I don’t think this is causing my stability issues?

You don’t need to (and indeed can’t) switch to the specific 580xx packages until you upgrade to Fedora 44.

On F42 and F43, the “mainline” RPM Fusion packages will keep you at driver version 580, which is fine for your GPU.

Try these kernel parameters:

sudo grubby --update-kernel=ALL --args="processor.max_cstate=2 nvme_core.default_ps_max_latency_us=0 pcie_aspm=off pcie_ports=native amdgpu.ppfeaturemask=0xfff73fff"

Be aware that these will limit the energy savings that your processor can achieve; if you have a laptop this might be an issue for you. If you have a desktop or never rely on battery power, you may not care in the least.

I missed the fact you are using f42 which is EOL.

Most users usually upgrade versions before the one they are running becomes EOL so it may be wise to do that soon. (EOL versions do not receive any new updates, bug fixes, or security fixes)

The upgrade can easily be done following this doc, and you can choose to upgrade to either f43 or f44. If upgrading to f44 then my comment about selecting the 580xx driver version for nvidia will be applicable.

Actually you can do so.
Simply enable the rpmfusion-nonfree repos then run sudo dnf swap akmod-nvidia akmod-nvidia-580xx --releasever=44 --allowerasing

As part of all this I ended up on f42 and as others have pointed out, I will be updating - I was only on 42 as a poor mans trying to troubleshoot. Will your suggestion be fine on the newer version? Sorry if this is a dumb question, but best to ask - just incase.

Jeff, if I do:

Then do the update and upgrade, will the correct driver stay in place? Do you know?

If you perform that swap so the 580xx packages are installed before you upgrade then, no, the newer driver will not replace it. The newest driver (610.43.03) is the one currently provided by the akmod-nvidia package from the rpmfusion-nonfree-updates repo, but the akmod-nvidia-580xx package does not get replaced with updates (the name is different).

Yep - it won’t do anything different on 42, 43 or 44, or 45.

It’s the non-atomic version of the “fixes” listed in the article you found by others digging into what these fabric errors are caused by.

Parameter Effect
processor.max_cstate=2 Limits CPU to C1/C2, disables CC6 deep sleep
nvme_core.default_ps_max_latency_us=0 Disables NVMe autonomous power state transitions (APST)
pcie_aspm=off Disables PCIe Active State Power Management entirely
pcie_ports=native OS takes full control of PCIe ports including error reporting
amdgpu.ppfeaturemask=0xfff73fff Disables GFXOFF (bit 15), ACG (bit 16), and stutter mode (bit 17)

If they make no difference for you, it’s a very similar command to remove them - think of them as switches provided to the kernel when it starts, about how to treat power management. In this case it restricts the power savings which could take place to turn some of them off, and that appears to have some influence on the fabric errors causing your crash.