Bug 14052 - Knot Resolver on IPFire Core Update 205 crashes when it is shut down.
Summary: Knot Resolver on IPFire Core Update 205 crashes when it is shut down.
Status: VERIFIED
Alias: None
Product: IPFire
Classification: Unclassified
Component: --- (show other bugs)
Version: 2
Hardware: all Unspecified
: - Unknown - Aesthetic Issue
Assignee: Michael Tremer
QA Contact:
URL:
Keywords:
Depends on:
Blocks:
 
Reported: 2026-10-03 11:52 UTC by Adam Gibbons
Modified: 2026-10-09 13:31 UTC (History)
3 users (show)

See Also:


Attachments
/var/log/messages (23.60 KB, text/plain)
2026-10-06 16:00 UTC, pscar13
Details
/var/log/messages (17.71 KB, text/plain)
2026-10-06 19:13 UTC, pscar13
Details

Note You need to log in before you can comment on or make changes to this bug.
Description Adam Gibbons 2026-10-03 11:52:05 UTC
Component
=========

knot-resolver 6.4.2

Steps
=====

Reboot the machine.
Or, on a running system, terminate one worker:

shkill -TERM $(supervisorctl -s unix:///run/knot-resolver/supervisord.sock pid kresd:kresd7)

Expected
========

The worker exits cleanly.

Actual
======

The worker dies with SIGSEGV at offset 0x27603. On reboot every worker does this. A single terminate only crashes that worker, and Supervisor starts a replacement.

Sep 29 19:46:32 shutdown: shutting down for system reboot
Sep 29 19:46:32 init: Switching to runlevel: 6
Sep 29 19:46:55 supervisord: received SIGTERM indicating exit request
Sep 29 19:46:56 kernel: kresd[2481]: segfault ... in kresd[27603,...]
Sep 29 19:46:56 supervisord: stopped: kresd0 through kresd7 (terminated by SIGSEGV)

The same line is in /var/log/messages after a manual terminate.
Comment 1 Adam Gibbons 2026-10-03 12:03:05 UTC
Edit...

Small spelling mistake on the last post the command should have been:

kill -TERM $(supervisorctl -s unix:///run/knot-resolver/supervisord.sock pid kresd:kresd7)

---

The following cause and fix is AI generated so perhaps take with a pinch of salt:

Cause
=====

src/patches/knot-resolver-6.5-keep-dst-ip.patch stores an io_handle_data_t in the UDP listener's handle->data. endpoint_close() still treats that pointer as a session and calls session2_close() on it. Shutdown then faults in session2_event_unwrap().

Fix
====

endpoint_close() needs the session from the wrapper. on_session2_handle_close() has the same assumption.
Comment 2 Michael Tremer 2026-10-06 09:17:56 UTC
Here is a fix for this problem:

> https://git.ipfire.org/?p=people/ms/knot-resolver.git;a=commitdiff;h=ed453515e8c880baa96a13e8ddef485f200fe6e2

It has also been updated in master/next:

> https://git.ipfire.org/?p=ipfire-2.x.git;a=commitdiff;h=284715c5b03a131feb2f2bef39c0e5fef264add2

Please let me know if this fixes the problem as this is currently the last remaining blocker for the release.
Comment 3 pscar13 2026-10-06 16:00:53 UTC
Created attachment 1760 [details]
/var/log/messages

Test CU204 Development Build: master/88a55eb0 

DNS Ko, see log messages errors

Oct  6 16:06:11 ipfireTest kresd[9994]: [system] assertion "(data->session->transport.type == SESSION2_TRANSPORT_IO && data->session->transport.io.handle == handle)" failed in on_session2_handle_close@../daemon/session2.c:1780 
Oct  6 16:06:11 ipfireTest supervisord: exited: kresd0 (terminated by SIGABRT; not expected)
Comment 4 Michael Tremer 2026-10-06 16:08:44 UTC
Thanks for the feedback. On the on_session2_handle_close part, how are you reproducing this? I have just following the advice from Adam's AI and have not been able to test the change.
Comment 5 pscar13 2026-10-06 16:11:05 UTC
The problem occurs immediately upon startup following installation or upgrade.
Comment 6 Michael Tremer 2026-10-06 16:13:06 UTC
Interesting. It didn't for me. Do you have a chance to revert the last commit, rebuild and check again?
Comment 7 pscar13 2026-10-06 16:22:58 UTC
(In reply to Michael Tremer from comment #6)
> Interesting. It didn't for me. Do you have a chance to revert the last
> commit, rebuild and check again?

Sorry, I don't know if you're addressing me, but I don't know how to do that on the IPFire repository.
My build platform is too slow to quickly perform a local rebuild.
Comment 8 Michael Tremer 2026-10-06 16:44:09 UTC
Yes, this was for you. I have reverted the commit (as it has not been a problem before), but the build will take a couple of hours to come through:

> https://git.ipfire.org/?p=ipfire-2.x.git;a=commitdiff;h=61c40dd31ccd7b3a8a7e0183ce3995a05a6fd3e6

This is only in next.
Comment 9 pscar13 2026-10-06 17:03:11 UTC
OK
I'll test your CU205 when the ISO is ready.

I'm also reverting on the master locally on my build system, but my build will take longer (> 3 hours).
Comment 10 pscar13 2026-10-06 19:13:42 UTC
Created attachment 1761 [details]
/var/log/messages

I've just finished local compil CU204 with a revert on master.

phili@PCBUREAU:~/projects/ipfire-2.x$ git status
On branch master
Your branch is behind 'origin/master' by 1 commit, and can be fast-forwarded.
  (use "git pull" to update your local branch)

nothing to commit, working tree clean
phili@PCBUREAU:~/projects/ipfire-2.x$ sudo ./make.sh build


I installed the ISO on my test VM
There are no more kresd error messages; here is the reboot log.
Comment 11 Michael Tremer 2026-10-06 19:14:45 UTC
Thank you very much, so I will revert the commit in the other places, too.
Comment 12 pscar13 2026-10-07 08:41:49 UTC
I just tested the installation of the x86_64 ISO for the CU205 Development Build (next/61c40dd3)
It works fine, and there are no more kresd error messages either.

Note: In VMware, I had to use the SATA HDD because installer couldn't find the SCSI drive.
Comment 13 Michael Tremer 2026-10-07 08:42:52 UTC
(In reply to pscar13 from comment #12)
> Note: In VMware, I had to use the SATA HDD because installer couldn't find
> the SCSI drive.

Could you file a separate ticket for this, please?
Comment 14 Adam Gibbons 2026-10-07 19:20:36 UTC
Confirmed on next/3e4b1ad3 that this issue is now resolved.
Comment 15 Adolf Belka 2026-10-08 17:04:28 UTC
(In reply to Adam Gibbons from comment #14)
> Confirmed on next/3e4b1ad3 that this issue is now resolved.

@Adam, can you also please confirm it on master/aefdb657 (CU204).
Comment 16 Adam Gibbons 2026-10-09 11:32:55 UTC
(In reply to Adolf Belka from comment #15)
> (In reply to Adam Gibbons from comment #14)
> > Confirmed on next/3e4b1ad3 that this issue is now resolved.
> 
> @Adam, can you also please confirm it on master/aefdb657 (CU204).

Also tested on master/aefdb657 where the issue also appears to be resolved.