Component ========= knot-resolver 6.4.2 Steps ===== Reboot the machine. Or, on a running system, terminate one worker: shkill -TERM $(supervisorctl -s unix:///run/knot-resolver/supervisord.sock pid kresd:kresd7) Expected ======== The worker exits cleanly. Actual ====== The worker dies with SIGSEGV at offset 0x27603. On reboot every worker does this. A single terminate only crashes that worker, and Supervisor starts a replacement. Sep 29 19:46:32 shutdown: shutting down for system reboot Sep 29 19:46:32 init: Switching to runlevel: 6 Sep 29 19:46:55 supervisord: received SIGTERM indicating exit request Sep 29 19:46:56 kernel: kresd[2481]: segfault ... in kresd[27603,...] Sep 29 19:46:56 supervisord: stopped: kresd0 through kresd7 (terminated by SIGSEGV) The same line is in /var/log/messages after a manual terminate.
Edit... Small spelling mistake on the last post the command should have been: kill -TERM $(supervisorctl -s unix:///run/knot-resolver/supervisord.sock pid kresd:kresd7) --- The following cause and fix is AI generated so perhaps take with a pinch of salt: Cause ===== src/patches/knot-resolver-6.5-keep-dst-ip.patch stores an io_handle_data_t in the UDP listener's handle->data. endpoint_close() still treats that pointer as a session and calls session2_close() on it. Shutdown then faults in session2_event_unwrap(). Fix ==== endpoint_close() needs the session from the wrapper. on_session2_handle_close() has the same assumption.
Here is a fix for this problem: > https://git.ipfire.org/?p=people/ms/knot-resolver.git;a=commitdiff;h=ed453515e8c880baa96a13e8ddef485f200fe6e2 It has also been updated in master/next: > https://git.ipfire.org/?p=ipfire-2.x.git;a=commitdiff;h=284715c5b03a131feb2f2bef39c0e5fef264add2 Please let me know if this fixes the problem as this is currently the last remaining blocker for the release.
Created attachment 1760 [details] /var/log/messages Test CU204 Development Build: master/88a55eb0 DNS Ko, see log messages errors Oct 6 16:06:11 ipfireTest kresd[9994]: [system] assertion "(data->session->transport.type == SESSION2_TRANSPORT_IO && data->session->transport.io.handle == handle)" failed in on_session2_handle_close@../daemon/session2.c:1780 Oct 6 16:06:11 ipfireTest supervisord: exited: kresd0 (terminated by SIGABRT; not expected)
Thanks for the feedback. On the on_session2_handle_close part, how are you reproducing this? I have just following the advice from Adam's AI and have not been able to test the change.
The problem occurs immediately upon startup following installation or upgrade.
Interesting. It didn't for me. Do you have a chance to revert the last commit, rebuild and check again?
(In reply to Michael Tremer from comment #6) > Interesting. It didn't for me. Do you have a chance to revert the last > commit, rebuild and check again? Sorry, I don't know if you're addressing me, but I don't know how to do that on the IPFire repository. My build platform is too slow to quickly perform a local rebuild.
Yes, this was for you. I have reverted the commit (as it has not been a problem before), but the build will take a couple of hours to come through: > https://git.ipfire.org/?p=ipfire-2.x.git;a=commitdiff;h=61c40dd31ccd7b3a8a7e0183ce3995a05a6fd3e6 This is only in next.
OK I'll test your CU205 when the ISO is ready. I'm also reverting on the master locally on my build system, but my build will take longer (> 3 hours).
Created attachment 1761 [details] /var/log/messages I've just finished local compil CU204 with a revert on master. phili@PCBUREAU:~/projects/ipfire-2.x$ git status On branch master Your branch is behind 'origin/master' by 1 commit, and can be fast-forwarded. (use "git pull" to update your local branch) nothing to commit, working tree clean phili@PCBUREAU:~/projects/ipfire-2.x$ sudo ./make.sh build I installed the ISO on my test VM There are no more kresd error messages; here is the reboot log.
Thank you very much, so I will revert the commit in the other places, too.
I just tested the installation of the x86_64 ISO for the CU205 Development Build (next/61c40dd3) It works fine, and there are no more kresd error messages either. Note: In VMware, I had to use the SATA HDD because installer couldn't find the SCSI drive.
(In reply to pscar13 from comment #12) > Note: In VMware, I had to use the SATA HDD because installer couldn't find > the SCSI drive. Could you file a separate ticket for this, please?
Confirmed on next/3e4b1ad3 that this issue is now resolved.
(In reply to Adam Gibbons from comment #14) > Confirmed on next/3e4b1ad3 that this issue is now resolved. @Adam, can you also please confirm it on master/aefdb657 (CU204).
(In reply to Adolf Belka from comment #15) > (In reply to Adam Gibbons from comment #14) > > Confirmed on next/3e4b1ad3 that this issue is now resolved. > > @Adam, can you also please confirm it on master/aefdb657 (CU204). Also tested on master/aefdb657 where the issue also appears to be resolved.