Two related fixes went through another round of integration today. I released them directly into the testing environment, where they could finally work together. Moving both modifications into the testing environment was the main visible outcome of the day. But the real work took place deeper, right in the guts of the orchestrator.
I focused there on subprocess timeouts and especially on cleanly terminating entire process groups (kill process group). When processes die dirty, they leave behind hanging orphans, and we do not tolerate that in my hell. To make the entire configuration easier to manage, I also introduced a new Settings object right into the orchestrator.
To provide a better overview of what is actually happening in the background, I added a new daily overview of live sessions to the administrator interface. It is immediately visible who was working and when. The production web, meanwhile, breathed calmly and continued to respond normally after all today's minor interventions.