- Bump submodule/base to the acpid AML-handler hardening (bounded stall,
mutex owner-check, static _PSS/_PSD/_CST/_CPC cache, panic-free scheme
path, observability). Proven NOT a regression: a mutex-only baseline
wedges identically under load in a 3/3 framebuffer-ground-truth test,
so the residual under-load boot wedge is head-of-line blocking in
initnsmgr, not acpid.
- Add local/docs/INITNSMGR-CONCURRENCY-DESIGN.md: the concrete
worker-offload design (Design A) that decouples the blocking openat
from the initnsmgr dispatch loop, plus Design B (kernel O_NONBLOCK +
deferred single-thread), a staged plan, and validation rules. Key
finding baked in: redox_rt's Mutex is a spinlock, so the cap_fd must be
resolved under a short lock and the openat run with the lock released.
- Update INIT-NAMESPACE-MANAGER-SCALABILITY-PLAN.md with the acpid
hardening section (done vs the still-deferred #1 transport decoupling).
- Add local/patches/wip-initnsmgr/step1-send-refactor.patch: the
compiled-but-not-yet-boot-validated Step 1 (Rc<RefCell> -> Arc<Mutex>
Send refactor) of initnsmgr, saved durably. It is intentionally NOT in
the base gitlink: bootstrap is the earliest-boot component and this
must be boot-validated on an idle host (the current host is under heavy
external load) before landing. Steps 2-5 (worker bring-up) likewise
need an idle host.
Systematic audit of all 18 scheme-serving daemons for the acpid failure
class (single-threaded daemon that can wait unboundedly in its serving
thread). Findings recorded in INIT-NAMESPACE-MANAGER-SCALABILITY-PLAN.md:
- acpid was the ONLY software unbounded wait in the boot-critical path
(fixed). No other boot-path handler has an unbounded software wait or a
reentrant blocking open of another scheme.
- Hardware busy-waits (rtcd/ps2d/audio/gpu register polls) are
hardware-bounded or in daemons pcid does not spawn without the device.
- ucsid is the one remaining co-victim: it reads /scheme/acpi in
build_state() before publishing its scheme and is the only blocking,
acpi-dependent boot unit in redbear-mini. It already degrades on EAGAIN
and works with acpid fixed. It must NOT be flipped to oneshot_async
(that breaks ucsi scheme registration, which init performs via the
{scheme=…} type); the correct hardening is a source refactor
(publish-then-discover), done with runtime validation, not blind.
redbear-mini now reaches a working brush login and executes commands
(framebuffer ground truth: login -> MOTD -> `user@redbear: $` ->
`echo RB=$((21*2))=OK` -> `RB=42=OK`; login ~12s, brush ~16s in QEMU
q35/KVM). This is the console-login floor of the desktop path.
Root cause was NOT in login/brush/pty/spawn (16+ prior sessions chased
those). cpufreqd reads /scheme/acpi/processor/CPUn/pss; evaluating that
AML ran `_ACQ` on an ACPI mutex whose acquire handler (a) multiplied
the millisecond timeout by 1000 (0xFFFF "wait forever" became ~18h) and
(b) tracked no owner, so a nested acquire by acpid's single AML thread
self-deadlocked. Because acpid is single-threaded and also serves the
`acpi` scheme socket, and because the single-threaded init namespace
manager does a blocking openat in its dispatch loop, a stuck acpid
froze EVERY open in the system. Fixed in submodule/base d78fd44a
(bumped here), verified against local/reference/linux-7.1 ACPICA.
Docs:
- Add local/docs/INIT-NAMESPACE-MANAGER-SCALABILITY-PLAN.md: the
residual architectural root (initnsmgr head-of-line blocking + kernel
ignoring O_NONBLOCK on open) that still lets one slow daemon wedge the
whole open path, with a worker-offload / deferred-open / kernel
O_NONBLOCK execution plan. This is the answer to "use SMP where it's
justified": the parallelism that matters is fault isolation of the
namespace-open path, not throughput.
- CONSOLE-TO-KDE-DESKTOP-PLAN.md v5.9: record the mini-login result and
point at the new plan.
Note: submodule/base worktree also carries in-progress i2c/driver-manager
work (not part of this commit); the acpid fix is isolated.