- Bump submodule/base to the acpid AML-handler hardening (bounded stall, mutex owner-check, static _PSS/_PSD/_CST/_CPC cache, panic-free scheme path, observability). Proven NOT a regression: a mutex-only baseline wedges identically under load in a 3/3 framebuffer-ground-truth test, so the residual under-load boot wedge is head-of-line blocking in initnsmgr, not acpid. - Add local/docs/INITNSMGR-CONCURRENCY-DESIGN.md: the concrete worker-offload design (Design A) that decouples the blocking openat from the initnsmgr dispatch loop, plus Design B (kernel O_NONBLOCK + deferred single-thread), a staged plan, and validation rules. Key finding baked in: redox_rt's Mutex is a spinlock, so the cap_fd must be resolved under a short lock and the openat run with the lock released. - Update INIT-NAMESPACE-MANAGER-SCALABILITY-PLAN.md with the acpid hardening section (done vs the still-deferred #1 transport decoupling). - Add local/patches/wip-initnsmgr/step1-send-refactor.patch: the compiled-but-not-yet-boot-validated Step 1 (Rc<RefCell> -> Arc<Mutex> Send refactor) of initnsmgr, saved durably. It is intentionally NOT in the base gitlink: bootstrap is the earliest-boot component and this must be boot-validated on an idle host (the current host is under heavy external load) before landing. Steps 2-5 (worker bring-up) likewise need an idle host.
14 KiB
Red Bear OS Init Namespace Manager Scalability Plan
Purpose
This document is the canonical analysis and execution plan for a systemic reliability defect
in the boot/console path: the single-threaded, blocking design of the init namespace manager
(initnsmgr) turns any transient slowness in one scheme provider into a whole-system wedge.
It was written after root-causing the long-standing "redbear-mini never reaches a usable brush
login" symptom (16+ prior investigation sessions with shifting, incorrect theories). The immediate
trigger was a concrete acpid bug (fixed — see below), but the reason a single stuck daemon could
freeze the entire system is the architecture described here. That amplifier is still present.
When another document discusses mini/console login reliability, boot slowness, "flaky boot hangs", or the namespace/scheme-open path, prefer this file for the current robustness judgment and the execution order.
Ground state (2026-07-21)
redbear-minilogs intobrushand executes commands. Verified by framebuffer ground truth:Red Bear login: user→ MOTD →user@redbear: $→echo RB=$((21*2))=OK→RB=42=OK. When it wins the race, it is fast (login prompt ~12s, brush prompt ~16s in QEMU q35/KVM).- The dominant blocker — an acpid AML-mutex self-deadlock — is fixed (
local/sources/basecommitd78fd44a). See "Trigger 1" below. - A residual, architectural flakiness remains: under load, boot occasionally wedges at the
fbcond→vesad display handoff (
display.vesa:v2/2). This is not a vesad bug; it is the namespace manager amplifying vesad's transient unreadiness. See "The amplifier" and "Execution plan".
The two layers
Trigger 1 — acpid AML mutex (FIXED, d78fd44a)
cpufreqd reads /scheme/acpi/processor/CPUn/pss (ACPI P-state data). Evaluating that AML runs
_ACQ on an ACPI mutex. The acquire handler in
local/sources/base/drivers/acpid/src/aml_physmem.rs had two bugs:
- Timeout units. ACPI
_ACQtimeouts are milliseconds;0xFFFFis the spec's "wait forever" (ACPICAACPI_WAIT_FOREVER,include/acpi/actypes.h:459). The code multiplied the value by 1000, treating it as seconds —0xFFFFbecame ~18 hours. - No ownership / no recursion. ACPI mutexes are recursive: the owning thread may re-acquire,
each
Acquirepaired with aRelease(ACPICAacpi_ex_acquire_mutex_object,drivers/acpi/acpica/exmutex.c:140,acquisition_depth). The code tracked only a "held" set with no owner, so a nested acquire by the single AML thread waited for a release that only the waiting thread itself could perform — a self-deadlock.
Fixed by tracking (owner ThreadId, recursion depth) per mutex, treating a nested acquire by the
owner as a depth bump, interpreting the timeout as milliseconds, and bounding any wait to 5s so a
misbehaving AML method can never freeze the scheme-serving thread. Verified against
local/reference/linux-7.1 ACPICA.
The amplifier — initnsmgr head-of-line blocking (NOT fixed)
local/sources/base/bootstrap/src/initnsmgr.rs runs the init namespace manager. Every process in a
restricted login namespace (i.e. every login shell, and much of init's own spawn path) resolves
paths through it: the kernel routes an open of scheme:path in that namespace to the manager, which
proxies to the real provider. Two properties make this fragile:
- It is single-threaded: one
run()loop,next_request→handle_sync→write_response. - Its
open_scheme_resourcedoes a blockingsyscall::openaton the provider's capability fd, inside the dispatch loop.
So when any provider is briefly not servicing its socket (descheduled under load, mid-tick, in a
one-shot init window), the manager's openat blocks, the loop stops, and every open in the
system queues behind it forever. This is textbook head-of-line blocking. It also explains mini's
historical slowness: ~2452 namespace opens serialize through that one thread per boot.
acpid was merely the trigger with ~100% hit probability. With acpid fixed, the next-most-likely trigger (vesad's display handoff) surfaces at a lower rate. Fixing individual daemons is a whack-a-mole; the amplifier is the real defect.
Diagnostic proof (reusable): the console/log stack is itself a victim of the wedge, so a diagnostic
that writes /scheme/debug or re-opens a scheme goes silent exactly when the wedge starts. Cut
through it with (1) the QEMU debug console (I/O port 0xE9, qemu -debugcon file:) after acquiring
port-IO rights via the pcid handshake, and (2) pre-opening the kernel scheme fds once and re-reading
via lseek(0)+read. A per-open trace in initnsmgr (to a debug fd from
UPPER_FDTBL_TAG + GlobalSchemes::Debug) pinpointed the wedged open as acpi:processor/CPU3/pss
(then, post-fix, display.vesa:v2/2).
Co-victim audit (2026-07-21): other daemons on the same chain
A systematic audit of all 18 scheme-serving daemons for the acpid failure class (a single-threaded daemon that can wait unboundedly in its serving thread) found:
-
acpid was the only software unbounded wait in the boot-critical path (the AML mutex bug — fixed). No other boot-path scheme handler contains an unbounded software wait or a reentrant blocking open/read/call of another scheme.
-
Hardware busy-waits (
rtcdRTC UIP bit,ps2dPS/2OUTPUT_FULL,sb16d/ihdadDSP/RIRB,ihdgd/ahcid/nvmed/bcm2835register polls) are either hardware-bounded (the bit clears in microseconds) or live in daemons that pcid does not spawn without the matching device (no audio on QEMU q35, etc.). Not the acpid class. (They are still a robustness item under the project WARNING POLICY — unbounded on paper — but adding timeouts to ~15 sites blind, without a reproducing hang, risks regressions and is out of scope here.) -
ucsidis the one remaining co-victim worth noting. It consumes/scheme/acpi/*inbuild_state()before it publishes its scheme, and inconfig/redbear-mini.tomlit is the only boot-critical, acpi-dependent unit with a blocking init type (type = { scheme = "ucsi" }, while every other acpi consumer there isoneshot_async). Before the acpid fix, a stalled acpid would have hung ucsid's init and thus boot. It already degrades gracefully onEAGAIN(discover_ucsi_devicesreturns empty), butEAGAINonly covers "acpid is busy", not "acpid is wedged" (open blocks, per the kernel corollary below). With acpid fixed it works.Do not simply flip ucsid to
oneshot_asyncto satisfy the "optional hardware must not block boot" rule (bare-boot-blocking-init-units):ready_sync_schemehands the scheme cap fd to init via the INIT_NOTIFY pipe (daemon/src/lib.rs:130→ready_with_fd), so init registers theucsischeme on the daemon's behalf — only the{scheme=…}init type does that.oneshot_asyncwould leaveucsiunregistered (consumers getENODEV). The correct hardening is a source refactor: publish the scheme immediately with a lazy/empty state, then run the heavy ACPI/i2c discovery in the background afterready_sync_scheme. It is defense-in-depth against an already-fixed trigger, not an active bug, so it must be done with runtime validation, not blind at session end.
The kernel corollary — open ignores O_NONBLOCK
The manager cannot even defer a slow open today, because in Redox an open() with O_NONBLOCK
still blocks: O_NONBLOCK governs later reads, not the open itself. This is documented in-tree at
local/sources/base/drivers/inputd/src/lib.rs:139 ("this can block indefinitely") and enforced by
local/sources/kernel/src/scheme/user.rs UserInner::call_inner, which unconditionally .block()s
the caller until the provider replies (no O_NONBLOCK check on the request/response path).
Notably, fbcond already has handoff-retry logic that would work if open returned EAGAIN — its
retry never fires today only because the first open blocks instead of returning.
Execution plan (fix EITHER to break the chain; both are worth doing)
Ordered by blast-radius / safety. Neither was attempted at the end of the discovering session:
both live in the most critical components (kernel, no_std bootstrap), where an error means "the
system does not boot at all", and the discovering host was under heavy uncontrolled external load,
making clean validation impossible. They must be done deliberately, with framebuffer-ground-truth
validation on an idle host or real hardware.
Option A — make initnsmgr not block the dispatch loop (preferred)
Two shapes, in increasing order of change:
- Deferred/event-driven open. Attempt the provider
openatnon-blocking; on would-block, park(tag, cap_fd, reference, flags)and return without responding, then complete andwrite_responsewhen the provider becomes ready. Requires the kernel corollary (Option C) so the open can actually report would-block, plus a redox-scheme deferred-response path. - Worker offload. Keep the dispatch loop responsive by handing the blocking
openatto a small worker pool; respond from the worker viasocket.write_response(Response::…, tag)(Socketis an fd wrapper →Send;CallRequest→TagisSend). Cost:bootstrapis#![no_std], so this needs a real thread primitive (redox_rt::thread::rlct_clone_impl— raw stack/TCB/TLS setup), which is the risky part.
This is the direct answer to the operator's SMP question: the parallelism that matters here is fault isolation of the namespace-open path, not throughput. AML execution itself must stay serialized (ACPI requires it; Linux does the same with a global interpreter mutex) — the thing that must not serialize is the transport: serving the scheme socket must not stall behind one slow open.
Option B — bound provider unreadiness at the source
Ensure boot-critical providers (notably vesad) are servicing their scheme socket before anything can open them (e.g. tighten the inputd→driver handoff so the handoff signal follows, not precedes, the provider entering its event loop). Narrower, but only closes known triggers; it does not remove the amplifier for the next slow daemon.
Option C — honor O_NONBLOCK on open in the kernel
Make UserInner::call_inner (open opcode) return EAGAIN instead of blocking when O_NONBLOCK is
set and the provider has not yet taken/answered the request (with the existing cancellation
protocol). This is the cleanest enabler for Option A.1 and immediately activates fbcond's existing
retry, but it changes scheme-open semantics system-wide — the widest blast radius, so it needs the
most care and the strongest validation.
acpid hardening (2026-07-21): the daemon-local half of the same problem
acpid is the same shape as initnsmgr in miniature: single-threaded, serving a scheme (acpi)
on the very thread that evaluates AML. A slow or stuck AML method stops it answering scheme
requests, which is what let the mutex bug wedge the whole system. Six hardening items were
identified; five are done (safe, verifiable, reduce or bound the ways acpid can stall the
transport), one is designed-but-deferred (the same deferred-response rewrite as Option A above,
in a boot-critical daemon).
Done (local/sources/base/drivers/acpid):
- AML mutex recursion + ownership + timeout units (
aml_physmem.rs, the original fix): a nested acquire by the owning thread bumps a depth instead of self-deadlocking; timeout read as milliseconds; any wait bounded to 5s. Mirrors ACPICAexmutex.c:140. stall()bounded per ACPICA (aml_physmem.rs): refuse>255usand warn>100usinstead of an unbounded CPU busy-spin on the serving thread (ACPICAexsystem.c:129-147).releaseowner-check (aml_physmem.rs): a release of a not-currently-held mutex, or by a non-owning thread, is logged (ACPICAAE_NOT_ACQUIRED/ owner-mismatch,exmutex.c:287,376) rather than silently corrupting depth.- Static processor-method cache (
scheme.rs):_PSS/_PSD/_CST/_CPCare fixed after boot but cpufreqd polls them; cache the first evaluation so later reads never re-run the AML interpreter under the global lock. This removes acpid's dominant recurring head-of-line source. - Panic-free scheme path (
acpi.rs): everyrelease_global_lock().expect(...)on the AML-evaluation path became log-and-continue, and aresult.ok()?.unwrap()that panicked onOk(None)(absent method) became?. A scheme-daemon panic kills theacpischeme and wedges every consumer, so the serving path must not panic on odd firmware. - Observability (
aml_physmem.rs,acpi.rs): log AML evaluations ≥50ms and mutex-acquire timeouts — the early-warning signal before a stall becomes a wedge.
Deferred (#1 — decouple AML execution from the transport): make acpid keep answering the acpi
scheme while a slow AML method runs, by rewriting its loop from process_requests_nonblocking onto
the low-level next_request → park the CallRequest Tag → run AML on a single serialized worker
→ write_response(tag) when done. AML must stay serialized (ACPI requires it); only the transport
is parallelized. This is exactly Option A's deferred-response shape, in a boot-critical daemon, and
cannot be cleanly validated under the current external load — so it is designed here, not done
blind. With the cache and bounds above, acpid's remaining exposure is a single slow/first AML
evaluation (bounded to seconds, not infinite), which is the tail this rewrite would close.
Validation expectations
- Framebuffer screendump (QMP) is ground truth; the serial mirror is racy and often silent on a
healthy boot. Harness pattern: QMP
send-keylogin + periodicscreendump. - Because the failure is a race, one boot proves nothing — always compare rates across N≥5 boots
on an idle host (external load, e.g. a background
opencodeat ~290% CPU, contaminates every measurement and inflates the failure rate). - Success criterion for this plan: N≥10 consecutive clean boots to a working brush login on an idle
host, and no
display.vesa:v2/*-class wedge under induced provider delay.
Related
local/docs/CONSOLE-TO-KDE-DESKTOP-PLAN.md— canonical desktop path; console/login is its floor.local/docs/IRQ-AND-LOWLEVEL-CONTROLLERS-ENHANCEMENT-PLAN.md— acpid/ACPI robustness context.- Memory:
mini-login-acpid-mutex-rootcause,fbcond-handoff-console-wedge.