# Red Bear OS Init Namespace Manager Scalability Plan ## Purpose This document is the **canonical analysis and execution plan** for a systemic reliability defect in the boot/console path: the single-threaded, blocking design of the init namespace manager (`initnsmgr`) turns any transient slowness in *one* scheme provider into a *whole-system* wedge. It was written after root-causing the long-standing "`redbear-mini` never reaches a usable brush login" symptom (16+ prior investigation sessions with shifting, incorrect theories). The immediate trigger was a concrete acpid bug (fixed — see below), but the reason a single stuck daemon could freeze the entire system is the architecture described here. That amplifier is still present. When another document discusses mini/console login reliability, boot slowness, "flaky boot hangs", or the namespace/scheme-open path, prefer this file for the current robustness judgment and the execution order. ## Ground state (2026-07-21) - **`redbear-mini` logs into `brush` and executes commands.** Verified by framebuffer ground truth: `Red Bear login: user` → MOTD → `user@redbear: $` → `echo RB=$((21*2))=OK` → `RB=42=OK`. When it wins the race, it is fast (login prompt ~12s, brush prompt ~16s in QEMU q35/KVM). - The dominant blocker — an acpid AML-mutex self-deadlock — is **fixed** (`local/sources/base` commit `d78fd44a`). See "Trigger 1" below. - A **residual, architectural** flakiness remains: under load, boot occasionally wedges at the fbcond→vesad display handoff (`display.vesa:v2/2`). This is not a vesad bug; it is the namespace manager amplifying vesad's transient unreadiness. See "The amplifier" and "Execution plan". ## The two layers ### Trigger 1 — acpid AML mutex (FIXED, `d78fd44a`) `cpufreqd` reads `/scheme/acpi/processor/CPUn/pss` (ACPI P-state data). Evaluating that AML runs `_ACQ` on an ACPI mutex. The `acquire` handler in `local/sources/base/drivers/acpid/src/aml_physmem.rs` had two bugs: 1. **Timeout units.** ACPI `_ACQ` timeouts are milliseconds; `0xFFFF` is the spec's "wait forever" (ACPICA `ACPI_WAIT_FOREVER`, `include/acpi/actypes.h:459`). The code multiplied the value by 1000, treating it as seconds — `0xFFFF` became ~18 hours. 2. **No ownership / no recursion.** ACPI mutexes are recursive: the owning thread may re-acquire, each `Acquire` paired with a `Release` (ACPICA `acpi_ex_acquire_mutex_object`, `drivers/acpi/acpica/exmutex.c:140`, `acquisition_depth`). The code tracked only a "held" set with no owner, so a nested acquire by the single AML thread waited for a release that only the waiting thread itself could perform — a self-deadlock. Fixed by tracking `(owner ThreadId, recursion depth)` per mutex, treating a nested acquire by the owner as a depth bump, interpreting the timeout as milliseconds, and bounding any wait to 5s so a misbehaving AML method can never freeze the scheme-serving thread. Verified against `local/reference/linux-7.1` ACPICA. ### The amplifier — `initnsmgr` head-of-line blocking (NOT fixed) `local/sources/base/bootstrap/src/initnsmgr.rs` runs the init namespace manager. Every process in a restricted login namespace (i.e. every login shell, and much of init's own spawn path) resolves paths through it: the kernel routes an open of `scheme:path` in that namespace to the manager, which proxies to the real provider. Two properties make this fragile: - It is **single-threaded**: one `run()` loop, `next_request` → `handle_sync` → `write_response`. - Its `open_scheme_resource` does a **blocking `syscall::openat`** on the provider's capability fd, *inside the dispatch loop*. So when any provider is briefly not servicing its socket (descheduled under load, mid-`tick`, in a one-shot init window), the manager's `openat` blocks, the loop stops, and **every open in the system queues behind it forever**. This is textbook head-of-line blocking. It also explains mini's historical slowness: ~2452 namespace opens serialize through that one thread per boot. acpid was merely the trigger with ~100% hit probability. With acpid fixed, the next-most-likely trigger (vesad's display handoff) surfaces at a lower rate. **Fixing individual daemons is a whack-a-mole; the amplifier is the real defect.** Diagnostic proof (reusable): the console/log stack is itself a victim of the wedge, so a diagnostic that writes `/scheme/debug` or re-opens a scheme goes silent exactly when the wedge starts. Cut through it with (1) the QEMU debug console (I/O port `0xE9`, `qemu -debugcon file:`) after acquiring port-IO rights via the pcid handshake, and (2) pre-opening the kernel scheme fds once and re-reading via `lseek(0)`+`read`. A per-open trace in `initnsmgr` (to a debug fd from `UPPER_FDTBL_TAG + GlobalSchemes::Debug`) pinpointed the wedged open as `acpi:processor/CPU3/pss` (then, post-fix, `display.vesa:v2/2`). ## Co-victim audit (2026-07-21): other daemons on the same chain A systematic audit of all 18 scheme-serving daemons for the acpid failure class (a single-threaded daemon that can wait unboundedly in its serving thread) found: - **acpid was the only *software* unbounded wait** in the boot-critical path (the AML mutex bug — fixed). No other boot-path scheme handler contains an unbounded software wait or a reentrant blocking open/read/call of another scheme. - **Hardware busy-waits** (`rtcd` RTC UIP bit, `ps2d` PS/2 `OUTPUT_FULL`, `sb16d`/`ihdad` DSP/RIRB, `ihdgd`/`ahcid`/`nvmed`/`bcm2835` register polls) are either hardware-bounded (the bit clears in microseconds) or live in daemons that pcid does not spawn without the matching device (no audio on QEMU q35, etc.). Not the acpid class. (They are still a robustness item under the project WARNING POLICY — unbounded on paper — but adding timeouts to ~15 sites blind, without a reproducing hang, risks regressions and is out of scope here.) - **`ucsid` is the one remaining co-victim worth noting.** It consumes `/scheme/acpi/*` in `build_state()` *before* it publishes its scheme, and in `config/redbear-mini.toml` it is the only boot-critical, acpi-dependent unit with a **blocking** init type (`type = { scheme = "ucsi" }`, while every other acpi consumer there is `oneshot_async`). Before the acpid fix, a stalled acpid would have hung ucsid's init and thus boot. It already degrades gracefully on `EAGAIN` (`discover_ucsi_devices` returns empty), but `EAGAIN` only covers "acpid is busy", not "acpid is wedged" (open blocks, per the kernel corollary below). With acpid fixed it works. **Do not simply flip ucsid to `oneshot_async`** to satisfy the "optional hardware must not block boot" rule (`bare-boot-blocking-init-units`): `ready_sync_scheme` hands the scheme cap fd to init via the INIT_NOTIFY pipe (`daemon/src/lib.rs:130` → `ready_with_fd`), so init registers the `ucsi` scheme *on the daemon's behalf* — only the `{scheme=…}` init type does that. `oneshot_async` would leave `ucsi` unregistered (consumers get `ENODEV`). The correct hardening is a source refactor: publish the scheme immediately with a lazy/empty state, then run the heavy ACPI/i2c discovery in the background after `ready_sync_scheme`. It is defense-in-depth against an already-fixed trigger, not an active bug, so it must be done with runtime validation, not blind at session end. ## The kernel corollary — `open` ignores `O_NONBLOCK` The manager cannot even *defer* a slow open today, because in Redox an `open()` with `O_NONBLOCK` still blocks: `O_NONBLOCK` governs later reads, not the open itself. This is documented in-tree at `local/sources/base/drivers/inputd/src/lib.rs:139` ("this can block indefinitely") and enforced by `local/sources/kernel/src/scheme/user.rs` `UserInner::call_inner`, which unconditionally `.block()`s the caller until the provider replies (no `O_NONBLOCK` check on the request/response path). Notably, **fbcond already has handoff-retry logic that would work if open returned `EAGAIN`** — its retry never fires today only because the first open blocks instead of returning. ## Execution plan (fix EITHER to break the chain; both are worth doing) Ordered by blast-radius / safety. Neither was attempted at the end of the discovering session: both live in the most critical components (kernel, `no_std` bootstrap), where an error means "the system does not boot at all", and the discovering host was under heavy uncontrolled external load, making clean validation impossible. They must be done deliberately, with framebuffer-ground-truth validation on an idle host or real hardware. ### Option A — make `initnsmgr` not block the dispatch loop (preferred) Two shapes, in increasing order of change: 1. **Deferred/event-driven open.** Attempt the provider `openat` non-blocking; on would-block, park `(tag, cap_fd, reference, flags)` and return without responding, then complete and `write_response` when the provider becomes ready. Requires the kernel corollary (Option C) so the open can actually report would-block, plus a redox-scheme deferred-response path. 2. **Worker offload.** Keep the dispatch loop responsive by handing the blocking `openat` to a small worker pool; respond from the worker via `socket.write_response(Response::…, tag)` (`Socket` is an fd wrapper → `Send`; `CallRequest`→`Tag` is `Send`). Cost: `bootstrap` is `#![no_std]`, so this needs a real thread primitive (`redox_rt::thread::rlct_clone_impl` — raw stack/TCB/TLS setup), which is the risky part. This is the direct answer to the operator's SMP question: the parallelism that matters here is **fault isolation of the namespace-open path**, not throughput. AML execution itself must stay serialized (ACPI requires it; Linux does the same with a global interpreter mutex) — the thing that must not serialize is the *transport*: serving the scheme socket must not stall behind one slow open. ### Option B — bound provider unreadiness at the source Ensure boot-critical providers (notably vesad) are servicing their scheme socket before anything can open them (e.g. tighten the inputd→driver handoff so the handoff signal follows, not precedes, the provider entering its event loop). Narrower, but only closes known triggers; it does not remove the amplifier for the *next* slow daemon. ### Option C — honor `O_NONBLOCK` on `open` in the kernel Make `UserInner::call_inner` (open opcode) return `EAGAIN` instead of blocking when `O_NONBLOCK` is set and the provider has not yet taken/answered the request (with the existing cancellation protocol). This is the cleanest enabler for Option A.1 and immediately activates fbcond's existing retry, but it changes scheme-open semantics system-wide — the widest blast radius, so it needs the most care and the strongest validation. ## acpid hardening (2026-07-21): the daemon-local half of the same problem `acpid` is the same shape as `initnsmgr` in miniature: single-threaded, serving a scheme (`acpi`) on the very thread that evaluates AML. A slow or stuck AML method stops it answering scheme requests, which is what let the mutex bug wedge the whole system. Six hardening items were identified; five are **done** (safe, verifiable, reduce or bound the ways acpid can stall the transport), one is **designed-but-deferred** (the same deferred-response rewrite as Option A above, in a boot-critical daemon). **Done (`local/sources/base/drivers/acpid`):** - **AML mutex recursion + ownership + timeout units** (`aml_physmem.rs`, the original fix): a nested acquire by the owning thread bumps a depth instead of self-deadlocking; timeout read as milliseconds; any wait bounded to 5s. Mirrors ACPICA `exmutex.c:140`. - **`stall()` bounded per ACPICA** (`aml_physmem.rs`): refuse `>255us` and warn `>100us` instead of an unbounded CPU busy-spin on the serving thread (ACPICA `exsystem.c:129-147`). - **`release` owner-check** (`aml_physmem.rs`): a release of a not-currently-held mutex, or by a non-owning thread, is logged (ACPICA `AE_NOT_ACQUIRED` / owner-mismatch, `exmutex.c:287,376`) rather than silently corrupting depth. - **Static processor-method cache** (`scheme.rs`): `_PSS`/`_PSD`/`_CST`/`_CPC` are fixed after boot but cpufreqd polls them; cache the first evaluation so later reads never re-run the AML interpreter under the global lock. This removes acpid's dominant *recurring* head-of-line source. - **Panic-free scheme path** (`acpi.rs`): every `release_global_lock().expect(...)` on the AML-evaluation path became log-and-continue, and a `result.ok()?.unwrap()` that panicked on `Ok(None)` (absent method) became `?`. A scheme-daemon panic kills the `acpi` scheme and wedges every consumer, so the serving path must not panic on odd firmware. - **Observability** (`aml_physmem.rs`, `acpi.rs`): log AML evaluations ≥50ms and mutex-acquire timeouts — the early-warning signal before a stall becomes a wedge. **Deferred (#1 — decouple AML execution from the transport):** make acpid keep answering the `acpi` scheme while a slow AML method runs, by rewriting its loop from `process_requests_nonblocking` onto the low-level `next_request` → park the `CallRequest` `Tag` → run AML on a single serialized worker → `write_response(tag)` when done. AML must stay serialized (ACPI requires it); only the transport is parallelized. This is exactly Option A's deferred-response shape, in a boot-critical daemon, and cannot be cleanly validated under the current external load — so it is designed here, not done blind. With the cache and bounds above, acpid's remaining exposure is a *single* slow/first AML evaluation (bounded to seconds, not infinite), which is the tail this rewrite would close. ## Validation expectations - Framebuffer screendump (QMP) is ground truth; the serial mirror is racy and often silent on a healthy boot. Harness pattern: QMP `send-key` login + periodic `screendump`. - Because the failure is a race, **one boot proves nothing** — always compare rates across N≥5 boots on an idle host (external load, e.g. a background `opencode` at ~290% CPU, contaminates every measurement and inflates the failure rate). - Success criterion for this plan: N≥10 consecutive clean boots to a working brush login on an idle host, and no `display.vesa:v2/*`-class wedge under induced provider delay. ## Related - `local/docs/CONSOLE-TO-KDE-DESKTOP-PLAN.md` — canonical desktop path; console/login is its floor. - `local/docs/IRQ-AND-LOWLEVEL-CONTROLLERS-ENHANCEMENT-PLAN.md` — acpid/ACPI robustness context. - Memory: `mini-login-acpid-mutex-rootcause`, `fbcond-handoff-console-wedge`.