diff --git a/local/docs/DRIVER-MANAGER-MIGRATION-PLAN.md b/local/docs/DRIVER-MANAGER-MIGRATION-PLAN.md index 8b8d8690f6..b3e22b3503 100644 --- a/local/docs/DRIVER-MANAGER-MIGRATION-PLAN.md +++ b/local/docs/DRIVER-MANAGER-MIGRATION-PLAN.md @@ -1,18 +1,21 @@ # Red Bear OS — `pci-spawner` → `driver-manager` Migration Plan -**Document status:** v5.3 canonical planning authority (supersedes v5.2). -v5.3 records W2 closure (broken KDE daemon activation .service files -removed). v5.2 closed C1 (OHCI bulk + interrupt transfers). v5.1 -closed G-A4 (iwlwifi spawned-mode channel contract). v5.0 closed v5.3 -(initnsmgr O_NONBLOCK) and the W1-W8 stub-fix pass. v4.9 recorded -v5.0, v5.1, v5.5 implementation. v4.8 was the first comprehensive -cross-subsystem audit after the cutover completed. +**Document status:** v5.4 canonical planning authority (supersedes v5.3). +v5.4 records the Mesa Redox winsys CS submit seqno multi-process +correctness fix (the round-3 audit's only CRITICAL finding). v5.3 +records W2 closure (broken KDE daemon activation .service files removed). +v5.2 closed C1 (OHCI bulk + interrupt transfers). v5.1 closed G-A4 +(iwlwifi spawned-mode channel contract). v5.0 closed v5.3 (initnsmgr +O_NONBLOCK) and the W1-W8 stub-fix pass. v4.9 recorded v5.0, v5.1, +v5.5 implementation. v4.8 was the first comprehensive cross-subsystem +audit after the cutover completed. -The round-2 audit (2026-07-26) found C1 was missed by W1-W8; the -remaining items (W1 btctl stub backend, W3 seatd incomplete, W4 -notifications stderr-only, W5 redox-drm relocations) are documented -limitations that require real kernel-stack work (BlueZ, seatd Redox -port, D-Bus display integration, i915 GEM relocations) beyond the +The round-3 audit (2026-07-26) found the Mesa CS seqno fake was a +real multi-process correctness bug and fixed it; the remaining items +(W1 btctl stub backend, W3 seatd incomplete, W4 notifications +stderr-only, W5 redox-drm relocations) are documented limitations +that require real kernel-stack work (BlueZ, seatd Redox port, D-Bus +display integration, i915 GEM relocations) beyond the scope of this plan. v4.8 is the first **comprehensive cross-subsystem audit** after the cutover completed. It documents five newly-discovered runtime-grade gaps that @@ -125,6 +128,68 @@ end-state once a freestanding thread-spawn helper is added to `redox_rt`. | **v5.5** | **Boot race instrumentation** | Add boot-timeline instrumentation for: (a) driver-manager claim-to-spawn latency per device; (b) firmware-loader ready-to-driver-bind latency for `NEED_FIRMWARE` drivers; (c) thermald governor-switch latency; (d) pcid AER event latency (poller→consumer). Surface as `/scheme/driver-manager/timing` and `/tmp/redbear-boot-timeline.json` extensions. | ✅ **DONE 2026-07-25** — commit `045aaa4579`. New `timing.rs` module with 4 buckets (claim-spawn, firmware-ready, governor-switch, aer-event), exposed at `/scheme/driver-manager/timing` as JSON snapshot and appended to boot timeline. 24 new tests, 112 total passing. | | **v5.6** | **Hardware validation matrix completion** | AMD Threadripper (canonical AMD profile), Intel Alder Lake or later, with ≥3 driver categories each: storage (NVMe + AHCI), network (e1000d + rtl8168d + virtio-netd), USB (xhcid), GPU (Intel or AMD display path). Per `local/docs/HARDWARE-VALIDATION-MATRIX.md`. | 🔴 **OPERATOR-ONLY** — D5 ratification gate. No agent can perform real-hardware validation. Status unchanged from v4.7. | +### v5.4 implementation summary (2026-07-26) + +**Mesa CS submit seqno multi-process correctness** (commit `9846f288b4`): + +The round-3 audit found the Mesa Redox winsys CS submit path +(`src/gallium/winsys/redox/drm/redox_drm_cs.c:156`) was faking the +seqno with `result.seqno = rws->cs->last_seqno + 1 : 1;` instead of +reading the kernel-assigned seqno. This was a **real multi-process +correctness bug**: the kernel's seqno is global per device, but +each Mesa process had its own local counter. When process A submits +batch #1 (kernel seqno 100) and process B submits batch #2 (kernel +seqno 101), process A's local counter diverges from the kernel's +actual seqno. Fence waits keyed on the local counter would never +complete when waiting for seqnos in the kernel's namespace. Common +case: any compositor + GPU client setup (KDE compositor plus browser, +video player, etc.). + +**Fix** (three coordinated changes following the standard DRM +bidirectional-ioctl pattern that `DrmAmdgpuCsWire` already uses): + +**redox-drm kernel side:** +- `driver.rs`: `RedoxPrivateCsSubmit` gains `seqno: u64` output + field (40 bytes total). `RedoxPrivateCsWait` gains `completed`, + `_pad`, `completed_seqno` response fields (32 bytes total). Both + structs gain `Default` derive for `..Default::default()` use. +- `scheme.rs`: CS_SUBMIT handler writes `resp.seqno` back into + `req.seqno` and serializes `req` (`bytes_of(&resp)` → + `bytes_of(&req)`). CS_WAIT handler similarly copies the + result fields into req. The trait's `RedoxPrivateCsSubmitResult` + and `RedoxPrivateCsWaitResult` types are retained as return + types; the wire response is carried in the bidirectional struct. +- `drivers/amd/mod.rs`, `intel/mod.rs`, `virtio/mod.rs`: each + cs_submit / cs_wait construction site uses + `..Default::default()` for the new response fields. No logic + changes (the driver impls still return the response via the + trait method, scheme.rs copies it into the bidirectional struct). + +**Mesa winsys side:** +- `source/src/gallium/winsys/redox/drm/redox_drm_cs.c`: merge + `RedoxCsSubmitWire` + `RedoxCsSubmitResultWire` into ONE + bidirectional `RedoxCsSubmitWire` (with seqno output field). + Same for wait. Remove the fake `result.seqno = rws->cs->...` + line. Instead read `submit.seqno` and `wait.completed_seqno` + from the same struct after `drmIoctl` returns. The file's + header comment now explicitly documents the bidirectional + pattern, the kernel ABI, and the multi-process correctness + consequence. + +**Durability:** +- `local/patches/mesa/26-cs-submit-bidirectional-seqno.patch` + (135 lines, new): persists the C-side merge across clean + re-extracts of the upstream Mesa 26.1.4 tarball. +- `local/recipes/libs/mesa/recipe.toml`: patch added to the + `patches = [...]` list with a comment noting it's required + for multi-process GPU fence correctness. + +**Verification gate (operator-side):** the runtime fix requires +real-hardware multi-process GPU testing (compositor + clients +all using the same `/scheme/drm/card0`). Fence waits between +processes must complete correctly. This is the operator-side +runtime gate, not compilable. + ### v5.0 implementation summary (2026-07-26) **v5.3 initnsmgr Design B (kernel + base paired change)**: