|
1 | 1 | --- |
2 | 2 | type: mistakes |
3 | 3 | project: Vol NixOS |
4 | | -last_updated: 2026-09-28 |
| 4 | +last_updated: 2026-10-01 |
5 | 5 | status: append-only |
6 | 6 | --- |
7 | 7 |
|
@@ -114,28 +114,32 @@ This file catalogs past bugs, configuration issues, and operational pitfalls enc |
114 | 114 | * **Prevention Rule:** If GTK/Electron file pickers or portal Settings fail with `AccessDenied` / `Unable to open /proc/<pid>/root`, do NOT chase portal backends, icons, or `GTK_USE_PORTAL`. Reproduce with `gdbus call --session --dest org.freedesktop.portal.Desktop --object-path /org/freedesktop/portal/desktop --method org.freedesktop.portal.Settings.ReadAll '[]'`; if it errors, the app-id step is broken. Compare against `dbus-run-session -- <same call>`. If the daemon works and the live broker bus does not, set `services.dbus.implementation = "dbus"`. |
115 | 115 | * **Rebuild caution:** Switching the dbus implementation restarts the message bus on `switch` and will tear down the running Wayland session (see Mistake #1). Apply via reboot, or run the rebuild detached (tmux / `systemd-run`). |
116 | 116 |
|
117 | | -### 2026-09-24 — Audio Module Activation: Codec Hardware Mute Bits Set, Headphone Jack-Sense Failing |
| 117 | +### 2026-09-30 — Full-Repo Audit Surfaces Multiple Unit, Script, and Config Defects |
118 | 118 |
|
119 | | -* **Symptom:** After audio module activation (make switch), onboard Realtek ALC256 speaker output completely silent despite software volume at 126% and no software mute. When investigating, found speaker and headphone pins had codec hardware mute bits set (Amp-Out vals: [0x80 0x80]). Headphones also failed to appear in output port list; jack-sense reports "not available" even with physical headphone insertion. |
| 119 | +* **Symptom/cause/fix, by component:** |
| 120 | + - `decapitate-fuse-mounts` systemd unit: `DefaultDependencies=` sat under `[Service]` instead of `[Unit]` (systemd only honors it there, so it was silently ignored); the unit called bare `umount` instead of the coreutils path, had no `|| true` fallback for already-unmounted targets, and hardcoded uid 1000 against a real uid of 1001. Net effect: shutdown-time unmounts were unreliable. Rewritten as a script unit wanted by `shutdown.target` with correct directive placement, full path, fallback, and matching uid. |
| 121 | + - `phone-proximity-daemon` was ordered under `default.target` (pre-graphical boot) but depends on `NIRI_SOCKET`, which only exists once niri is running. Moved to `graphical-session.target`. |
| 122 | + - fish `rmspcs` was missing its closing `end`; `gpgkey` used the nonexistent `read -S` instead of `read -s` for silent input. Both are fish syntax/runtime errors invisible to `nix build` — only an interactive fish smoke-test catches them. |
| 123 | + - `ingest-sync` and `phone-mcp-call.sh` hand-built JSON via shell string interpolation; `ingest-sync` also didn't reject filenames containing `/`. Both now build payloads with `jq` (added to the client's PATH); ingest-sync rejects `/` in names. |
| 124 | + - `anon-selftest` had a mangled multi-substitution `printf` producing malformed diagnostic output, and the jail seal check only asserted the IPv4 blackhole, leaving the IPv6 leg of decision #42's L4 ladder unverified. Both fixed. |
| 125 | + - `nix-ld` referenced a hardcoded `linuxPackages.nvidia_x11` instead of `config.hardware.nvidia.package`, risking a version mismatch against whichever NVIDIA driver package the host actually configures. Fixed to reference the config option. |
120 | 126 |
|
121 | | -* **Root cause (dual failure):** (1) **Speaker:** Codec hardware mute bits set at bootstrap, downstream of software volume controls. This prevented any audio from flowing regardless of software settings. Cause unknown — possibly wireplumber initialization ordering or audio module configuration not clearing mute on startup. (2) **Headphones:** Jack-sense detection failed to trigger on physical insertion. Codec correctly identified headphones at boot (hp_outs=1, node 0x21), but jack-sense kcontrol never updates state from "not available" to "available". Possible causes: missing jack-detect kcontrol setup in wireplumber config, BIOS/ACPI DSDT issue, or kernel driver configuration missing. |
| 127 | +* **Prevention rule:** Config-as-code checks (`nix build`, `nix flake check`) do not parse the *bodies* of shell/fish scripts or validate that a systemd directive sits in the section that actually reads it — both classes of bug here passed every automated gate while broken. Any script or unit touched during a refactor needs an actual runtime smoke-test (fish call, `systemd-analyze verify` at minimum, ideally a live exercise), not just a green build. Per decision #43 (silence is not evidence), the IPv6 seal gap is the same pattern again: the negative-path check simply never asked the question, and its absence looked identical to a passing check. |
122 | 128 |
|
123 | | -* **Resolution (partial):** Speaker mute bit cleared manually via `pactl set-sink-mute alsa_output.pci-0000_66_00.6.analog-stereo false`; audio restored immediately (Amp-Out flipped to [0x00 0x00]). Headphone issue unresolved — jack-sense still non-functional. Requires deeper investigation. |
| 129 | +* **Status:** All fixes applied and built locally (`nix build --no-link`, exit 0) as of 2026-09-30; uncommitted, unswitched. |
124 | 130 |
|
125 | | -* **Prevention rule:** (1) Audio module or wireplumber initialization should explicitly clear codec mute bits on startup (not rely on defaults). Add `amixer -c <N> sset Master unmute` or wireplumber post-startup hook if not present. (2) Jack-detect kcontrol must be explicitly enabled in wireplumber config if kernel driver does not auto-enable; verify `amixer -c <N> scontents | grep -i jack-detect` post-activation and ensure state is ON. (3) Test both speaker and headphones with actual sound output and physical headphone insertion immediately after audio module changes — silence does not mean success; measure actual audio flow. |
| 131 | +### 2026-09-30 — nvidia-drm.modeset=1 Incorrectly Flagged as Redundant, Restored After Verification |
126 | 132 |
|
127 | | -### 2026-09-24 — smartmontools Missing, Backup Drive Health Unmonitored Until Error Surfaces |
| 133 | +* **Symptom:** A delegated sub-task during the audit proposed removing `nvidia-drm.modeset=1` from `kernelParams`, on the claim that `hardware.nvidia` adds it automatically. |
128 | 134 |
|
129 | | -* **Symptom:** External Seagate 2TB USB backup drive logged unrecovered read error (sector 3574956888) during routine restic backup run on 2026-09-24. No prior SMART monitoring configured; `smartctl` not installed on host. Error surfaced at block layer, not file-level; `restic check` completed with "no errors" despite medium failure underneath. |
| 135 | +* **Root cause:** False on this config — evaluation showed the nvidia module does not add this parameter here. Removing it would have disabled DRM kernel modesetting for the NVIDIA driver. |
130 | 136 |
|
131 | | -* **Root cause:** No SMART health monitoring infrastructure. smartmontools not installed; no long-running SMART daemon (`services.smartd`). Reliance on `restic check` as sole health signal is insufficient — logical verification passes while physical medium degrades silently. |
| 137 | +* **Prevention rule:** Verify claims about what a NixOS module "already sets" by tracing the actual module source or `nix eval` output, not by assumption — especially when a delegated/sub-agent worker proposes removing a line as redundant cleanup. The removal was caught before being applied (system was rebuilt with the param restored). |
132 | 138 |
|
133 | | -* **Prevention rule:** (1) Install `pkgs.smartmontools` on any host using external USB backup drives. (2) Configure `services.smartd` with external drive explicitly included (or excluded if external-only monitoring deferred). (3) On backup-drive attach, run `smartctl -a -d sat /dev/sdX` to read SMART attributes (Reallocated_Sector_Ct, Current_Pending_Sector, Offline_Uncorrectable). (4) Trend SMART history across multiple backup runs; do not rely on single-run `restic check` to validate medium health. (5) Replace drive if reallocation or pending-sector counts are non-zero and climbing — restic repo is only as durable as the medium it lives on. |
| 139 | +### 2026-10-01 — ~/.omo Wiped by Impermanence Reboot, Pre-Seeding Step Not Executed |
134 | 140 |
|
135 | | -### 2026-09-25 — Cachix CI Token Expired, Silent CI Outage for 36h |
| 141 | +* **Symptom:** ~/.omo directory completely wiped at 2026-10-01 boot (tmpfs-root wipe). Omo configuration, community plugins, wallpaper state, and any persisted data lost. Impermanence activation was scheduled but pre-seeding backup never created. |
136 | 142 |
|
137 | | -* **Symptom:** GitHub Actions CI failed 2026-09-25/26 (runs 36094076894, 36231008195) at `cachix/cachix-action@v16` step (~1m45s into build) with error: `Binary cache volnixos doesn't exist or you don't have access. Error: Cachix Auth token "CI deployment" has expired.` No prior warning; token expiry is silent (Cachix does not send revocation alerts or pre-expiry warnings). Builds were red for 36 hours before investigation. |
| 143 | +* **Root cause:** Audit fix pass declared ~/.omo persistence binding in `home/persist.nix` with a todo item to pre-seed via `cp -r ~/.omo /persist/home/lowcache/` before `make switch`. This pre-seeding was deferred (in the audit fix pass todo checklist). A system reboot occurred before the pre-seeding step was executed. At boot, tmpfs root was wiped as designed, but ~/.omo had never been copied to /persist, so the live configuration and extensions were lost completely. The persistence binding would have protected the data if it had been seeded first. |
138 | 144 |
|
139 | | -* **Root cause:** Cache-scoped Cachix token created 2026-08-26T10:11:34Z with **30-day default expiry**. Token reached expiry at 2026-09-25T10:11Z (exactly 30 days later). Token was not revoked; expiry time was simply reached. Prior documentation incorrectly labeled token as "account-scoped" (commit 3f2e952 corrected this — token was always cache-scoped, which is correct). The 30-day default expiry was unnecessarily aggressive for a production CI credential that is infrequently rotated. |
140 | | - |
141 | | -* **Prevention rule:** (1) Cachix tokens for CI should be created with **1-year expiry, not 30-day default**. (2) Set calendar reminder or automated sweep task to **proactively rotate tokens ~60 days before expiry** — do not wait for expiry to occur and cause outage. (3) Do NOT rely on Cachix notifications (silent expiry; no pre-expiry alerts exist). (4) When rotating tokens via `gh secret set CACHIX_AUTH_TOKEN`, use **interactive masked prompt** (`gh secret set CACHIX_AUTH_TOKEN`, then paste in prompt) rather than piping via echo — `echo "$TOKEN" | gh secret set` introduces a trailing newline (cli/cli#5031) that breaks the token. (5) Verify token rotation by confirming next CI build completes successfully with new token (check cachix-action push log in workflow run). (6) Secondary: formatting errors found during CI (e.g., flake.nix blank lines) should be caught locally (`make check`) to avoid burning 1h39m CI cycles on trivial fixes. |
| 145 | +* **Prevention rule:** Impermanence-bound directories must be pre-seeded (live state copied to /persist) and that backup must be verified BEFORE any system reboot. If a reboot must happen before seeding, document and accept the data loss explicitly. When scheduling pre-seeding steps, enforce ordering: (1) declare persistence, (2) pre-seed from live state, (3) verify backup exists, (4) reboot. Use blocking task dependencies or clear reminders to prevent skipping steps 2-3. Test the full pre-seed→activate→reboot sequence on non-critical state before applying to important directories. |
0 commit comments