This guide defines the default testing layers for Alera and the commands that should be used before shipping features, UI changes, and refactors.
Native text rendering is covered by flutter test integration_test/typography_rendering_test.dart -d <linux|macos|windows>. Set ALERA_VISUAL_REVIEW_DIR to an absolute output directory to save PNG captures at 1x, 1.5x, and 2x text scales. Desktop Builds uploads these as typography-<platform> artifacts; review them alongside the goldens because the goldens obscure text and cannot validate native glyph rendering.
terminal_input_native_test.dart verifies terminal selection, copy/paste shortcuts, Unicode clipboard round-tripping and terminal key sequences through the native Flutter runner. It requires ALERA_NATIVE_TEST_CLIPBOARD=1 because it owns and clears the clipboard. Run it only on disposable CI desktops or a separate Xvfb display, never against a user's desktop session. Desktop Builds supplies this opt-in on its native runners.
- Unit tests cover pure domain logic, controllers, repositories, command construction, parsers, and platform branches with the smallest possible setup.
- Widget tests cover user-visible UI state, layout contracts, shortcuts, and interactions inside a focused widget tree.
- Golden tests use
alchemistto snapshot important UI states. They are best for design-system components and stable product surfaces where visual regressions matter. - E2E tests use Flutter
integration_testto run complete desktop flows through the real app shell with temporary storage and fake external boundaries. - Manual desktop builds still matter when a change touches packaging, release behavior, native plugins, platform file handling, or terminal process behavior.
After a batch that changes generated inputs or generators, run dart run build_runner build, dart tool/ci/normalize_generated_eof.dart, then the fast checks. From mobile/, use dart ../tool/ci/normalize_generated_eof.dart . after its own one-shot generation. Never run a build-runner watcher. The PR generation job repeats this sequence and rejects any difference from the committed result.
Run the fast checks first:
dart format --set-exit-if-changed lib test integration_test tool packages
dart run tool/quality/check_max_lines.dart
flutter analyze
flutter test --exclude-tags goldentool/quality/check_max_lines.dart enforces the AGENTS.md ~500-line file guidance as a ratchet: new files over the limit (or growth past a baseline entry) fail. Existing debt lives in tool/quality/max_lines_baseline.txt; refresh only with --write-baseline when intentionally accepting oversized files.
Run unit and widget tests with line coverage:
flutter test --coverage --exclude-tags golden
dart run tool/quality/coverage_report.dart --input coverage/lcov.info --min-lines 100 --worst 25Run golden tests:
flutter test --tags goldenUpdate golden files only when the visual change is intentional:
flutter test --update-goldens test/goldenRun desktop E2E locally on the current platform:
flutter test integration_test/alera_smoke_flow_test.dart -d macos
flutter test integration_test/rust_process_runner_test.dart -d macosUse -d linux or -d windows on those platforms. Invoke each integration file separately because the desktop launcher cannot reliably restart multiple suites in one invocation. The checked-in E2E flow must use temporary directories, temporary databases, fake process runners, and fake terminal runtimes unless the test explicitly needs a native boundary.
Shared packages under packages/alera_configuration each run dependency resolution from their lockfile, formatting, analysis and tests. Resolve both native plugin manifests independently without modifying Cargokit. The standalone runtime packager uses dart pub get from tool/release/runtime_packager; dart tool/ci/verify_runtime_packager.dart verifies all six archive variants with temporary inputs and a Dart-only dependency graph.
Mobile Builds validates changes to Flutter, mobile, and shared native dependencies on pull requests. It builds an arm64 Android release APK with the debug signing fallback, verifies its bundled native dependencies and 16 KB ELF alignment, and builds both an unsigned iOS device release and a debug simulator app on macOS. These checks use no distribution credentials and publish no applications. Desktop release compilation remains in the existing Desktop Builds workflow.
Run the Android equivalent from mobile/ with flutter build apk --release --target-platform android-arm64 -Pdisable-abi-filtering=true -PaleraAbiFilters=arm64-v8a, then run bash tool/release/verify_android_native_dependencies.sh mobile/build/app/outputs/flutter-apk from the repository root. --target-platform alone still packs plugin JNI for other ABIs; omit -PaleraAbiFilters and -Pdisable-abi-filtering only for emulator flutter run. On macOS, run flutter build ios --release --no-codesign and flutter build ios --simulator --debug --no-codesign from mobile/. Keep APKs signed with the debug fallback off release channels; they cannot update release-signed installations.
Run the Linux profile startup harness from a graphical Linux session:
make perf-linuxThe harness performs five launches, records startup marks plus first-frame build/raster/total timings, and writes raw samples with median, p95, p99, and median absolute deviation to .dart_tool/performance/startup_linux.json. Use dart tool/performance/alera_performance.dart --runs 5 --enforce to fail when p95 exceeds tool/performance/linux_startup_budget.json; keep the default local command report-only while hardware and runner variance are being calibrated. The Startup Performance workflow runs three samples under Xvfb on main and on its daily schedule as an informative, non-blocking smoke and uploads the JSON report.
For macOS CPU and memory profiling, launch make app-profile, then run PERF_SCENARIO=<name> PERF_APP_PID=<profile-pid> make perf-macos-resources from another terminal while exercising the scenario. The JSON report separates app, runtime host, Flutter tooling, code generation, terminal descendants, and agent CLIs. Use a 250 ms interval for short-lived provider processes and stop the build runner before the final comparison.
Compare measurements only on the same machine, power mode, display configuration, Flutter revision, and build mode. Run at least five samples for a decision, use median for the typical result, p95/p99 for tails, and MAD to spot noisy runs. Do not tighten the checked-in budget from a single capture.
tool/quality/coverage_report.dart reads coverage/lcov.info and enforces 100% line coverage for maintained domain sources under lib/src/features/**/domain/. Generated *.g.dart and *.mapper.dart files are excluded. Presentation code is validated by widget, golden, and desktop E2E suites; application and infrastructure code remains covered by focused unit and integration tests; generated flutter_rust_bridge bindings and the native implementation are validated by the Rust workspace and native build jobs.
When coverage drops, use the "worst files by missed lines" section to decide whether to add focused unit tests, widget tests, or an E2E path. Do not chase coverage by snapshotting implementation details; cover behavior that would catch a real regression.
Golden tests live under test/golden/ and use alchemist. The project config disables platform-readable goldens and keeps CI goldens stable across hosts. The first snapshots cover core design-system controls and the welcome dashboard in desktop and compact states.
Good golden candidates:
- Design-system components in
lib/src/design_system/. - Stable shell/dashboard states.
- Dialogs with meaningful layout variants.
- Error and empty states that are easy to regress visually.
Poor golden candidates:
- Highly animated or cursor-heavy states.
- Real terminal rendering.
- Native file picker flows.
- Screens that depend on wall-clock time, network data, or host fonts outside the configured test theme.
E2E tests live under integration_test/. They should prove full product flows that cross multiple widgets and application providers, such as adding a project, selecting a workspace, and opening terminal tabs.
Keep E2E tests deterministic:
- Use temporary project folders.
- Override
aleraDatabaseProviderwith a temporary or in-memory database. - Override runtime-backed repositories (
projectRepositoryProvider,workbenchRepositoryProvider,projectConfigRepositoryProvider,settingsRepositoryProvider) with Drift (or other in-process) implementations so the smoke flow does not require a live terminal-host sidecar. - Override
processRunnerProviderwhen a flow should not execute real commands. - Override
terminalRuntimeProviderwhen a flow only needs terminal UI behavior. - Avoid network access.
- Avoid native file pickers; paste paths directly into dialogs.
The runner's native regression tests use real Windows processes and windows without loading Flutter or user data. Run them from a shell with CMake and the Visual Studio C++ toolchain available:
cmake -S windows/runner/tests -B build/windows-runner-tests
cmake --build build/windows-runner-tests --config Debug
ctest --test-dir build/windows-runner-tests -C Debug --output-on-failureThe suite covers second launches before window creation, simultaneous cold launches, dev/release separation, restart after quit or crash, hidden and minimized window activation, maximization preservation, and synchronization failures. For the packaged app smoke, enable the tray, close the window, and launch the same shortcut again: the original process and tray icon must remain and its window must appear in front. Repeat while minimized and while already visible, then quit through the tray and confirm a fresh launch works. Keep Alera Dev open alongside Alera to verify flavor isolation.
Terminal persistence changes should include focused unit tests for the host client/session boundary and at least one manual or integration smoke on the current desktop platform: start a long-running terminal command, close Alera, reopen it, and confirm the terminal output continues under the same workspace tab. Explicit tab or workspace close must be checked separately because it should terminate the durable session instead of detaching.
Lifecycle changes must cover both host timeout paths. With the app closed and no running sessions, the host should stop after the configured empty-host delay. With the app closed and at least one running session, the host should keep the session alive until the configured detached-session delay, then terminate the PTY, write a final checkpoint, and delete host.json. Use small values from Settings during manual smoke tests so the behavior can be observed without waiting for the production defaults.
History and command-admission changes must test saturated work/control budgets, reliable output/resync timers, blocked storage without blocking the actor, ordered recovery of every accepted batch, idempotent retries after ambiguous commits, and configure/retention serialization. A graceful shutdown must retain sessions while history is pending, then finish after storage recovers. Shutdown and restart retries must preserve the first accepted exit intent. Oversized completion payloads must return a bounded error with the original request identity. Duplicate close requests must not multiply retry tasks, disconnect must release its pending barrier, and owner lifecycle operations must only be persisted after the history barrier succeeds. Relay fixtures must use the pinned workspace toolchain and retain workspace feature unification.
Scrollback changes must check both rendering and host memory behavior. The terminal row scrollback controls xterm history in the app. The host scrollback size controls how many bytes are retained for detached-session snapshots and checkpoint restore. Tests should prove incremental output chunks are trimmed to the configured byte limit, oversized chunks keep only their tail, shrinking the configured limit trims existing retained output, and checkpoints remain restorable after restart. Schema changes that intentionally discard old terminal history should include a focused test for the legacy checkpoint database shape being reset.
Output visibility and backpressure changes must prove the PTY keeps running while a hidden or slow terminal pauses only that client's output delivery. Cover the host protocol with two clients for the same session, confirm the paused client stops receiving output frames while another client continues, then resume and verify it is served only the bytes it missed, on the output lane and ahead of the resume reply, so the emulator is appended to rather than rebuilt. A saturated outbound queue must remain bounded, emit outputResyncRequired after capacity returns, and recover through the same delta path; only a client whose gap the ring has already dropped falls back to a full snapshot. Assert that a dropped frame does not advance the client's delivery cursor, since that gap is exactly what the next resume has to resend. Exit and error delivery should remain independent from output pause state.
For local sidecar smoke tests, build the Rust CLI sidecar with the makefile (which drives cargo and stages the binary):
make cli-build
make cli-helpmake cli-build runs cargo build --release -p alera-cli and stages the single binary into .dart_tool/alera/alera (.dart_tool/alera/alera.exe on Windows); make cli-help runs the staged binary's --help. The makefile debug targets (init-submodules, app-debug, cli-build, host-debug, and the rest) go through alera-xtask so they do not need a matching Dart SDK. The Rust workspace also has its own checks via make rust-test (cargo fmt --check, cargo clippy --workspace --all-targets -- -D warnings, cargo test --workspace).
make rust-test, the workspace test script and Linux CI checks use ALERA_BUILD_COMMIT=unknown so a new HEAD does not invalidate check binaries solely through their build stamp. make rust-test uses target-specific exports to clear an inherited ALERA_BUILD_VERSION and invokes Cargo directly, so it does not require Bash on Windows. The CI workspace script also clears an inherited version. Normal app and CLI builds still embed the real commit unless explicitly overridden; use those builds when diagnosing a running host. bash tool/ci/test_rust_build_cache.sh verifies both identity paths and the test runner's execution modes using disposable fixtures. bash tool/ci/run_rust_workspace_tests.sh --no-run compiles the same two test groups without running them; warm-rust-cache.yml uses this mode on pushes to main that change Rust sources or shared cache inputs. Rust-only pushes do not run the desktop warm-cache matrix; scheduled and manual warm-cache.yml runs still warm both desktop builds and Rust checks. This prepares compiler-cache inputs and does not replace test execution in PR checks. Compiler caching continues to use sccache.
The repository makefile exposes cross-platform debug targets around the same flow. make help lists available targets. For foreground host debugging, make host-debug accepts ALERA_HOST_EMPTY_SHUTDOWN_SECONDS, ALERA_HOST_DETACHED_SHUTDOWN_SECONDS, and ALERA_HOST_SCROLLBACK_BYTES, which are forwarded to the runtime host. alera terminal-host remains a compatibility alias, but new product behavior should be validated through alera runtime-host and the project, workspace, tag, tab, and ssh-target CLI groups.
Cross-version host conformance is intentionally one pinned Linux combination rather than a release matrix. tool/ci/host_compatibility.sh shallow-fetches tag v0.49.0 into a disposable repository, verifies that it resolves to commit e60c96ec7522052e9af81ab15ae5d6da2443dac4, and downloads the published alera-runtime-0.49.0 tarball after checking its pinned SHA-256. That binary reports crate version 0.1.0 and commit 17a183f51debfc29114c0e682bc917ed4cdc58ae on status.get (the parent of the release tag; ALERA_BUILD_VERSION was unset at packaging time). Product 0.49.0 is the tag plus the tarball, not runtimeHostVersion. Rebuilding that tag from source in CI compiled a second copy of alera-cli after the workspace tests and dominated the rust-test job. The ignored client test keeps --workspace so it does not relink after tool/ci/run_rust_workspace_tests.sh. The current protocol client then covers status.get, capability negotiation, agent profile upsert/list, and terminal launch/output. It also proves that the older host does not advertise agentProfileOrderingV1 and returns a named error if that newer verb is sent accidentally; runtime_agent_profile_repository_test.dart separately holds the Flutter client contract that capability absence produces a user-facing newer-host message without sending the unsupported verb. The normal Rust suite drives the current host with the same field set accepted by v0.49.0, covering the reverse direction without another build.
Run the historical combination locally from a checkout with access to origin:
bash tool/ci/host_compatibility.shRuntime-owned Projects, Workspaces, Tabs, Layouts, tags, relations, and SSH targets should include Rust store tests plus Dart repository/provider tests. Relation tests must cover self-link rejection, cycle prevention, cross-project links, tag assignment, and cascade previews for descendants and tags. SSH target tests should use fakes or local fixtures unless the test is explicitly marked as a manual remote-host smoke.
Alera currently favors small hand-written fakes for repositories, process runners, and terminal runtimes because those boundaries are domain-specific and easy to inspect. mocktail is still a good Dart package when a test needs many interaction assertions or when a collaborator has a broad interface that would make a fake noisy. Prefer explicit fakes for durable behavior tests and use mocks sparingly for call verification.
Run bash tool/ci/run_rust_workspace_tests.sh from the repository with inherited ALERA_* variables removed. Keep cargo commands inside rust/ so the pinned toolchain applies, and retain --workspace to avoid changing feature unification. The script runs PTY orchestration regressions separately with one test thread.
Admission regressions must preserve the existing 25 MiB dictation-audio contract after base64 encoding, keep small control requests available when large requests occupy work capacity, and reject requests larger than the whole work budget without waiting forever. Saturate real TCP and mobile WebSocket readers to prove they wait for capacity, preserve message order and notify disconnect on exit. Relay cancellation must retain a reserved disconnect slot even when other control admission is full. Preserve the existing 1 MiB encrypted relay envelope limit.
History tests must prove FIFO retries, idempotent exact replay, rejection of conflicting data at the same sequence, checkpoint/retention barriers and graceful close only after durable completion. Failed storage retains accepted batches and pauses the producer; forced process exit is not proof of persistence. Recovery fixtures require a fresh sidecar built with ALERA_BUILD_VERSION=isolated-test; build it after the default Rust suite so that version does not contaminate normal version assertions.
Remote-retirement regressions must preserve a working terminal after transport failure, remove only the confirmed task on retry, preserve shared files and keep the originally captured shutdown guard across history barriers. captured_shutdown_survives_history_retry_without_recapture checks guard ownership through deferred persistence, and the headless home-owner retirement fixture checks the real transport and processes.