Skip to content

fix: harden performance and stability across the stack - #888

Merged
leynier merged 27 commits into
mainfrom
perf/performance-stability-audit-release
Oct 2, 2026
Merged

leynier merged 27 commits into
mainfrom
perf/performance-stability-audit-release

Conversation

@leynier

@leynier leynier commented Oct 2, 2026

Copy link
Copy Markdown
Owner

What changed

  • Bound runtime admission, process capture, agent output, logging, diagnostics, updater transport, and relay resources.
  • Preserve ordered history, durable coordinator checkpoints, transport ownership, and compatibility under pressure or cancellation.
  • Serialize workspace/editor refreshes, cancel stale agent work, recover mobile transport, and improve workspace attention indexing.
  • Harden cloud push delivery with bounded parallelism, request deadlines, durable per-device claims, stale-claim recovery, account-scoped duplicate handling, and FCM cleanup.
  • Split required PostgreSQL migrations from allowlisted online performance indexes, with advisory-lock coordination, deadlines, checksum validation, dirty-state protection, and invalid-index repair.
  • Add ChatGPT model discovery, thinking-effort settings, Normal/Fast speed controls, runtime capability checks, and legacy-runtime compatibility handling.
  • Gate stable landing deployment on public release assets and keep package publication independent of non-critical cleanup failures.
  • Validate the vendored updater boundary in CI and add native workspace editor coverage to desktop builds.
  • Record audit findings, validation evidence, performance measurements, and known platform or live-acceptance limitations.

Why

These changes reduce unbounded work, prevent stale or duplicate state from overwriting current state, and keep cancellation and shutdown paths from leaking resources or losing accepted data. Cloud delivery can recover interrupted requests without introducing an always-running retry worker, while online indexes improve query performance without blocking service startup. Release gating prevents landing pages from advertising assets before they are publicly available.

The ChatGPT controls preserve account- and model-specific behavior without hardcoded model assumptions or silent fallback, while capability gates prevent older runtimes from ignoring explicit options. The accompanying audit documentation distinguishes local validation from deployment, production verification, and live ChatGPT acceptance.

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Important

Do not merge this as a stability hardening until the lifecycle, admission, and push-claim failures inline are fixed. Several are regressions from the new bounds, and the later ChatGPT commits do not address them.

Reviewed changes across the runtime, desktop, cloud, mobile, edge, updater, and landing gate, including the vendored updater fork.

  • Runtime admission and history. Work and control commands are now capacity-limited, accepted terminal history is written in order, and a paused PTY session holds its reader until the durable buffer can take more bytes.
  • Process capture. Native runs use a combined stdout/stderr budget, cancel by killing the tree, and on Windows assign the child to a job before resume.
  • Desktop lifecycle. Explorer refreshes are generation-guarded, the editor rebuilds chrome only on a clean/dirty change, logs and diagnostics are budgeted, and the Linux and Windows runners release services on shutdown.
  • Cloud push and schema. Each device is claimed in delivery_attempts before FCM, sends are capped at four inside a 12 second budget, and indexes 0005-0020 run after bind instead of blocking startup.
  • Mobile and edge. Malformed socket callbacks fail the transport, workspace attention is indexed, and each relay connection is capped at 16 MiB/s.
  • Updater. A local desktop_updater 2.7.0 fork exposes cancellation and cleanup hooks, and artifact downloads abort on stall, oversize, or dispose.
  • ChatGPT options. Model discovery carries reasoning metadata, and thinking effort plus Normal/Fast are sent only when the running host advertises the capability.
  • Release gate. A production landing build waits until the pinned public release assets exist.

ℹ️ Edge authorization limit is unchanged

edge/src/index.ts still keys the HTTP limiter on the SHA-256 of Authorization whenever that header is present, including an invalid token. Rotating fake credentials therefore gets a fresh bucket and is not limited by client IP. This pull request does not touch that function. It is recorded in docs/performance-stability-second-review.md and is still open.

Pullfrog  | Fix all ➔ | Fix 👍s ➔ | View workflow run | Using grok-4.7 | 𝕏

Comment thread lib/src/features/agent_status/infra/agent_hook_receiver.dart
Comment thread lib/src/features/ai_assist/application/ai_assist_agent_runner.dart Outdated
Comment thread rust/alera-cli/src/terminal_host/server/server_command_inbox.rs Outdated
Comment thread rust/alera-cli/src/terminal_host/server/server_command_inbox.rs Outdated
Comment thread cloud/src/push_delivery.rs
Comment thread cloud/src/push_delivery.rs Outdated
Comment thread third_party/desktop_updater/lib/src/io/file_update_transport.dart

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Important

Completion delivery can still stall after the queue drains, because one wakeup is shared across admission classes.

Reviewed changes since ba192b1: the follow-up commits repair the seven selected threads and record the validation.

  • Hook shutdown. Native producer stop now runs before the FRB subscription cancel, and a gated test fails on the old order.
  • AI process cleanup. Decode, stream, and output-limit failures kill the tree, cancel both readers, and reap before the original error propagates.
  • Completion lane. Deferred results wait on a separate 128-command / 16 MiB lane instead of being discarded when control admission is full.
  • Voice bytes. PCM counts against the work budget, and turn or synthesis results count as completions.
  • Push claims. A fresh in-progress claim returns 503, persistence failures release the lease as retryable, and quota is reserved once per claim.
  • Local artifact cancel. A file: copy checks cancellation between chunks and restores the previous destination if cancel wins before the commit boundary.
  • Validation notes. The repair report records the commands and the limits that were not re-tested.

Pullfrog  | Fix all ➔ | Fix 👍s ➔ | View workflow run | Using grok-4.7 | 𝕏

Comment thread rust/alera-cli/src/terminal_host/server/server_command_inbox_admission.rs Outdated

@pullfrog pullfrog Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ No new issues found.

Reviewed changes since f87fb053: one commit wakes every admission waiter when any class releases a slot.

  • Admission wakeup. release now uses Condvar::notify_all and Notify::notify_waiters, so a free completion slot is not stranded behind a control or work waiter that still cannot acquire.
  • Regression coverage. New tests fill all three lanes, park ineligible senders first, dequeue one completion, and require that completion to proceed while the others stay blocked.

Pullfrog  | View workflow run | Using grok-4.7 | 𝕏

@leynier
leynier merged commit fb36942 into main Oct 2, 2026
30 checks passed
@leynier
leynier deleted the perf/performance-stability-audit-release branch October 2, 2026 17:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant