Repository navigation
fix: harden performance and stability across the stack - #888
Conversation
There was a problem hiding this comment.
Important
Do not merge this as a stability hardening until the lifecycle, admission, and push-claim failures inline are fixed. Several are regressions from the new bounds, and the later ChatGPT commits do not address them.
Reviewed changes across the runtime, desktop, cloud, mobile, edge, updater, and landing gate, including the vendored updater fork.
- Runtime admission and history. Work and control commands are now capacity-limited, accepted terminal history is written in order, and a paused PTY session holds its reader until the durable buffer can take more bytes.
- Process capture. Native runs use a combined stdout/stderr budget, cancel by killing the tree, and on Windows assign the child to a job before resume.
- Desktop lifecycle. Explorer refreshes are generation-guarded, the editor rebuilds chrome only on a clean/dirty change, logs and diagnostics are budgeted, and the Linux and Windows runners release services on shutdown.
- Cloud push and schema. Each device is claimed in
delivery_attemptsbefore FCM, sends are capped at four inside a 12 second budget, and indexes 0005-0020 run after bind instead of blocking startup. - Mobile and edge. Malformed socket callbacks fail the transport, workspace attention is indexed, and each relay connection is capped at 16 MiB/s.
- Updater. A local
desktop_updater2.7.0 fork exposes cancellation and cleanup hooks, and artifact downloads abort on stall, oversize, or dispose. - ChatGPT options. Model discovery carries reasoning metadata, and thinking effort plus Normal/Fast are sent only when the running host advertises the capability.
- Release gate. A production landing build waits until the pinned public release assets exist.
ℹ️ Edge authorization limit is unchanged
edge/src/index.ts still keys the HTTP limiter on the SHA-256 of Authorization whenever that header is present, including an invalid token. Rotating fake credentials therefore gets a fresh bucket and is not limited by client IP. This pull request does not touch that function. It is recorded in docs/performance-stability-second-review.md and is still open.
grok-4.7 | 𝕏
There was a problem hiding this comment.
Important
Completion delivery can still stall after the queue drains, because one wakeup is shared across admission classes.
Reviewed changes since ba192b1: the follow-up commits repair the seven selected threads and record the validation.
- Hook shutdown. Native producer stop now runs before the FRB subscription cancel, and a gated test fails on the old order.
- AI process cleanup. Decode, stream, and output-limit failures kill the tree, cancel both readers, and reap before the original error propagates.
- Completion lane. Deferred results wait on a separate 128-command / 16 MiB lane instead of being discarded when control admission is full.
- Voice bytes. PCM counts against the work budget, and turn or synthesis results count as completions.
- Push claims. A fresh in-progress claim returns 503, persistence failures release the lease as
retryable, and quota is reserved once per claim. - Local artifact cancel. A
file:copy checks cancellation between chunks and restores the previous destination if cancel wins before the commit boundary. - Validation notes. The repair report records the commands and the limits that were not re-tested.
grok-4.7 | 𝕏
There was a problem hiding this comment.
✅ No new issues found.
Reviewed changes since f87fb053: one commit wakes every admission waiter when any class releases a slot.
- Admission wakeup.
releasenow usesCondvar::notify_allandNotify::notify_waiters, so a free completion slot is not stranded behind a control or work waiter that still cannot acquire. - Regression coverage. New tests fill all three lanes, park ineligible senders first, dequeue one completion, and require that completion to proceed while the others stay blocked.
grok-4.7 | 𝕏

What changed
Why
These changes reduce unbounded work, prevent stale or duplicate state from overwriting current state, and keep cancellation and shutdown paths from leaking resources or losing accepted data. Cloud delivery can recover interrupted requests without introducing an always-running retry worker, while online indexes improve query performance without blocking service startup. Release gating prevents landing pages from advertising assets before they are publicly available.
The ChatGPT controls preserve account- and model-specific behavior without hardcoded model assumptions or silent fallback, while capability gates prevent older runtimes from ignoring explicit options. The accompanying audit documentation distinguishes local validation from deployment, production verification, and live ChatGPT acceptance.