Skip to content

Wait up to 10 minutes for a translation job, not 60 polls - #125

Merged
arvida merged 1 commit into
mainfrom
fix/job-poll-budget
Oct 9, 2026
Merged

arvida merged 1 commit into
mainfrom
fix/job-poll-budget

Conversation

@arvida

@arvida arvida commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Problem

The CLI gave up on a translation job after 60 status checks of about 3 seconds each, so roughly 3 minutes, and recorded that locale as failed. The 10-minute limit in monitorJobStatus never fired: its start time reset on every check, and so did the backoff (always 1s). Jobs whose validation takes longer, which the backend allows for larger jobs, were dropped even when they completed later.

Fix

  • A job is polled until 10 minutes after the CLI started monitoring the batch, however many checks that takes (JOB_WAIT_MINUTES). Past that it takes the same give-up path as before: final progress check, same messages, locale added to failedLanguages.
  • Checks run every 3 seconds for the first minute (as before), then every 10 seconds.
  • monitorJobStatus does one check and returns; the do/while that never looped, its startTime/retries and the outer 2s wait are gone.
  • A failed status still fails on the first check.
  • One message changed: "exceeded maximum retries (60)" is now "exceeded maximum wait (10 minutes)".

Open question

Server-side, a large job's validation plus a retry cycle can exceed 10 minutes (initial cycle 2s per validation, retry cycle ~60s per candidate). 15 minutes may be the safer budget; kept at 10 here to match the existing intent.

Backend side: localheroai/localhero-ai#884.

Testing

  • New: a job that stays validating for 5 minutes of fake time, then completes, is applied; waits are 3s then capped at 10s.
  • Rewritten: the give-up test asserts the new limit (10:00-10:10), the message and failedLanguages.
  • Full suite 1496 tests passing, lint and tsc clean. Codex review: no findings.

- The poll cap was 60 checks of about 3 seconds, so the CLI gave up on a
  job after roughly 3 minutes and recorded its locale as failed. The
  10-minute limit in monitorJobStatus never fired: its start time reset on
  every check
- A job is now polled until 10 minutes after the CLI started monitoring
  it, however many checks that takes. Past that it takes the same
  give-up path as before; the message names the wait instead of a count
- Checks run every 3 seconds for the first minute, as before, then every
  10 seconds
- A failed status still fails at once
@arvida
arvida merged commit 97464f1 into main Oct 9, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant