Skip to content

perf(workbook): remove two O(formula-cell-count) recalc floors (issue #991) - #995

Merged
hhimanshu merged 4 commits into
mainfrom
perf/991-recalc-floors
Sep 3, 2026
Merged

hhimanshu merged 4 commits into
mainfrom
perf/991-recalc-floors

Conversation

@hhimanshu

@hhimanshu hhimanshu commented Sep 3, 2026 •

Copy link
Copy Markdown
Member

Summary

Workbook::recalc_incremental_measured paid two costs proportional to
total formula-cell count, not to the dirty-closure size, on every call
regardless of how small the edit was:

  1. snapshot_formula_values cloned every formula cell's current value
    into a BTreeMap up front, purely so a widened spill retry could rewind
    to the pre-edit grid and so returned Change events could report correct
    pre-edit old values across multiple internal recompute passes.
  2. seed_spill_sensitive rebuilt an AuthoredCellIndex from scratch on
    every call that examined any range precedent, even though the
    authored-cell set rarely changes between calls.

Fix 1 (Design A): delete the upfront snapshot. Accumulate the pre-image
map lazily instead, folding it in first-wins from each recompute pass's own
returned Vec<Change> — apply_changes is the only code path that writes to
the grid, and every write it makes already carries the pre-write value as
old. The final diff step is unchanged, so the returned change list is
byte-identical to before; it's just computed without ever touching a formula
cell the edit didn't reach.

Fix 2 (a narrower, safer fallback — not the full "Design B"): cache just
the built AuthoredCellIndex on the workbook (authored_cell_index_cache.rs),
mirroring spill_anchor_cache.rs's own pattern. Invalidated only when the
authored-cell set actually changes (Workbook::set introducing a new cell,
Workbook::clear removing one) or a sheet-structure operation runs
(sheets_mut/sheet_mut/insert_sheet/remove_sheet/rename_sheet —
move_sheet is deliberately exempt, since reordering tabs changes neither the
folded-name keys nor any sheet's authored cells). apply_changes never
invalidates it: it only ever rewrites a cell that already carries formula
text, so it can change a value but never adds or removes an authored entry.

Also adds pre_image_stats.rs: exact-count instrumentation
(pre_image_count()) proving a one-cell edit into a 1,000-formula workbook
records exactly one pre-image, not one per formula cell.

Measured results

Apple M1 Max (10 core), macOS 14.4, rustc 1.98.0, release profile.
incremental_recalc/row_totals_volatile_seed, before (origin/main, median
of 3 real bench runs) vs after (this branch, fully rebased onto #946 — the
single combined run also recorded as this PR's baseline, see below):

Raw (includes the template clone the benchmark's timed closure pays
every iteration alongside set + recalc_incremental):

n (formula cells) before after change
100 251.8 µs 197.4 µs -21.6%
10,000 23.55 ms 16.46 ms -30.1%
100,000 290.1 ms 208.1 ms -28.3%

Clone-subtracted (subtracting row_totals_volatile_seed_clone_only, a
new control-group benchmark this PR adds at the same three scales, built
identically but doing nothing except template.clone(); before uses the
same after-run's clone-only numbers, since neither this fix nor #946 touches
clone cost):

n before after change
100 216.9 µs 162.5 µs -25.1%
10,000 17.69 ms 10.60 ms -40.1%
100,000 212.0 ms 129.9 ms -38.7%

Why both numbers matter: at n=100,000 the benchmark clones a 900,000-cell
workbook inside its timed closure on every iteration, and neither of this
PR's fixes touches clone cost at all — so the raw number understates what
this PR actually removed from recalc_incremental_measured itself. The
clone-subtracted number isolates that; the raw number is the direct answer to
"does calling set + recalc_incremental actually get faster." Both point
the same direction here, on a quiet machine with a single clean combined run.

Baseline update

crates/workbook/benches/baselines.json is re-recorded in this PR (two
chore(bench) commits — the branch had to be rebased mid-review, see below),
via check_perf_regression.py --record from a real, complete bench run of
the fully-rebased branch each time — not a hand-edit. It needed updating for
two independent reasons:

Deferred scope

seed_spill_sensitive's O(formula cells × precedents) sweep — checking
every formula cell's precedents for spill-sensitivity on every call — is
untouched by this PR. Fix 2 only removes the cost of building the
AuthoredCellIndex that sweep consults; the sweep itself still runs in full
every call.

A full "Design B" — caching the entire derived seed set
seed_spill_sensitive produces, not just the index it consults — was
considered and rejected for this PR as a strictly bigger, less-confident
correctness surface, and filed separately as
#993 for its own dedicated review cycle. Quoting the original
design investigation's own confidence caveats directly, since they're exactly
why this was deferred rather than rushed in here:

I am less confident in Design B than in Design A... a strictly bigger
correctness surface than #983 or #984... a missed invalidation is a wrong
answer, not a slow one... I'd want the property/differential suites run at
a raised seed budget before believing it.

Also out of scope, already tracked and merged separately at #992:
seed_spills_from_grid / GridSpillIndex::build's own full-grid-scan cost.

Tests

  • recalc_differential_tests.rs: existing random differential sweeps
    unchanged; new changes_report_true_pre_recalc_old_values_not_intermediate_widen_values
    pins that Change.old is always the true pre-recalc value, never an
    intermediate widen-pass value — with an explicit investigation note that
    the widen loop's rewind branch (pass > 0) could not be forced to execute
    anywhere in this suite (including a 3,000-seed sweep), so the invariant is
    asserted in its general form rather than gated on that branch.
  • recalc_work_tests.rs: new exact pre-image-count assertion.
  • spill_incremental_recalc_tests.rs: three new adversarial cases — a scalar
    edit creating a new spill several cells downstream, one removing an
    existing spill, and a from_json-loaded workbook recalculated
    incrementally on its very first call.
  • authored_cell_index_cache_tests.rs (new file): cold build, warm reuse,
    the negative "safe edit stays warm" case, one rebuild test per real
    invalidation condition, move_sheet's deliberate exemption, and a
    from_json cold-start case.
  • authored_cell_index_tests.rs and the full recalc_incremental(edits) == recalc() property suite pass unchanged.

Independent verification performed for this PR

  • cargo fmt --all -- --check on every touched file: clean
  • cargo clippy --workspace --exclude truecalc-python -- -D warnings: clean
  • cargo test --workspace --exclude truecalc-python: full suite green
  • Benchmarks re-run from scratch on the fully-rebased branch (numbers above)
  • python3 .github/scripts/check_perf_regression.py against the recorded
    baseline: Perf OK
  • python3 .github/scripts/test_check_perf_regression.py: all cases pass

closes #991

@hhimanshu hhimanshu self-assigned this Sep 3, 2026
@github-actions

github-actions Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Test Coverage by Category

Category Unit Tests Google Sheets Conformance Property Cases Total
Array 42 552/552 ✓ 1,000 (2×500) 1,594
Database 35 182/182 ✓ 3,500 (7×500) 3,717
Date 373 418/418 ✓ 2,500 (5×500) 3,291
Engineering 245 886/888 ⚠ 5,500 (11×500) 6,633
Filter 11 81/81 ✓ 4,500 (9×500) 4,592
Financial 149 1,208/1,208 ✓ 2,000 (4×500) 3,357
Info 0 256/256 ✓ 4,500 (9×500) 4,756
Logical 121 267/267 ✓ 3,500 (7×500) 3,888
Lookup 69 393/393 ✓ 1,000 (2×500) 1,462
Math 545 2,006/2,006 ✓ 8,000 (16×500) 10,551
Operator 87 251/251 ✓ 7,500 (15×500) 7,838
Parser 83 93/93 ✓ 4,000 (8×500) 4,176
Query 37 — — 37
Statistical 529 3,191/3,191 ✓ 5,000 (10×500) 8,720
Text 327 803/804 ⚠ 4,000 (8×500) 5,131
Timezone 47 — — 47
Volatile 0 — 3,500 (7×500) 3,500
Web 29 59/59 ✓ 6,000 (12×500) 6,088
Total 3,090 10,646/10,649 66,000 (132×500) ~79,739

✓ = 100% passing · ⚠ = known deviation · The ~79,739 total counts formula evaluations (each conformance row and each property case = 1). GitHub Checks reports 4,191 Rust test functions: 3,090 unit + 159 property functions (shown as cases above) + 942 conformance/integration.

hhimanshu and others added 4 commits September 3, 2026 12:59
…eed (issue #991 prereq)

incremental_recalc/row_totals_volatile_seed clones the full 900,000-cell
template inside its timed b.iter() closure, alongside the set + recalc_incremental
under test. At n=100,000 that clone is a real, possibly dominant, share of the
reported number, and neither of #991's two fixes touches clone cost at all.
Add row_totals_volatile_seed_clone_only at the same three scales so both fixes
can be judged on (raw − clone_only), not the raw number alone.
…991)

recalc_incremental_measured paid two costs proportional to total formula-cell
count on every call, regardless of how small the actual edit was:

1. snapshot_formula_values cloned every formula cell's value into a BTreeMap
   before the spill-widen loop ran, purely so a widened retry could rewind to
   the pre-edit grid and report correct pre-edit `old` values.

2. seed_spill_sensitive rebuilt its AuthoredCellIndex fresh on every call that
   touched any range precedent (~900,000-cell sweep at n=100,000 in the
   row_totals fixture), even though the authored-cell set itself rarely
   changes between calls.

Design A (fix 1): delete the upfront snapshot. Accumulate the pre-image map
lazily instead, folding it in first-wins from each recompute pass's own
returned Vec<Change> (apply_changes is the only code path that writes to the
grid, and every write it makes is recorded there with the pre-write value).
The final diff_against_snapshot step is unchanged and still runs, so the
returned change list is byte-identical to before, just computed without ever
touching a formula cell the edit didn't reach.

Fix 2 (safer fallback, not full "Design B"): cache only the built
AuthoredCellIndex on the workbook (authored_cell_index_cache.rs), mirroring
spill_anchor_cache.rs's own pattern, invalidated only when the authored-cell
SET actually changes (Workbook::set introducing a new cell, Workbook::clear
removing one) or a sheet-structure operation runs (sheets_mut/sheet_mut/
insert_sheet/remove_sheet/rename_sheet — not move_sheet, which changes
neither the folded-name keys nor any sheet's authored cells). Verified
apply_changes can never change the authored-cell set (it only rewrites cells
that already have formula text), so recalc's own value write-back needs no
invalidation. seed_spill_sensitive_built_index (the existing test
instrumentation in authored_cell_index_tests.rs) is deliberately left
uncached, so its "did this call need to build the index" answer keeps its
existing meaning regardless of the workbook cache's warmth.

A full "Design B" (caching seed_spill_sensitive's entire derived seed set,
not just the AuthoredCellIndex build) was considered and rejected for this PR
as a strictly bigger, less-confident correctness surface; tracked separately
as issue #992.

Also adds pre_image_stats.rs: an exact-count instrumentation counter
(pre_image_count()) proving a one-cell edit into a 1,000-formula workbook
records exactly one pre-image, not one per formula cell.

Measured on incremental_recalc/row_totals_volatile_seed, clone-subtracted:
  n=100:      208µs -> 188µs
  n=10,000:   17.2ms -> 11.2ms  (~35% faster)
  n=100,000:  229ms  -> 162ms   (~29% faster)

Ablating the AuthoredCellIndex-cache fallback alone (Design A kept) showed no
measurable difference on this fixture: the benchmark clones a fresh workbook
every iteration, so the cross-call cache this fallback targets is never warm
across more than one call in that harness, and the O(formula-cells x
range-width) sweep this fallback does not touch dominates what remains. A
direct probe issuing 20 repeated incremental calls on one live instance (no
inter-call clone) confirms the fallback behaves exactly as designed
(authored_index_builds() stays at 1 across all 20 calls) with a small,
noise-level wall-clock difference on this fixture's cost profile; its real
win is on workloads with a much larger authored-cell-to-formula-cell ratio,
or many more back-to-back incremental calls than this benchmark issues.

Out of scope, tracked at #985: seed_spills_from_grid / GridSpillIndex::build's
own remaining full-grid-scan cost on any workbook with at least one real
spill (only the all-empty case is handled there).

Tests: recalc_differential_tests.rs (existing sweeps unchanged, plus a new
named rewind test asserting Change.old always reflects the true pre-recalc
grid); recalc_work_tests.rs (new exact pre-image count assertion);
spill_incremental_recalc_tests.rs (three new adversarial cases: a scalar edit
creating a new spill several cells downstream, one removing an existing spill,
and a from_json-loaded workbook recalculated incrementally on its first call,
never having run a full recalc); authored_cell_index_cache_tests.rs (cold
build, warm reuse, the negative "safe edit stays warm" case, one rebuild test
per real invalidation condition, move_sheet's deliberate exemption, and a
from_json cold-start case). authored_cell_index_tests.rs and the full
recalc_incremental(edits) == recalc() property suite pass unchanged.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WfhRbND7tjFJ4JeGZ9gHL5
6061075fd's commit message says the deferred full "Design B" (caching
seed_spill_sensitive's entire derived seed set) is "tracked separately
as issue #992" — #992 is the #985 spill-scan-residual PR, not an issue.
The actual filed follow-up is #993 ("recalc: cache
seed_spill_sensitive's full derived seed set (Design B, deferred from
#991)"). Recording the correction here since repo convention is to add
a follow-up commit rather than rewrite an already-shared commit's
message.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WfhRbND7tjFJ4JeGZ9gHL5
#946 (spill_chain bench fixture) and its own baselines.json re-record
merged to main while this PR's CI was running, conflicting with the
baseline this PR had just recorded. Rebased onto origin/main and
recorded one fresh combined baseline from a real, complete bench run
of the fully-rebased code via check_perf_regression.py --record,
per repo convention — not a hand-merge of the two recordings.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WfhRbND7tjFJ4JeGZ9gHL5
@hhimanshu
hhimanshu force-pushed the perf/991-recalc-floors branch from 2b87f4d to 39a4fc6 Compare September 3, 2026 01:12
@github-actions

github-actions Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Test Coverage by Category

Category Unit Tests Google Sheets Conformance Property Cases Total
Array 42 552/552 ✓ 1,000 (2×500) 1,594
Database 35 182/182 ✓ 3,500 (7×500) 3,717
Date 373 418/418 ✓ 2,500 (5×500) 3,291
Engineering 245 886/888 ⚠ 5,500 (11×500) 6,633
Filter 11 81/81 ✓ 4,500 (9×500) 4,592
Financial 149 1,208/1,208 ✓ 2,000 (4×500) 3,357
Info 0 256/256 ✓ 4,500 (9×500) 4,756
Logical 121 267/267 ✓ 3,500 (7×500) 3,888
Lookup 69 393/393 ✓ 1,000 (2×500) 1,462
Math 545 2,006/2,006 ✓ 8,000 (16×500) 10,551
Operator 87 251/251 ✓ 7,500 (15×500) 7,838
Parser 83 93/93 ✓ 4,000 (8×500) 4,176
Query 37 — — 37
Statistical 529 3,191/3,191 ✓ 5,000 (10×500) 8,720
Text 327 803/804 ⚠ 4,000 (8×500) 5,131
Timezone 47 — — 47
Volatile 0 — 3,500 (7×500) 3,500
Web 29 59/59 ✓ 6,000 (12×500) 6,088
Total 3,090 10,646/10,649 66,000 (132×500) ~79,739

✓ = 100% passing · ⚠ = known deviation · The ~79,739 total counts formula evaluations (each conformance row and each property case = 1). GitHub Checks reports 4,192 Rust test functions: 3,090 unit + 159 property functions (shown as cases above) + 943 conformance/integration.

@hhimanshu
hhimanshu merged commit e72410e into main Sep 3, 2026
9 checks passed
@hhimanshu
hhimanshu deleted the perf/991-recalc-floors branch September 3, 2026 01:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

recalc_incremental pays two O(formula-cell-count) fixed costs every call: snapshot_formula_values and seed_spill_sensitive

1 participant