Skip to content

branch-3.1: [fix](cloud) Refresh base tablet before schema change V1 #64312 - #67608

Open
Yukang-Lian wants to merge 1 commit into
apache:branch-3.1from
Yukang-Lian:codex/pick-64312-doris-3.1-20260907
Open

branch-3.1: [fix](cloud) Refresh base tablet before schema change V1 #64312#67608
Yukang-Lian wants to merge 1 commit into
apache:branch-3.1from
Yukang-Lian:codex/pick-64312-doris-3.1-20260907

Conversation

@Yukang-Lian

Copy link
Copy Markdown
Collaborator

What problem does this PR solve?

Issue Number: None

Related PR: #64312

Problem Summary: Backport the Cloud schema-change base-tablet refresh fix from #64312 to branch-3.1.

Before capturing rowset readers and registering schema change V1, the base tablet now performs an uncapped rowset refresh. This prevents a stale request alter version from hiding newer visible rowsets across retries. The implementation is adapted to this branch's older CloudTablet::sync_rowsets() API.

The master regression fixture depends on later cross-V1 test infrastructure that is absent from this branch, so this backport keeps the production fix and its diagnostic debug point without importing unsupported test scaffolding.

The branch is rebased onto the latest branch-3.1.

Release note

Refresh the Cloud schema-change base tablet before calculating and registering V1.

Check List (For Author)

  • Test: Manual / static validation
    • clang-format 16 --dry-run --Werror: passed for all changed C++ files
    • git diff --check: passed
    • Exact-base linked BE tests were not run locally; hosted CI is requested with run buildall
  • Behavior changed: Yes
    • Schema change uses the latest visible base-tablet rowsets instead of a query-version-capped refresh
  • Does this need documentation: No

Related PR: apache#62272

Problem Summary: Cloud schema change retries could keep registering a
stale alter version after a cross-V1 compaction conflict. The retry path
synced the base tablet only up to the request alter version, so a local
base tablet whose max version already exceeded that request could skip
refreshing newer visible rowsets. When a new schema-change tablet had a
compaction rowset crossing the stale V1, every retry reused that V1 and
eventually exhausted the retry limit.

This change refreshes the base tablet without capping sync by the
request alter version before computing schema change V1, allowing
retries to register a fresh boundary. It also adds a docker regression
case that blocks schema change, creates a cross-V1 retry, injects stale
local max-version state for query-version sync, and verifies the retry
finishes after the base tablet is refreshed.

Fix cloud schema change retry stale V1 failure.
@hello-stephen

Copy link
Copy Markdown
Contributor

Thank you for your contribution to Apache Doris.
Don't know what should be done next? See How to process your PR.

Please clearly describe your PR:

  1. What problem was fixed (it's best to include specific error reporting information). How it was fixed.
  2. Which behaviors were modified. What was the previous behavior, what is it now, why was it modified, and what possible impacts might there be.
  3. What features were added. Why was this function added?
  4. Which code was refactored and why was this part of the code refactored?
  5. Which functions were optimized and what is the difference before and after the optimization?

@Yukang-Lian

Copy link
Copy Markdown
Collaborator Author

run buildall

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants