fix(restore): resume follow mode when saved TXID is ahead of latest s…#1385
Open
atmin wants to merge 1 commit into
Open
fix(restore): resume follow mode when saved TXID is ahead of latest s…#1385atmin wants to merge 1 commit into
atmin wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Follow-mode crash-recovery resume validated the saved
-txidagainst the latest snapshot'sMaxTXIDand refused to resume when the saved TXID was greater. Compare against the newest TXID across all levels instead, so a follow file that is ahead of the last snapshot but still backed by existing deltas resumes correctly. A stale/foreign-txidpast all replicated data is still rejected; the "behind the earliest snapshot" check is unchanged.Motivation and Context
A live follow file continuously applies level-0 deltas past the latest snapshot, while snapshots are only written every
SnapshotInterval(default 24h). So the saved-txidis almost always ahead of the last snapshot, and resume was rejected in the normal case — includinglitestream restore -followcrash recovery — forcing a full re-restore instead of an incremental catch-up.How Has This Been Tested?
New regression test
TestReplica_Restore_Follow_ResumeAheadOfSnapshot(fails before this change withsaved TXID … is ahead of latest snapshot, passes after):All pass locally (full race suite green across all packages;
go fmt,go vet,goimports,staticcheckon the changed files clean).Types of changes
Checklist
go fmt,go vet)go test ./...)