Skip to content

--resume with --gtid can skip the rest of a large transaction (checkpoint written mid-transaction): silent data loss #1781

Description

@kotigor

Summary

With --gtid --checkpoint, a checkpoint can be written while a multi-row transaction is only partly applied to the
ghost table. --resume then starts the binlog stream from a GTID set that already contains that transaction, so the
server never sends it again and the rest of its rows are silently lost. The migration completes normally (exit code
0, Done migrating) and the cut-over puts the stale ghost table live.

How it happens (v1.1.11)

  • GoMySQLReader advances its coordinates on the GTIDEvent, i.e. before the transaction's row events
    (go/binlog/gomysql_reader.go, case *replication.GTIDEvent): the running GTID set already includes the
    transaction being streamed.
  • Migrator.Checkpoint() takes eventsStreamer.GetCurrentBinlogCoordinates() as LastTrxCoords and writes the
    checkpoint as soon as coords.SmallerThanOrEquals(applier.CurrentCoordinates) (go/logic/migrator.go, around
    line 1815).
  • The applier sets CurrentCoordinates = eventStruct.coords after every DML batch (onApplyEventStruct, around line
    1797) — after the first batch of a large transaction its coordinates already equal the streamer's.
  • So a checkpoint taken during a large transaction records its GTID after only --dml-batch-size rows were applied.
    On --resume, InitialStreamerCoords = lastCheckpoint.LastTrxCoords and the replica stream starts with that set:
    the transaction is considered executed and is skipped entirely.

Reproduction

MySQL 9.7 (Percona Server 9.7.1-1; nothing version specific is involved), single node, GTID on, a table t of
250,000 rows (id PK, chat_id, v), gh-ost v1.1.11 built from the tag:

gh-ost ... --database=vt --table=t --alter="ADD INDEX chat_v (chat_id, v)" \
  --allow-on-master --assume-rbr --gtid --checkpoint --checkpoint-seconds=10 --dml-batch-size=10 \
  --cut-over=atomic --postpone-cut-over-flag-file=/tmp/postpone --execute
  1. Let the row copy complete (cut-over postponed).
  2. Run one big transaction: UPDATE vt.t SET v = v + 1000; (271,299 rows, GTID …:999999).
  3. Wait until the newest row of _t_ghk has gh_ost_chk_coords containing that GTID while _t_gho has only part of
    the update (here: checkpoint …:1-53495:999999 written when 14,360 of 271,299 rows were updated in _t_gho).
  4. kill -9 gh-ost.
  5. Restart the same command with --resume:
    Resuming from checkpoint coords=…:1-53495:999999 range_min=262190 range_max=262190 iteration=249.
  6. Compare: 256,889 rows of _t_gho still have the old v while t has the new one; no rows are missing.
  7. Remove the postpone flag: cut-over succeeds, exit code 0 — the live table now has 256,889 rows with the update
    lost (the old table _t_<ts>_del has them right).

Before the kill, during the copy with concurrent inserts/updates/deletes, t and _t_gho matched by checksum, so
the loss comes from the resume only.

Impact

Silent data loss on resume after any crash or kill that happens while a large transaction is being applied: bulk
UPDATE/DELETE, backfills, archival jobs. A lost DELETE brings deleted rows back.

Possible fix

Checkpoint only on a transaction boundary: record the coordinates of the last transaction applied completely (e.g.
advance the applier's checkpointable coordinates on the transaction's XID/commit event rather than per batch), or
make Checkpoint() use the coordinates of the previous transaction while one is in progress. Until then, a note in
doc/resume.md that --resume is unsafe after a crash in the middle of a large transaction would help.

Related: #1780 (a different cut-over issue found while testing the same setup).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions