Summary
With --gtid --checkpoint, a checkpoint can be written while a multi-row transaction is only partly applied to the
ghost table. --resume then starts the binlog stream from a GTID set that already contains that transaction, so the
server never sends it again and the rest of its rows are silently lost. The migration completes normally (exit code
0, Done migrating) and the cut-over puts the stale ghost table live.
How it happens (v1.1.11)
GoMySQLReader advances its coordinates on the GTIDEvent, i.e. before the transaction's row events
(go/binlog/gomysql_reader.go, case *replication.GTIDEvent): the running GTID set already includes the
transaction being streamed.
Migrator.Checkpoint() takes eventsStreamer.GetCurrentBinlogCoordinates() as LastTrxCoords and writes the
checkpoint as soon as coords.SmallerThanOrEquals(applier.CurrentCoordinates) (go/logic/migrator.go, around
line 1815).
- The applier sets
CurrentCoordinates = eventStruct.coords after every DML batch (onApplyEventStruct, around line
1797) — after the first batch of a large transaction its coordinates already equal the streamer's.
- So a checkpoint taken during a large transaction records its GTID after only
--dml-batch-size rows were applied.
On --resume, InitialStreamerCoords = lastCheckpoint.LastTrxCoords and the replica stream starts with that set:
the transaction is considered executed and is skipped entirely.
Reproduction
MySQL 9.7 (Percona Server 9.7.1-1; nothing version specific is involved), single node, GTID on, a table t of
250,000 rows (id PK, chat_id, v), gh-ost v1.1.11 built from the tag:
gh-ost ... --database=vt --table=t --alter="ADD INDEX chat_v (chat_id, v)" \
--allow-on-master --assume-rbr --gtid --checkpoint --checkpoint-seconds=10 --dml-batch-size=10 \
--cut-over=atomic --postpone-cut-over-flag-file=/tmp/postpone --execute
- Let the row copy complete (cut-over postponed).
- Run one big transaction:
UPDATE vt.t SET v = v + 1000; (271,299 rows, GTID …:999999).
- Wait until the newest row of
_t_ghk has gh_ost_chk_coords containing that GTID while _t_gho has only part of
the update (here: checkpoint …:1-53495:999999 written when 14,360 of 271,299 rows were updated in _t_gho).
kill -9 gh-ost.
- Restart the same command with
--resume:
Resuming from checkpoint coords=…:1-53495:999999 range_min=262190 range_max=262190 iteration=249.
- Compare: 256,889 rows of
_t_gho still have the old v while t has the new one; no rows are missing.
- Remove the postpone flag: cut-over succeeds, exit code 0 — the live table now has 256,889 rows with the update
lost (the old table _t_<ts>_del has them right).
Before the kill, during the copy with concurrent inserts/updates/deletes, t and _t_gho matched by checksum, so
the loss comes from the resume only.
Impact
Silent data loss on resume after any crash or kill that happens while a large transaction is being applied: bulk
UPDATE/DELETE, backfills, archival jobs. A lost DELETE brings deleted rows back.
Possible fix
Checkpoint only on a transaction boundary: record the coordinates of the last transaction applied completely (e.g.
advance the applier's checkpointable coordinates on the transaction's XID/commit event rather than per batch), or
make Checkpoint() use the coordinates of the previous transaction while one is in progress. Until then, a note in
doc/resume.md that --resume is unsafe after a crash in the middle of a large transaction would help.
Related: #1780 (a different cut-over issue found while testing the same setup).
Summary
With
--gtid --checkpoint, a checkpoint can be written while a multi-row transaction is only partly applied to theghost table.
--resumethen starts the binlog stream from a GTID set that already contains that transaction, so theserver never sends it again and the rest of its rows are silently lost. The migration completes normally (exit code
0,
Done migrating) and the cut-over puts the stale ghost table live.How it happens (v1.1.11)
GoMySQLReaderadvances its coordinates on theGTIDEvent, i.e. before the transaction's row events(
go/binlog/gomysql_reader.go,case *replication.GTIDEvent): the running GTID set already includes thetransaction being streamed.
Migrator.Checkpoint()takeseventsStreamer.GetCurrentBinlogCoordinates()asLastTrxCoordsand writes thecheckpoint as soon as
coords.SmallerThanOrEquals(applier.CurrentCoordinates)(go/logic/migrator.go, aroundline 1815).
CurrentCoordinates = eventStruct.coordsafter every DML batch (onApplyEventStruct, around line1797) — after the first batch of a large transaction its coordinates already equal the streamer's.
--dml-batch-sizerows were applied.On
--resume,InitialStreamerCoords = lastCheckpoint.LastTrxCoordsand the replica stream starts with that set:the transaction is considered executed and is skipped entirely.
Reproduction
MySQL 9.7 (Percona Server 9.7.1-1; nothing version specific is involved), single node, GTID on, a table
tof250,000 rows (
idPK,chat_id,v), gh-ost v1.1.11 built from the tag:UPDATE vt.t SET v = v + 1000;(271,299 rows, GTID…:999999)._t_ghkhasgh_ost_chk_coordscontaining that GTID while_t_ghohas only part ofthe update (here: checkpoint
…:1-53495:999999written when 14,360 of 271,299 rows were updated in_t_gho).kill -9gh-ost.--resume:Resuming from checkpoint coords=…:1-53495:999999 range_min=262190 range_max=262190 iteration=249._t_ghostill have the oldvwhilethas the new one; no rows are missing.lost (the old table
_t_<ts>_delhas them right).Before the kill, during the copy with concurrent inserts/updates/deletes,
tand_t_ghomatched by checksum, sothe loss comes from the resume only.
Impact
Silent data loss on resume after any crash or kill that happens while a large transaction is being applied: bulk
UPDATE/DELETE, backfills, archival jobs. A lostDELETEbrings deleted rows back.Possible fix
Checkpoint only on a transaction boundary: record the coordinates of the last transaction applied completely (e.g.
advance the applier's checkpointable coordinates on the transaction's
XID/commit event rather than per batch), ormake
Checkpoint()use the coordinates of the previous transaction while one is in progress. Until then, a note indoc/resume.mdthat--resumeis unsafe after a crash in the middle of a large transaction would help.Related: #1780 (a different cut-over issue found while testing the same setup).