Repository navigation
Migration fails with "ghc table does not exist" #1622
Description
Activity
Ran into the same exact issue today, was there any investigation into the root cause? Thanks!
+1 to " not running in either "revert" or "resume" mode"
@ajm188 @saimitta-looker what gh-ost version are you using? I couldn't reproduce this. Does it also error with
--execute?I was using the 1.18 release candidate. If I get some time this week I can try with 1.18 GA to see if it is happening there as well.
Unfortunately I don't remember if this also errored with
--executebut I will try to remember to grab that info next time aroundHas this issue been fixed?
#1664this happened again yesterday, version 1.1.11. i'm going to poke around a bit at logs to see if i can find anything useful and will report back!
i'm back, and i think this is an intermittent but ultimately harmless error log noise. it's a race between the cleanup and the throttler's lag collection goroutine.
The race between
Throttler.collectReplicationLagandMigrator.finalCleanup:collectReplicationLag's ticker (default 100ms) spawnsgo collectFunc()on every tick,
gated only onthlr.finishedMigrating— which is set inThrottler.Teardown(), called from
Migrator.teardown()afterfinalCleanup()(and the_ghcdrop) has already run.collectFunc()does checkCleanupImminentFlagbefore querying, but it's a check-then-act
race: if the goroutine reads the flag as0a moment beforefinalCleanupsets it, it
proceeds anyway.- That query (
readChangelogState("heartbeat")) runs on the inspector connection, i.e. the
read replica — not the primary whereDROP TABLE _ghcactually executes. So the race window
is the full replication lag needed for the drop to propagate to the replica. - In my most recent example,
HeartbeatLagspikes from0.06sto1.07sat the exact moment of the
error, and that spike gave the stale in-flight SELECT enough time to land on the replica
after the drop had already replicated there.
it's pretty rare since it needs a
collectFunc()goroutine spawned in the small window right beforeCleanupImminentFlagis set, and enough replica lag for the drop to beat it there.Planning to look at a fix (tighter check in the collection loop, and/or having
finalCleanupwait out in-flight throttler queries before dropping_ghc).
This was running off of
masteragainst an 8.4.6 database.Interestingly, I was not running in either "revert" or "resume" mode, so I have no idea why gh-ost didn't actually try to create the changelog table. After the gh-ost process exited, there were no gh-ost tables on the database at all: