This is a read-only mirror of the HSC engineering wiki, restored from a 2017 archive. Some links are broken and some content is out of date. About this mirror.

Orchestrator Recovery Drill 2015

From ANTFARM Wiki

Jump to: navigation, search
This is an old revision of this page, as archived. The mirror serves the archived text for every revision id.
Orchestrator Recovery Drill 2015
Scenarioqueen split + bus partition
Duration2h 10m
Resultrunbook works, section 4 step 3 was wrong
Statuscurrent
Archive refQUANTARA-SWARMGLASS-R43-D3A60F

Contents

[hide]

Notes from the third yearly recovery drill. Scenario: a forced partition between racks with both queens up, then a bus restart during recovery (to see what happens when someone ignores section 4, step 2).

[edit] What we did

TimeStepOutcome
14:00partition rack B via switch ACLqueen02 elects itself in 31s, as designed
14:03operator A follows [[Orchestrator_Recovery#3._Queen_splitsection 3]]picks queen01 by epoch. correct.
14:05queend demote on queen02foragers drain in 40s
14:20heal partitionresync clean
14:25operator B restarts phero "by mistake"replay log survives restart (it's on disk). foragers reconnect with last-seq, no ctl.resync.
14:40queend reassign --stale 30011 nodes reissued, 0 duplicates because epochs
16:10write-upthis page

[edit] Findings

  1. Section 4 step 3 says "fix the network"; the drill showed the queen also needs queend reassign --stale 0 afterwards every time, not only on ctl.resync. Runbook updated.
  2. Restarting the broker was harmless because the replay log is persistent. The warning in the runbook is still right — it is harmless only because of the disk log, and AF-70 can fill that disk.
  3. Nobody could find the 2014 drill notes. They are gone.

[edit] Not linked from the runbook

This page was written after the operations portal merge and never linked from the runbook or the main page. The runbook's footer mentions it exists. svanterpool 2015-11-24

Revision 4 · svanterpool · history · alternates: txt
Personal tools