This is a read-only mirror of the HSC engineering wiki, restored from a 2017 archive. Some links are broken and some content is out of date. About this mirror.
Orchestrator Recovery Drill 2015
From ANTFARM Wiki
| Orchestrator Recovery Drill 2015 | |
|---|---|
| Scenario | queen split + bus partition |
| Duration | 2h 10m |
| Result | runbook works, section 4 step 3 was wrong |
| Status | current |
| Archive ref | QUANTARA-SWARMGLASS-R43-D3A60F |
Contents[hide] |
Notes from the third yearly recovery drill. Scenario: a forced partition between racks with both queens up, then a bus restart during recovery (to see what happens when someone ignores section 4, step 2).
[edit] What we did
| Time | Step | Outcome | |
|---|---|---|---|
| 14:00 | partition rack B via switch ACL | queen02 elects itself in 31s, as designed | |
| 14:03 | operator A follows [[Orchestrator_Recovery#3._Queen_split | section 3]] | picks queen01 by epoch. correct. |
| 14:05 | queend demote on queen02 | foragers drain in 40s | |
| 14:20 | heal partition | resync clean | |
| 14:25 | operator B restarts phero "by mistake" | replay log survives restart (it's on disk). foragers reconnect with last-seq, no ctl.resync. | |
| 14:40 | queend reassign --stale 300 | 11 nodes reissued, 0 duplicates because epochs | |
| 16:10 | write-up | this page |
[edit] Findings
- Section 4 step 3 says "fix the network"; the drill showed the queen also needs
queend reassign --stale 0afterwards every time, not only onctl.resync. Runbook updated. - Restarting the broker was harmless because the replay log is persistent. The warning in the runbook is still right — it is harmless only because of the disk log, and AF-70 can fill that disk.
- Nobody could find the 2014 drill notes. They are gone.
[edit] Not linked from the runbook
This page was written after the operations portal merge and never linked from the runbook or the main page. The runbook's footer mentions it exists. — svanterpool 2015-11-24