SOP · ENGINEERING
Comb Data Fabric Failover
ActiveEngineeringRev C
Purpose
To govern detection of and failover for the Comb data fabric that ingests and serves all NectarNet telemetry, colony ledgers, and mission data. The procedure ensures continuity of data services and integrity of records during node or site failure.
Scope
All failover, failback, and integrity verification of the Comb data fabric.
In scope:
- Health monitoring and failure detection
- Automated and manual failover to the standby site
- Data integrity and replication verification
- Controlled failback to the primary
Out of scope:
- Sensor node field commissioning, which follows NectarNet Sensor Node Deployment & Calibration
- Physical network hardware repair, which follows facility engineering work orders
Definitions
Comb Fabric
The distributed, replicated data platform storing all HIVE-13 operational records.
Failover
The promotion of the standby site to serve traffic when the primary is degraded.
Replication Lag
The delay between a write on the primary and its appearance on the standby.
Responsibilities
Data Platform Engineer (Owner)
Owns the procedure and authorises manual failover and failback decisions.
On-Call Operator
Monitors fabric health, initiates the runbook, and executes failover steps.
Director of Operations
Is notified of any failover and approves extended single-site operation.
Procedure
Detection
- Confirm the alert against fabric health dashboards and quorum status.
- Measure replication lag and assess primary recoverability.
- Declare an incident and open the failover runbook.
Failover
- Fence the degraded primary to prevent split-brain writes.
- Promote the standby site and redirect ingestion and query traffic.
- Verify telemetry and ledger writes resume against the new primary.
- Notify the Director of Operations of the failover state.
Integrity & Failback
- Reconcile record counts and checksums across replicas.
- Restore and resynchronise the recovered site before failback.
- Execute a controlled failback and confirm full replication before closing the incident.
Records Generated
- Failover incident and timeline log
- Data integrity reconciliation report
- Failback verification record
References
- NectarNet Sensor Node Deployment & Calibration
- Drone Swarm Pre-Flight & Tasking
- HIVE Data Continuity Doctrine
Revision History
| Rev | Date | Note |
|---|---|---|
| B | 2091-05-05 | Added primary fencing step to prevent split-brain writes. |
| C | 2091-06-19 | Required checksum reconciliation before incident closure. |