Prove VMS archiver failover under full load—not just configuration state

A configured standby does not prove that all cameras will transfer, record, remain viewable, merge cleanly, and return under production load. Validate the exact failure, takeover, degraded operation, recovery, and evidence path.

Paired infrastructure paths converging on a stable recovered service.
DSE visual intelligenceContinuity & recoveryPlaybook · 3 min read
Executive summary

What you need to know

A configured standby does not prove that all cameras will transfer, record, remain viewable, merge cleanly, and return under production load. Validate the exact failure, takeover, degraded operation, recovery, and evidence path.

Potentially affected

VMS recording and archiving servers, standby or failover recorders, edge storage, camera streams, management services, storage, certificates, service accounts, networks, operators, retention, alarms, exports, and recovery procedures.

DSE recommendation

Document the supported failover design and capacity, establish pretest evidence and stop criteria, invoke only vendor-approved failure methods, measure every phase under representative load, verify recordings and exports, and close all gaps before relying on resilience claims.

Source facts: failover has product-specific behavior and limitations

Milestone Systems’ failover recording server documentation describes hot-standby and cold-standby operation, device-level choices for full, live-only, or no failover support, takeover, and later merging of recordings. It explicitly says failover does not provide complete redundancy and describes periods in which some older, archived, or failover-period recordings may be unavailable to clients.

The documentation also specifies a supported method for validating a merge: make the recording server unavailable by stopping its service or shutting down its computer. It says manually pulling a network cable or blocking the network with a test tool is not a valid method for that validation. The exact behavior applies to the documented Milestone product and version.

Axis similarly describes primary, redundant, and failover storage concepts, including edge or alternative storage when primary storage is unavailable. Different platforms use different triggers, delays, supported codecs, merge behavior, capacity rules, and licenses. RAID, server failover, camera edge recording, backup, and disaster recovery address different failure modes and must not be treated as interchangeable.

DSE recommendation: test every state the operator and investigator depend on

Start with an approved design record: protected cameras and streams, primary recorder, standby assignment, mode, maximum supported load, storage, edge-storage dependencies, management services, certificates, accounts, network paths, alarms, retention, and recovery owner. Compare deployed configuration with current vendor documentation and licensing before scheduling a test.

  1. Define success and stop conditions. Set maximum acceptable detection and takeover times, permitted recording gap, cameras that must remain live, required playback and export behavior, capacity thresholds, merge completion, and return-to-primary criteria. Establish when the test leader must restore service.
  2. Build representative load. Use ordinary production streams and an authorized window that reflects busy recording, operator viewing, analytics, and archive activity. A one-camera laboratory test cannot qualify a standby expected to receive hundreds of streams.
  3. Capture pretest evidence. Verify time, health, recording, retention, current alarms, storage headroom, service state, and a short retrievable clip from every priority class. Confirm backups and an out-of-band communication path.
  4. Invoke the documented failure. Follow the product’s supported validation procedure. Record the exact action and time. Observe detection, alarm delivery, stream reconnection, recording start, live viewing, playback, client errors, standby processor and storage load, and network utilization.
  5. Operate in the degraded state. Find and export events created after takeover. Test authorized user permissions, maps, bookmarks, audio and metadata where required. State clearly which historical or archived footage is unavailable rather than masking the limitation.
  6. Restore and reconcile. Return the primary system under the approved process. Measure handback, recording merge or edge retrieval, duplicate or missing intervals, timeline availability, retention calculations, alarms, and final standby readiness. Export samples that span before, during, and after the event.

Reconcile each camera against expected intervals; aggregate “server healthy” status is insufficient. Preserve logs and screenshots, note any deliberate test gap, and assign corrective work with retest criteria. Include a standby already occupied by another failed recorder, exhausted edge storage, a certificate problem, and management-service loss in later risk-based scenarios if the product supports those designs.

This test proves only the documented topology, load, version, and scenario. Repeat it after material camera growth, bitrate change, software upgrade, storage change, certificate or account change, network redesign, or failover reassignment. Resilience exists when usable recordings survive a controlled failure and recovery—not when the configuration page says they should.

Do not begin if the standby is already degraded, capacity is unknown, monitoring cannot distinguish primary from failover, or restoration authority is unavailable. Pause immediately if priority cameras stop without the approved alternate, storage approaches the defined safety threshold, or handback threatens existing evidence. Reschedule after the prerequisite is corrected; an unsafe resilience test creates the outage it was meant to prevent.

Official references

Primary reference

Review the official source

Milestone Systems: Failover recording server explained · Verified August 17, 2026

Open official reference ↗
Plan the next step

Need help applying this guidance safely?

DSE can help confirm applicability, protect service continuity, and validate the result across physical security and IT systems.

Talk with DSE