# Check protection and repair service state before expecting an unhealthy instance to recover

> Why can a scale-set instance remain unhealthy even though automatic repairs are enabled?

- Canonical URL: https://update.dsesecurity.com/updates/dse-20260909-140-check-protection-and-repair-service-state-before-expecting-an-unhealthy-instance/
- Publisher: Detection Systems & Engineering (DSE Security)
- Author: DSE Security Editorial Team
- Published: 2026-09-10T00:29:36+00:00
- Modified: 2026-09-10T00:52:39+00:00
- Last reviewed by DSE: 2026-09-09
- Resource type: Guide
- DSE priority: Information
- Topics: Business Continuity, IT
- Reading time: 2 minutes

## What you need to know

Why can a scale-set instance remain unhealthy even though automatic repairs are enabled?

## Potentially affected

Azure scale sets with automatic repairs and application health monitoring, excluding unsupported Service Fabric scale sets.

## DSE recommendation

Inspect instance protection, grace timing, and the repair service's live state before changing the repair policy.

## Article

## Source facts

Automatic repairs skip instances protected from either scale-in or scale-set actions. After a state-changing operation completes, the instance’s grace period delays repair. Repeatedly unhealthy replacements can cause the platform to suspend repairs as a safety measure. The live serviceState distinguishes Running, Suspended, and Not Running. Provisioning failures are outside automatic repairs’ supported recovery scope; instances must first initialize successfully. [Microsoft Learn](https://learn.microsoft.com/en-us/azure/virtual-machine-scale-sets/virtual-machine-scale-sets-automatic-instance-repairs).

## Applicability

Investigate a scale set with a configured, valid application health endpoint and an enabled repair policy. Confirm that it is not a Service Fabric scale set, which the source excludes. Identify whether the instance ever initialized successfully before interpreting its unhealthy state as an automatic-repair candidate.

## DSE recommendation

Inspect instance protection, grace timing, and the repair service’s live state before changing the repair policy. Ask the workload owner why protection exists and what must be preserved before removing it. For a suspended service, investigate the common reason replacement instances remain unhealthy before resuming. Keep an enabled model property separate from evidence that repair orchestration is currently running.

## Verification

Compare instance-view health and repair service state with the configuration and recent state-change timeline. Rehearse the intended recovery on a disposable instance only after the owner approves its repair action and data consequences. Observe whether the replacement or repaired instance becomes genuinely healthy. Retain any repeated failure as unresolved evidence rather than continually resuming repairs against an unchanged underlying problem.

## Official references

[Microsoft Learn: Automatic instance repairs for scale sets](https://learn.microsoft.com/en-us/azure/virtual-machine-scale-sets/virtual-machine-scale-sets-automatic-instance-repairs). Source reviewed September 9, 2026.

## Primary reference

- Name: Automatic instance repairs with Azure Virtual Machine Scale Sets - Azure Virtual Machine Scale Sets | Microsoft Learn
- Authority: Microsoft Learn
- URL: https://learn.microsoft.com/en-us/azure/virtual-machine-scale-sets/virtual-machine-scale-sets-automatic-instance-repairs
- Source publication date: Not stated by the source

## Citation and use

Preferred citation: “Check protection and repair service state before expecting an unhealthy instance to recover,” DSE Security, https://update.dsesecurity.com/updates/dse-20260909-140-check-protection-and-repair-service-state-before-expecting-an-unhealthy-instance/
Publishing principles: https://update.dsesecurity.com/updates/dse-updates-editorial-methodology/
Usage and citation policy: https://update.dsesecurity.com/usage/
Copyright © 2026 Detection Systems & Engineering. All rights reserved.
