SQL Server Always On Replica Unhealthy Alert


Use the Always On Replica Unhealthy alert in Mini DBA to monitor SQL Server instances and make this condition visible before it becomes a wider database incident.

Screenshot pending: Mini DBA SQL Server Always On Replica Unhealthy alert screenshot placeholder

Alert Summary

  • Platform: SQL Server
  • Alert category: Availability
  • Default enabled: false
  • Default evaluation frequency: Minute

What Mini DBA Checks

Mini DBA describes this alert as: Reflects a rollup of the database synchronization state (synchronization_state)of all joined availability databases (also known as database replicas) and the availability mode of the availability replica (synchronous-commit or asynchronous-commit mode). The rollup will reflect the least healthy accumulated state the databases on the availability replica. Below are the possible values and their descriptions. Not healthy. At least one joined database is in the NOT SYNCHRONIZING state. Partially healthy. Some replicas are not in the target synchronization state: synchronous - commit replicas should be synchronized, and asynchronous - commit replicas should be synchronizing. Healthy.All replicas are in the target synchronization state: synchronous - commit replicas are synchronized, and asynchronous-commit replicas are synchronizing. The default evaluation frequency is Minute, so the alert is intended to be close enough to operational reality for live triage.

Why This Alert Is Helpful

This alert protects high availability and disaster recovery by warning when replicas, replication slots, or apply processes fall behind. Lag can turn a planned failover into data loss risk or leave reporting users looking at stale data.

When To Enable It

Enable it where replication, read replicas, failover, or reporting copies are part of the service design. Disable it on standalone systems that do not use replication so the alert list stays focused.

Threshold Guidance

This alert is not primarily driven by a numeric threshold in Mini DBA. Tune the schedule, scope, severity, and notification route so that the alert matches the importance of the instance. Use higher thresholds on batch-heavy, development, or intentionally bursty systems where brief pressure is expected. Use lower thresholds on latency-sensitive production systems, small instances with little headroom, and services with strict recovery or availability commitments.

Remediation For An Active Alert

Check the sender, receiver, apply process, network path, and oldest retained log position. Restart failed replication workers only after capturing the error, then address the root cause such as disk pressure, long transactions, schema drift, network latency, or insufficient replica capacity.

Investigation Workflow

  1. Confirm the alert is still active and note the first seen time, affected instance, and severity.
  2. Review the primary, replica, transport path, apply process, retained log position, and oldest open transaction in Mini DBA before changing configuration or ending sessions.
  3. Compare the current value with the normal baseline for the same time of day or maintenance window.
  4. Record the cause, corrective action, and whether thresholds or routing should be adjusted after the incident.

Avoiding Alert Noise

For recovery and replication alerts, route notifications to the people who own recovery objectives. Higher thresholds can be reasonable for reporting replicas, but production failover paths normally need tighter settings and explicit escalation.

Related Pages