Problem overview
Microsoft Exchange powers business-critical email communications. Service disruptions can cause missed emails and productivity losses. This policy ensures Exchange services remain operational by automatically restarting stopped instances and notifying your team immediately.
Monitors
- Exchange CPU Usage (CPUUsed monitor) Exchange CPU Usage (CPUUsed monitor)
- Exchange Information Store (Service monitor) Exchange CPU Usage (CPUUsed monitor)
- Exchange Transport (Service monitor) Exchange CPU Usage (CPUUsed monitor)
- Exchange Frontend Transport (Service monitor) Exchange CPU Usage (CPUUsed monitor)
- Exchange Mailbox Transport Submission (Service monitor) Exchange CPU Usage (CPUUsed monitor)
- Exchange Mailbox Transport Delivery (Service monitor) Exchange CPU Usage (CPUUsed monitor)
- Exchange Health Manager (Service monitor) Exchange CPU Usage (CPUUsed monitor)
- Exchange Search (Service monitor) Exchange CPU Usage (CPUUsed monitor)
- Exchange Replication (Service monitor) Exchange CPU Usage (CPUUsed monitor)
- Exchange DAG Management (Service monitor) Exchange CPU Usage (CPUUsed monitor)
This monitoring policy targets devices tagged with “Exchange” and checks the status of Microsoft Exchange services. If the service stops or fails to start, it automatically triggers a restart and sends a real-time alert. This reduces downtime and ensures uninterrupted email functionality for your users.
Use cases
- Ensuring uptime for on-premises Microsoft Exchange servers.
- Proactively monitoring critical email and calendaring services.
- Supporting business continuity in hybrid Exchange environments.
- Avoiding SLA breaches related to email availability.
Recommendations
- Tagging: Apply the “Exchange” tag to all relevant devices. We recommend automatically tagging to avoid missing key devices. See “Service Based Tagging” automation as an example.
- Testing: Simulate service failures to validate the monitor’s ability to restart services and send alerts.
- Regular Updates: Keep Exchange servers updated to avoid service interruptions caused by outdated software.
- Resource Monitoring: Pair this policy with monitors for CPU, memory, and storage usage to catch underlying issues.