Root Cause Analysis – Datto RMM Merlot Intermittent User Interface Degradation
Summary
Between 2026-07-21 07:55 UTC and 2026-07-21 17:00 UTC, some customers on the Merlot platform experienced intermittent performance degradation within Datto RMM. During this period, users may have encountered slow-loading or failing pages, missing Global Search results, delayed or incomplete device summary information, and intermittent issues accessing certain functionality within the web interface. Service stability was restored following a series of remediation activities targeting the affected messaging and application services.
Root Cause
Our investigation determined that the incident was caused by instability within a core messaging service dependency path used by several Datto RMM components. This instability impacted communication between application services, resulting in elevated backend resource utilization and intermittent failures of customer-facing workflows. The issue was further compounded by the persistence of unhealthy service connections from the recent incident the day prior, which contributed to message processing delays and reduced application responsiveness. Recovery required restarting the affected messaging infrastructure and dependent services, along with infrastructure adjustments to restore normal operation.
Incident Timeline
Preventative Measures
To reduce the likelihood and impact of similar incidents in the future, we are taking the following actions:
Enhancing infrastructure resiliency by improving service recovery mechanisms so dependent services can recover automatically when messaging infrastructure is restored, reducing the risk of cascading failures across connected services.
Enhancing monitoring and alerting by implementing additional monitoring for messaging service health, queue backlog growth, and end-to-end workflow execution so potential issues can be identified and resolved before customer impact occurs.
Enhancing customer experience monitoring by introducing workflow-based monitoring and dashboards focused on key customer journeys, allowing earlier identification of degraded user experiences within the platform.
Enhancing incident management and response by developing and documenting recovery procedures for messaging-path degradation scenarios, including validation steps to ensure all related services are operating normally before an incident is considered resolved.
Enhancing post-incident validation by adding additional dependency-aware validation checks following major recovery activities to confirm the health of interconnected services and reduce the likelihood of follow-on incidents.