Root Cause Analysis – VSA10 Intermittent Service Degradation Following 10.28 Release
Summary
Between 2026-07-24 02:00 UTC and 2026-07-24 20:25 UTC, some customers using VSA10 experienced intermittent service degradation, including temporary agent connectivity issues, delayed agent check-ins, and HTTP 503 errors. During periods of elevated automation activity, affected agents could appear to repeatedly transition between online and offline states before reconnecting automatically. The issue was fully resolved on July 24, 2026.
Root Cause
A software change introduced in the 10.28 release altered how certain automation-related database operations were processed. Under high levels of concurrent automation activity, this change resulted in resource contention within the application environment, leading to intermittent service interruptions and delayed responses for agent communications. A hotfix was developed, tested, and deployed the same day to restore the previous transaction handling behavior and eliminate the contention condition.
Incident Timeline
Identified: 2026-07-24 02:00 UTC
Public Notification: 2026-07-24 13:27 UTC
Resolved: 2026-07-24 20:25 UTC
Preventative Measures
To reduce the likelihood and impact of similar incidents in the future, we are implementing the following improvements:
Enhancements to Monitoring and Alerting
Introduce additional monitoring and alerting for database contention and long-running transactions to further accelerate detection of similar conditions.
Expand service health dashboards to further improve visibility into application response performance.
Enhancements to Release Management Practices
Implement expanded canary deployment procedures for higher-risk platform changes, allowing updates to be validated on a limited subset of infrastructure before broader rollout, further reducing the risk for these deployments.
Add additional review requirements for changes affecting high-volume agent communication paths, further reducing risk for these deployments.
Enhancements to Validation and Testing
Strengthen load and concurrency testing for components that support large-scale automation activity, further increasing scalability at the time of deployment.
Increase validation coverage for scenarios that simulate production-scale agent workloads to catch scalability issues prior to deployment, further reducing risk.
Enhancements to Platform Resiliency
Continue development of automation workload distribution improvements designed to reduce concentrated processing spikes during large-scale task execution, further increasing scalability at the time of deployment.
Review additional service isolation and capacity improvements to further reduce the impact of localized workload surges.
Completed Corrective Action