Degradation on accessing Modules Pages and Modules search
Opened
Trainers and Learners are unable to access the modules OR search for the modules in search bar as the page is keepon loading OR ended up in an error page. We are actively investigating the reason for the same.
The issue is fixed and the module pages are loading as expected, we are actively analyzing the root cause and also monitoring it.
Trainers and Learners are unable to access the modules OR search for the modules in search bar as the page is keepon loading OR ended up in an error page. We are actively investigating the reason for the same.
Hosts all microservices, learner landing page widgets, and live provider modules with managed application runtime and scaling.
Virutal Machines10 components
100.00% uptime
Hosts compute workloads requiring full OS control, including legacy services or specialized processing tasks.
Container Apps183 components
99.98% uptime
Executes containerized workloads for microservices, learner widgets, and live provider modules with auto-scaling and simplified container management.
SQL Read Replica187 components
100.00% uptime
API Management9 components
100.00% uptime
Application Gateway12 components
100.00% uptime
Acts as the entry point for all incoming requests, handling routing, security, and load balancing before forwarding traffic to monolith or microservices.
CDN27 components
100.00% uptime
Delivers static content such as images, CSS, JavaScript, and documents globally with low latency using content caching.
Cloud Service7 components
100.00% uptime
Runs core platform functionalities including user authentication, trainer workflows, and administrative operations.
Event Grid17 components
100.00% uptime
Event-driven messaging service used to route events to background and downstream services in near real time.
Front Door8 components
100.00% uptime
Provides global HTTP/HTTPS load balancing, traffic acceleration, and high availability for frontend applications.
Kubernetes Service10 components
99.98% uptime
Orchestrates all microservices, learner landing page widgets, and live provider modules with advanced container management and scalability.
Mobile9 components
100.00% uptime
Service Bus15 components
100.00% uptime
Enables reliable asynchronous messaging for reporting, certificate generation, leaderboards, email notifications, and other background processing tasks.
SQL Databases920 components
100.00% uptime
Stores structured, transactional application data such as user profiles, course data, enrollments, progress tracking, and system configurations.
Storage Accounts149 components
100.00% uptime
Provides scalable blob storage for documents, images, media files, and other application content.
VM Scale Sets24 components
100.00% uptime
Supports scalable compute for core functionalities like login, trainer workflows, and admin modules with automatic instance scaling.
App Services29 components
99.83% uptime
Hosts all microservices, learner landing page widgets, and live provider modules with managed application runtime and scaling.
Container Apps47 components
99.98% uptime
Executes containerized workloads for microservices, learner widgets, and live provider modules with auto-scaling and simplified container management.
Application Gateway1 components
100.00% uptime
Acts as the entry point for all incoming requests, handling routing, security, and load balancing before forwarding traffic to monolith or microservices.
CDN1 components
100.00% uptime
Delivers static content such as images, CSS, JavaScript, and documents globally with low latency using content caching.
Kubernetes Service1 components
99.98% uptime
Orchestrates all microservices, learner landing page widgets, and live provider modules with advanced container management and scalability.
Mobile1 components
100.00% uptime
Service Bus1 components
100.00% uptime
Enables reliable asynchronous messaging for reporting, certificate generation, leaderboards, email notifications, and other background processing tasks.
SQL Databases21 components
100.00% uptime
Stores structured, transactional application data such as user profiles, course data, enrollments, progress tracking, and system configurations.
Storage Accounts5 components
100.00% uptime
Provides scalable blob storage for documents, images, media files, and other application content.
VM Scale Sets3 components
100.00% uptime
Supports scalable compute for core functionalities like login, trainer workflows, and admin modules with automatic instance scaling.
App Services33 components
99.83% uptime
Hosts all microservices, learner landing page widgets, and live provider modules with managed application runtime and scaling.
API Management1 components
100.00% uptime
Application Gateway1 components
100.00% uptime
Acts as the entry point for all incoming requests, handling routing, security, and load balancing before forwarding traffic to monolith or microservices.
Cloud Service1 components
100.00% uptime
Runs core platform functionalities including user authentication, trainer workflows, and administrative operations.
Container Apps18 components
99.98% uptime
Executes containerized workloads for microservices, learner widgets, and live provider modules with auto-scaling and simplified container management.
Event Grid1 components
100.00% uptime
Event-driven messaging service used to route events to background and downstream services in near real time.
Kubernetes Service1 components
99.98% uptime
Orchestrates all microservices, learner landing page widgets, and live provider modules with advanced container management and scalability.
Mobile1 components
100.00% uptime
Service Bus1 components
100.00% uptime
Enables reliable asynchronous messaging for reporting, certificate generation, leaderboards, email notifications, and other background processing tasks.
SQL Databases67 components
100.00% uptime
Stores structured, transactional application data such as user profiles, course data, enrollments, progress tracking, and system configurations.
SQL Read Replica90 components
100.00% uptime
Storage Accounts16 components
100.00% uptime
Provides scalable blob storage for documents, images, media files, and other application content.
VM Scale Sets3 components
100.00% uptime
Supports scalable compute for core functionalities like login, trainer workflows, and admin modules with automatic instance scaling.
CDN14 components
100.00% uptime
Delivers static content such as images, CSS, JavaScript, and documents globally with low latency using content caching.
Front Door8 components
100.00% uptime
Provides global HTTP/HTTPS load balancing, traffic acceleration, and high availability for frontend applications.
Mobile1 components
100.00% uptime
App Services209 components
99.83% uptime
Hosts all microservices, learner landing page widgets, and live provider modules with managed application runtime and scaling.
SQL Read Replica94 components
100.00% uptime
Container Apps39 components
99.98% uptime
Executes containerized workloads for microservices, learner widgets, and live provider modules with auto-scaling and simplified container management.
API Management4 components
100.00% uptime
Application Gateway5 components
100.00% uptime
Acts as the entry point for all incoming requests, handling routing, security, and load balancing before forwarding traffic to monolith or microservices.
CDN4 components
100.00% uptime
Delivers static content such as images, CSS, JavaScript, and documents globally with low latency using content caching.
Cloud Service1 components
100.00% uptime
Runs core platform functionalities including user authentication, trainer workflows, and administrative operations.
Event Grid10 components
100.00% uptime
Event-driven messaging service used to route events to background and downstream services in near real time.
Kubernetes Service3 components
99.98% uptime
Orchestrates all microservices, learner landing page widgets, and live provider modules with advanced container management and scalability.
Mobile1 components
100.00% uptime
Service Bus7 components
100.00% uptime
Enables reliable asynchronous messaging for reporting, certificate generation, leaderboards, email notifications, and other background processing tasks.
SQL Databases502 components
100.00% uptime
Stores structured, transactional application data such as user profiles, course data, enrollments, progress tracking, and system configurations.
Storage Accounts72 components
100.00% uptime
Provides scalable blob storage for documents, images, media files, and other application content.
Virutal Machines7 components
100.00% uptime
Hosts compute workloads requiring full OS control, including legacy services or specialized processing tasks.
VM Scale Sets5 components
100.00% uptime
Supports scalable compute for core functionalities like login, trainer workflows, and admin modules with automatic instance scaling.
App Services45 components
99.83% uptime
Hosts all microservices, learner landing page widgets, and live provider modules with managed application runtime and scaling.
Virutal Machines1 components
100.00% uptime
Hosts compute workloads requiring full OS control, including legacy services or specialized processing tasks.
Application Gateway1 components
100.00% uptime
Acts as the entry point for all incoming requests, handling routing, security, and load balancing before forwarding traffic to monolith or microservices.
Cloud Service1 components
100.00% uptime
Runs core platform functionalities including user authentication, trainer workflows, and administrative operations.
Container Apps2 components
99.98% uptime
Executes containerized workloads for microservices, learner widgets, and live provider modules with auto-scaling and simplified container management.
Event Grid2 components
100.00% uptime
Event-driven messaging service used to route events to background and downstream services in near real time.
Kubernetes Service1 components
99.98% uptime
Orchestrates all microservices, learner landing page widgets, and live provider modules with advanced container management and scalability.
Mobile1 components
100.00% uptime
Service Bus1 components
100.00% uptime
Enables reliable asynchronous messaging for reporting, certificate generation, leaderboards, email notifications, and other background processing tasks.
SQL Databases25 components
100.00% uptime
Stores structured, transactional application data such as user profiles, course data, enrollments, progress tracking, and system configurations.
Storage Accounts4 components
100.00% uptime
Provides scalable blob storage for documents, images, media files, and other application content.
VM Scale Sets3 components
100.00% uptime
Supports scalable compute for core functionalities like login, trainer workflows, and admin modules with automatic instance scaling.
App Services25 components
99.83% uptime
Hosts all microservices, learner landing page widgets, and live provider modules with managed application runtime and scaling.
Container Apps16 components
99.98% uptime
Executes containerized workloads for microservices, learner widgets, and live provider modules with auto-scaling and simplified container management.
API Management1 components
100.00% uptime
Application Gateway1 components
100.00% uptime
Acts as the entry point for all incoming requests, handling routing, security, and load balancing before forwarding traffic to monolith or microservices.
CDN2 components
100.00% uptime
Delivers static content such as images, CSS, JavaScript, and documents globally with low latency using content caching.
Event Grid1 components
100.00% uptime
Event-driven messaging service used to route events to background and downstream services in near real time.
Kubernetes Service1 components
99.98% uptime
Orchestrates all microservices, learner landing page widgets, and live provider modules with advanced container management and scalability.
Mobile1 components
100.00% uptime
Service Bus1 components
100.00% uptime
Enables reliable asynchronous messaging for reporting, certificate generation, leaderboards, email notifications, and other background processing tasks.
SQL Databases5 components
100.00% uptime
Stores structured, transactional application data such as user profiles, course data, enrollments, progress tracking, and system configurations.
SQL Read Replica2 components
100.00% uptime
Storage Accounts3 components
100.00% uptime
Provides scalable blob storage for documents, images, media files, and other application content.
App Services35 components
99.83% uptime
Hosts all microservices, learner landing page widgets, and live provider modules with managed application runtime and scaling.
API Management1 components
100.00% uptime
Application Gateway1 components
100.00% uptime
Acts as the entry point for all incoming requests, handling routing, security, and load balancing before forwarding traffic to monolith or microservices.
CDN2 components
100.00% uptime
Delivers static content such as images, CSS, JavaScript, and documents globally with low latency using content caching.
Cloud Service2 components
100.00% uptime
Runs core platform functionalities including user authentication, trainer workflows, and administrative operations.
Container Apps18 components
99.98% uptime
Executes containerized workloads for microservices, learner widgets, and live provider modules with auto-scaling and simplified container management.
Event Grid1 components
100.00% uptime
Event-driven messaging service used to route events to background and downstream services in near real time.
Kubernetes Service1 components
99.98% uptime
Orchestrates all microservices, learner landing page widgets, and live provider modules with advanced container management and scalability.
Mobile1 components
100.00% uptime
Service Bus1 components
100.00% uptime
Enables reliable asynchronous messaging for reporting, certificate generation, leaderboards, email notifications, and other background processing tasks.
SQL Databases226 components
100.00% uptime
Stores structured, transactional application data such as user profiles, course data, enrollments, progress tracking, and system configurations.
SQL Read Replica1 components
100.00% uptime
Storage Accounts21 components
100.00% uptime
Provides scalable blob storage for documents, images, media files, and other application content.
Virutal Machines1 components
100.00% uptime
Hosts compute workloads requiring full OS control, including legacy services or specialized processing tasks.
VM Scale Sets4 components
100.00% uptime
Supports scalable compute for core functionalities like login, trainer workflows, and admin modules with automatic instance scaling.
API Management2 components
100.00% uptime
App Services33 components
99.83% uptime
Hosts all microservices, learner landing page widgets, and live provider modules with managed application runtime and scaling.
Application Gateway1 components
100.00% uptime
Acts as the entry point for all incoming requests, handling routing, security, and load balancing before forwarding traffic to monolith or microservices.
Cloud Service1 components
100.00% uptime
Runs core platform functionalities including user authentication, trainer workflows, and administrative operations.
Container Apps25 components
99.98% uptime
Executes containerized workloads for microservices, learner widgets, and live provider modules with auto-scaling and simplified container management.
Event Grid1 components
100.00% uptime
Event-driven messaging service used to route events to background and downstream services in near real time.
Kubernetes Service1 components
99.98% uptime
Orchestrates all microservices, learner landing page widgets, and live provider modules with advanced container management and scalability.
Mobile1 components
100.00% uptime
Service Bus1 components
100.00% uptime
Enables reliable asynchronous messaging for reporting, certificate generation, leaderboards, email notifications, and other background processing tasks.
SQL Databases28 components
100.00% uptime
Stores structured, transactional application data such as user profiles, course data, enrollments, progress tracking, and system configurations.
Storage Accounts8 components
100.00% uptime
Provides scalable blob storage for documents, images, media files, and other application content.
Virutal Machines1 components
100.00% uptime
Hosts compute workloads requiring full OS control, including legacy services or specialized processing tasks.
VM Scale Sets3 components
100.00% uptime
Supports scalable compute for core functionalities like login, trainer workflows, and admin modules with automatic instance scaling.
App Services31 components
99.83% uptime
Hosts all microservices, learner landing page widgets, and live provider modules with managed application runtime and scaling.
Application Gateway1 components
100.00% uptime
Acts as the entry point for all incoming requests, handling routing, security, and load balancing before forwarding traffic to monolith or microservices.
CDN4 components
100.00% uptime
Delivers static content such as images, CSS, JavaScript, and documents globally with low latency using content caching.
Cloud Service1 components
100.00% uptime
Runs core platform functionalities including user authentication, trainer workflows, and administrative operations.
Container Apps18 components
99.98% uptime
Executes containerized workloads for microservices, learner widgets, and live provider modules with auto-scaling and simplified container management.
Event Grid1 components
100.00% uptime
Event-driven messaging service used to route events to background and downstream services in near real time.
Kubernetes Service1 components
99.98% uptime
Orchestrates all microservices, learner landing page widgets, and live provider modules with advanced container management and scalability.
Mobile1 components
100.00% uptime
Service Bus2 components
100.00% uptime
Enables reliable asynchronous messaging for reporting, certificate generation, leaderboards, email notifications, and other background processing tasks.
SQL Databases46 components
100.00% uptime
Stores structured, transactional application data such as user profiles, course data, enrollments, progress tracking, and system configurations.
Storage Accounts20 components
100.00% uptime
Provides scalable blob storage for documents, images, media files, and other application content.
VM Scale Sets3 components
100.00% uptime
Supports scalable compute for core functionalities like login, trainer workflows, and admin modules with automatic instance scaling.
Resolved
Unable to access Modules Pages
Incident window
to
Trainers and Learners are unable to access the modules as the page is keepon loading and ended up in an error page. We are actively investigating the reason for the same.
Started
Resolved
Affected regions
CUAT, Dubai, Global, India, Indonesia, QA, Saudi, SEA, Staging, UK, US
The issue is fixed and the module pages are loading as expected, we are actively analyzing the root cause and also monitoring it
Trainers and Learners are unable to access the modules as the page is keepon loading and ended up in an error page. We are actively investigating the reason for the same.
Post-incident report
Root cause
A certificate used by shared internal services passed its validity period, causing secure connections to fail. During replacement, the certificate chain initially could not be validated by some application services. This prevented the Modules page from retrieving the information it needed until the chain was corrected.
Mitigation steps
The certificate chain was corrected, and the affected service path was successfully validated after recovery. Our certificate monitoring, pre-cutover checks, staged rollout and rollback, customer-journey monitoring, and recovery controls were already in place and operating as expected. We will continue routine service and certificate monitoring. No action is required from customers.
Resolved
Performance Degradation in SEA Region
Incident window
to
The Databases in SEA region are saturated which caused the elastic pool to reach its max limit. The users will see the performance issue. The Scaling is currently in progress.
The elastic pool is scaled and the performance is back to normal
The Databases in SEA region are saturated which caused the elastic pool to reach its max limit. The users will see the performance issue. The Scaling is currently in progress.
Post-incident report
Root cause
A simultaneous increase in workload across multiple databases caused the combined demand to temporarily exceed the available shared capacity. Individual databases remained within their expected operating ranges.
Mitigation steps
We are enhancing our proactive capacity management and autoscaling controls to better account for simultaneous increases in overall demand. This will enable earlier scaling and reduce the likelihood of similar performance degradation recurring.
The root cause was an Azure Container Registry service disruption in Southeast Asia affecting image manifest retrieval. During this period, appsvc-servicediscovery-sea attempted to start/restart using a container image hosted in the affected registry region. Because ACR manifest retrieval was failing or unstable, App Service could not reliably pull or validate the container image. As a result, the container did not become healthy, and Azure App Service front-end returned 503 Site Unavailable. The issue first impacted servicediscovery-sea. Other services were not necessarily down because of their own container issue, but they were affected because they depend on Service Discovery for service lookup/routing. Once Service Discovery was unavailable, dependent services experienced high response time, failed calls, or degraded behavior.
Mitigation steps
Immediate recovery: - Pointed appsvc-servicediscovery-sea to a known working image from an unaffected registry.
Resolved
Reports downloading is degraded in India region
Incident window
to
We are seeing slowness in downloading reports in india region. Reports will be in not started or in progress state in report hub for longer than expected. We are investigating the same
The issue is addressed and the reports are now processing as expected. We are actively monitoring and investigating the cause for the issue.
We are seeing slowness in downloading reports in india region. Reports will be in not started or in progress state in report hub for longer than expected. We are investigating the same
Post-incident report
Root cause
The new application build caused large analytics and report requests to create more concurrent database processing than expected. Several operations then required substantially more compute and processing time than normal, causing worker contention and queueing in the report processing path. The underlying platform and storage remained healthy.
Mitigation steps
- Existing monitoring alerts identified elevated report completion times and processing pressure. - The investigation correlated the start of the degradation with the newly introduced application build. - The affected build was reverted immediately, and queued report work was cleared in a controlled manner. - New report requests were validated after recovery to confirm that completion times had returned to the normal range.
The new application build was reverted immediately after it was identified as the source of the degradation. Active work was then allowed to drain in a controlled manner, which reduced contention and restored normal report generation. Processing health, completion time, and queue behavior were monitored throughout recovery.
Resolved
we are seeing some degradation in INDIA region
Incident window
to
we are seeing some degradation in INDIA region and actively working on it.
Started
Resolved
Affected regions
India
Affected servicesvmss-monolith-india-api (India)
Customer updates
Services back to normal , investigating further to identify the issue.
The issue is mitigated. We are actively monitoring the instances.
we are seeing some degradation in INDIA region and actively working on it.
Post-incident report
Root cause
One India monolith API VMSS instance cpu went down. Because of this, requests queued and shifted to the remaining healthy instances, increasing load on them. The remaining instances became slow/overloaded, causing request cancellations, healthcheck failures, and HTTP 500 responses. Redis timeout logs were observed, but Redis metrics were healthy, so Redis was a secondary symptom and not the primary cause. The issue was resolved after reimaging the unhealthy instance.
Mitigation steps
Monitor VMSS instance health and load balancer health probe status more closely. Add alerting for unhealthy backend count and sudden traffic redistribution across instances. Review auto-healing/reimage automation for unhealthy monolith instances. Validate that failed/unhealthy instances are removed from rotation quickly. Continue monitoring Redis timeout logs, but treat them as application-side backlog symptoms unless Redis platform metrics show saturation.
Resolved
we are seeing some degradation in INDIA region and actively working on it.
Incident window
to
we are seeing some degradation in INDIA region and actively working on it.
Started
Resolved
Affected regions
India
Affected servicesvmss-monolith-india-api (India)
Customer updates
the failure was caused by Azure SQL worker/request exhaustion on the SQL elastic pool in INDIA region. The issue is resolved and actively monitoring
we are seeing some degradation in INDIA region and actively working on it.
Post-incident report
Root cause
The immediate cause of the incident was a failure in the internal autoscaling automation responsible for scaling regional Azure SQL elastic pools. The automation service became unhealthy because one of its background maintenance workers encountered memory exhaustion while processing a large recovery workload. This caused repeated application process exits and health probe failures in the automation service runtime. Because the autoscaling service was not healthy during the India incident window, scale-up actions that would normally add database pool compute capacity were not completed in time. The affected pool therefore operated with insufficient headroom while production traffic continued, causing elevated compute and worker utilization and resulting in slower responses for some users. The issue was not caused by a customer-side change, a specific tenant database, Azure SQL storage saturation, transaction log saturation, or an Azure platform outage. The trigger was an internal automation service reliability issue combined with insufficient regional pool headroom while the automation service was unavailable.
Mitigation steps
Optimize Autoscaler Background Recovery Processing: Avoid memory exhaustion by paging large recovery queries and preventing full materialization of large operation and log datasets. Status: In Progress. Separate Autoscaling Control Path from Maintenance Workers: Ensure unrelated background maintenance workloads cannot impact the autoscaling function. Status: Planned. Add High-Severity Alerts for Autoscaler Crash Loops and Failed Health Probes: Detect autoscaler unavailability before scale actions are missed. Status: In Progress. Add Alerts for Skipped, Delayed, or Stuck Database Pool Scale Executions: Identify when expected scale actions are not completed within the required window. Status: Planned. Increase Autoscaler Runtime Headroom and Review Replica Strategy: Reduce the risk of CPU and memory saturation in the automation service and improve service resilience. Status: Planned. Define Regional Manual Scale Runbook and Escalation Path: Ensure rapid manual mitigation when automation is unavailable. Status: Completed. Review Regional Pool Headroom Thresholds: Reduce the risk of recurrence during peak traffic while autoscaling is unavailable or delayed. Status: In Progress. Post-Incident Monitoring Window: Continue monitoring the India pool after mitigation to confirm sustained stability. Status: In Progress.
Resolved
we are seeing some degradation in SEA region and actively working on it.
Incident window
to
we are seeing some degradation in SEA region and actively working on it.
Started
Resolved
Affected regions
SEA
Affected servicesvmss-modularmonolith-sea (SEA)
Customer updates
the failure was caused by Azure SQL worker/request exhaustion on the SQL elastic pool in SEA region. The issue is resolved and actively monitoring
we are seeing some degradation in SEA region and actively working on it.
Post-incident report
Root cause
The immediate cause of the incident was a failure in the internal autoscaling automation responsible for scaling regional Azure SQL elastic pools. The automation service became unhealthy because one of its background maintenance workers encountered memory exhaustion while processing a large recovery workload. This caused repeated application process exits and health probe failures in the automation service runtime. Because the autoscaling service was not healthy during the India incident window, scale-up actions that would normally add database pool compute capacity were not completed in time. The affected pool therefore operated with insufficient headroom while production traffic continued, causing elevated compute and worker utilization and resulting in slower responses for some users. The issue was not caused by a customer-side change, a specific tenant database, Azure SQL storage saturation, transaction log saturation, or an Azure platform outage. The trigger was an internal automation service reliability issue combined with insufficient regional pool headroom while the automation service was unavailable.
Mitigation steps
Optimize Autoscaler Background Recovery Processing Action: Avoid memory exhaustion by paging large recovery queries and preventing full materialization of large operation and log datasets. Status: In Progress Separate Autoscaling Control Path from Maintenance Workers Action: Ensure unrelated background maintenance workloads cannot impact the autoscaling function. Status: Planned Add High-Severity Alerts for Autoscaler Crash Loops and Failed Health Probes Action: Detect autoscaler unavailability before scale actions are missed. Status: In Progress Add Alerts for Skipped, Delayed, or Stuck Database Pool Scale Executions Action: Identify when expected scale actions are not completed within the required window. Status: Planned Increase Autoscaler Runtime Headroom and Review Replica Strategy Action: Reduce the risk of CPU and memory saturation in the automation service and improve service resilience. Status: Planned Define Regional Manual Scale Runbook and Escalation Path Action: Ensure rapid manual mitigation when automation is unavailable. Status: Completed Review Regional Pool Headroom Thresholds Action: Reduce the risk of recurrence during peak traffic while autoscaling is unavailable or delayed. Status: In Progress Post-Incident Monitoring Window Action: Continue monitoring the SEA regional pool after mitigation to confirm sustained stability. Status: In Progress
Resolved
Currently observing performance degradation in the india region.
Incident window
to
we are seeing some degradation in India region and actively working on it.
The issue is mitigated. We are actively monitoring the instances.
we are seeing some degradation in India region and actively working on it.
Post-incident report
Root cause
The root cause was an Azure platform-side availability degradation on the public Azure Load Balancer alb-monolith-india-api, which fronts the backend service path used by the Application Gateway apipool configuration. Azure Resource Health classified the event as Unavailable / Downtime, PlatformInitiated, Unplanned, and Transient. As the Load Balancer data path degraded, Application Gateway health and routing for the apipool upstream became unavailable. The gateway therefore returned 502 responses with ERRORINFO_UPSTREAM_NO_LIVE instead of forwarding those requests to the backend service instances. Backend VM availability metrics did not show a regional compute outage as the initiating condition. VMSS autoscale activity occurred after the Load Balancer health event had already started, and automatic VM repair activity occurred after service recovery was already underway. This timing indicates that autoscale and repair activity were response or recovery signals, not the initial root cause.
Mitigation steps
The India region availability incident is resolved. Monitoring remains active for Application Gateway failed requests, backend health, Azure Load Balancer availability, VMSS availability, and regional synthetic probes.
Resolved
Currently observing performance degradation in Indonesia region. Investigation is in progress.
Incident window
to
Currently observing performance degradation in the Indonesia region container services.
Starting at 05:24 UTC on 21 May 2026, Azure experienced a service disruption affecting the Azure Container Registry GetToken API in Southeast Asia region. This caused the authentication failures or delays when obtaining access tokens for registry operations and affected our services.
We had temporarily changed the registry region and the services are up and working now.
Services back to normal we are actively monitoring it
Currently observing performance degradation in the Indonesia region container services.
Post-incident report
Root cause
The issue was caused by an Azure-side underlying storage dependency failure, which degraded backend operations and resulted in increased server errors for Azure Container Registry requests.
Mitigation steps
Azure has mitigated the issue, and we will monitor the service.
Resolved
Service Degradation Impacting Some Microservices Across Regions
Incident window
to
Some microservices may experience slowness, intermittent failures, or degraded performance across affected regions.
Started
Resolved
Affected regions
Global
Affected servicesvmss-monolith-india-api (India)
Customer updates
The SSL certificate pi configuration is updated to all the services, and access is restored. The app is submitted for review and waiting for the app get approved and we will release the same once it is approved
Web access has been restored and is working normally across the affected services. We are continuing to work with the mobile team to resolve SSL pinning failures impacting some white-labeled mobile apps. Further updates will be shared as soon as the mobile fix is validated.
We have identified an additional impact affecting some mobile applications that use SSL certificate pinning.
The affected endpoints are now serving the updated valid TLS certificate. However, some white-labeled mobile applications appear to have pinned the previous certificate or public key, causing SSL pinning validation failures after the certificate reissue.
Browser access and non-pinned clients are recovering after the certificate binding refresh. Mobile applications with certificate pinning may continue to experience connectivity failures until their pin configuration is updated to trust the new certificate/public key.
Our teams are working on remediation options for impacted white-labeled mobile apps and will provide further updates.
Some microservices may experience slowness, intermittent failures, or degraded performance across affected regions.
Post-incident report
Root cause
The incident was caused by a mismatch between the newly issued TLS certificate served by the affected endpoints and the SSL certificate/public key pinned in some white-labelled mobile applications.
Although the updated valid TLS certificate was deployed successfully, certain mobile apps continued to trust the previously pinned certificate or public key. This resulted in SSL pinning validation failures for those apps, causing mobile connectivity issues until the pinning configuration was updated.
Mitigation steps
All remediation steps have been completed. The SSL certificate pinning configuration has been updated across the required services, and access has been restored.
The remaining actions are:
App review and release The updated mobile app build has been submitted for review. Once approved, the app will be released. Post-release monitoring Continue monitoring application availability, SSL handshake errors, mobile login/access issues, and customer-reported incidents after the app release. Client communication Inform impacted clients once the updated build is approved and released, and request them to update/upload the latest white-labelled app version where applicable. Certificate rotation checklist Maintain a checklist of all services and mobile applications using SSL certificate pinning to ensure they are validated during future certificate updates. Preventive improvement Review the SSL pinning approach and consider using public key pinning or backup pins where supported, to reduce the risk of similar issues during future certificate renewals or reissues.
Resolved
Currently we are facing Performance Degradation in India region, and we are investigating the same
Incident window
to
Currently observing performance degradation in the India region. Investigation is in progress.
Started
Resolved
Affected regions
India
Affected servicesvmss-monolith-india-api (India)
Customer updates
Currently observing performance degradation in the India region. Investigation is in progress.
Resolved
Performance Degradation in India Region
Incident window
to
Users in the India region experienced temporary performance slowness in multiple pages like module listing page, app login.
the performance is back to normal, We are analysing the reason for this degradation. Also we are actively monitoring the service
Users in the India region experienced temporary performance slowness in multiple pages like module listing page, app login.
Post-incident report
Root cause
The degradation was caused by an unexpected service instability triggered by changes in the runtime environment, which impacted the availability of a critical component in the production region.
Mitigation steps
1. Introduce additional safeguards and validation checks before runtime changes 2. Improve failover and recovery mechanisms to minimize impact duration 3. Schedule critical changes with enhanced validation and rollback strategies
Resolved
SSO service is degraded. Users unable to login via sso
Incident window
to
SSO service is degraded. Users unable to login via sso. We are investigating the same
Started
Resolved
Affected regions
India
Affected servicesappsvc-disprzsso (India)
Customer updates
The impacted component was removed from the production environment, following which the service was restored. The service is now up and running, and the system is being actively monitored to ensure stability.
A deviation from the established change management process was identified in the production environment, resulting in service downtime. Our team is working on mitigating the same.
SSO service is degraded. Users unable to login via sso. We are investigating the same
Post-incident report
Root cause
As part of ongoing efforts to deprecate the legacy SSO integration, a production change related to the older SSO configuration led to service instability for clients still using the legacy authentication mechanism. This resulted in temporary access issues for those clients.
Mitigation steps
• Affected clients will be advised to migrate from the legacy SSO to the new SSO integration to ensure continued stability and support.
• The legacy SSO will be progressively deprecated in a controlled manner to avoid future disruptions.
• Additional validation and monitoring will be implemented during the deprecation process to minimize client impact.
Resolved
Performance Degradation in India Region
Incident window
to
Users in the India region experienced temporary performance slowness in multiple pages like module listing page, app login.