Snapshot 30770
Normalized text
Scripts and page chrome removed; this is what change detection compares.
{
"pages": [
{
"end_time": "2026-10-31T23:59:59Z",
"months": [
{
"incidents": [],
"name": "October",
"year": 2026
},
{
"incidents": [
{
"code": "f7h09n7kt0lg",
"impact": "minor",
"message": "Between approximately 10:30 AM IST and 1:10 PM IST today, some placeOrder calls made via the Plum Pro API failed with an internal validation error. All other Plum flows, storefronts, and clients were unaffected. The issue has been identified and resolved as of 1:10 PM IST.",
"name": "Partial degradation - Plum Pro API placeOrder",
"timestamp": "Sep \u003cvar data-var='date'\u003e18\u003c/var\u003e, \u003cvar data-var='time'\u003e07:50\u003c/var\u003e - \u003cvar data-var='time'\u003e07:50\u003c/var\u003e UTC"
},
{
"code": "p1gr8zj064lz",
"impact": "critical",
"message": "multiple vendor facing issue with order processing due to wrong vendor config got updated config and resolved.",
"name": "Order Processing Failure for multiple vendors",
"timestamp": "Sep \u003cvar data-var='date'\u003e4\u003c/var\u003e, \u003cvar data-var='time'\u003e08:30\u003c/var\u003e - \u003cvar data-var='time'\u003e08:30\u003c/var\u003e UTC"
},
{
"code": "65g22s2ntymy",
"impact": "minor",
"message": "This incident has been resolved.",
"name": "Delay in link processing",
"timestamp": "Sep \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e12:45\u003c/var\u003e - \u003cvar data-var='time'\u003e14:45\u003c/var\u003e UTC"
},
{
"code": "plqzhc3s0z94",
"impact": "major",
"message": "The intermittent order processing issue has been resolved. The root cause was an accumulation of MySQL connections that had not been released, which was affecting order processing reliability. We have released the affected connections and confirmed order processing is operating normally.",
"name": "Order Processing Intermitently not working",
"timestamp": "Sep \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e06:45\u003c/var\u003e - \u003cvar data-var='time'\u003e06:45\u003c/var\u003e UTC"
}
],
"name": "September",
"year": 2026
},
{
"incidents": [
{
"code": "44tv45f3mw6v",
"impact": "maintenance",
"message": "The scheduled maintenance has been completed. All services restored after successful upgrade.",
"name": "[Scheduled] Scheduled Maintenance",
"timestamp": "Aug \u003cvar data-var='date'\u003e30\u003c/var\u003e, \u003cvar data-var='time'\u003e00:55\u003c/var\u003e - \u003cvar data-var='time'\u003e05:12\u003c/var\u003e UTC"
},
{
"code": "57khh6kxmk2z",
"impact": "maintenance",
"message": "The scheduled maintenance has been completed.",
"name": "[Scheduled] Redis upgrade on empuls UAE",
"timestamp": "Aug \u003cvar data-var='date'\u003e29\u003c/var\u003e, \u003cvar data-var='time'\u003e00:45\u003c/var\u003e - \u003cvar data-var='time'\u003e01:00\u003c/var\u003e UTC"
},
{
"code": "nf81rtnnsg5n",
"impact": "minor",
"message": "Between August 22 and August 24, 2026, some users experienced failures when attempting to redeem reward points via UPI payout on the Plum storefront, receiving an error during the redemption process. Other redemption options (vouchers, gift cards, points-based redemptions) were not affected.\n\nThis was caused by an internal infrastructure/deployment issue following a scheduled maintenance window, which impacted the service responsible for validating UPI payout requests.\n\nThe issue has been identified and fully resolved. UPI payout redemptions are now processing normally. We apologize for the inconvenience this may have caused.",
"name": "Partial Degradation: UPI Payout Redemption",
"timestamp": "Aug \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e10:45\u003c/var\u003e - \u003cvar data-var='time'\u003e10:45\u003c/var\u003e UTC"
},
{
"code": "xj5p403thjys",
"impact": "maintenance",
"message": "The scheduled maintenance has been completed.",
"name": "[Scheduled] Scheduled Maintenance: Redis Server Upgrade",
"timestamp": "Aug \u003cvar data-var='date'\u003e22\u003c/var\u003e, \u003cvar data-var='time'\u003e00:30\u003c/var\u003e - \u003cvar data-var='time'\u003e00:45\u003c/var\u003e UTC"
},
{
"code": "zckxtn6ffy8w",
"impact": "none",
"message": "A set of redemption requests failed due to a database connectivity issue affecting the redemption service. Users attempting a redemption during this window may have received an error.The issue was identified and remediated, and all redemption traffic has been operating normally",
"name": "Resolved — Partial degradation of redemption service",
"timestamp": "Aug \u003cvar data-var='date'\u003e16\u003c/var\u003e, \u003cvar data-var='time'\u003e06:00\u003c/var\u003e - \u003cvar data-var='time'\u003e06:00\u003c/var\u003e UTC"
},
{
"code": "7ncm54vqh7nz",
"impact": "major",
"message": "We recently identified a performance issue affecting Plum Pro APIs, which resulted in intermittent errors and slower response times for some requests.",
"name": "Plum Pro Apis are timing out",
"timestamp": "Aug \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e05:00\u003c/var\u003e - \u003cvar data-var='time'\u003e05:00\u003c/var\u003e UTC"
},
{
"code": "0t43bp0hf0ng",
"impact": "none",
"message": "Identified (10:50 IST): We received reports of slowness and intermittent \"Something went wrong\" errors when accessing Empuls. We identified elevated resource contention on a backend service.\n\nMonitoring (11:10 IST): We applied a capacity increase to the affected service and adjusted its health-check configuration. Error rates returned to normal.\n\nResolved: The issue is fully resolved and the platform is operating normally. A root cause review has been completed and preventive measures, including improved capacity headroom and earlier detection alerting, are being implemented. We apologise for the disruption.",
"name": "Resolved — Intermittent errors and slowness on Empuls",
"timestamp": "Aug \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e05:20\u003c/var\u003e - \u003cvar data-var='time'\u003e05:20\u003c/var\u003e UTC"
}
],
"name": "August",
"year": 2026
}
],
"start_time": "2026-08-01T00:00:00Z",
"time_zone": "UTC"
},
{
"end_time": "2026-07-31T23:59:59Z",
"months": [
{
"incidents": [
{
"code": "nw1nbxqml38m",
"impact": "none",
"message": "We identified an issue causing order failures across multiple vendors and brands. The root cause has been traced to a dependency incompatibility in our vendor integration service, which prevented vendor configuration data from loading correctly. This affected approximately 1,500 orders. We deployed a fix its fine now.",
"name": "Multiple Vendor orders were failing",
"timestamp": "Jul \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e15:30\u003c/var\u003e - \u003cvar data-var='time'\u003e15:30\u003c/var\u003e UTC"
},
{
"code": "ml6ywy5ylwmv",
"impact": "major",
"message": "We got alerts for place order api saying its not working",
"name": "Storefront and PlumPro order were failed",
"timestamp": "Jul \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e12:30\u003c/var\u003e - \u003cvar data-var='time'\u003e12:30\u003c/var\u003e UTC"
}
],
"name": "July",
"year": 2026
},
{
"incidents": [
{
"code": "jrbf63xq2q9t",
"impact": "minor",
"message": "Due to a system issue due to a recent change the order report count fetching was failign and hence the entire report was not appearing.",
"name": "Order reports was failing to show details",
"timestamp": "Jun \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e16:30\u003c/var\u003e - \u003cvar data-var='time'\u003e16:30\u003c/var\u003e UTC"
},
{
"code": "pz7j38htfmqh",
"impact": "major",
"message": "Due to a change in the pipieline sync code, there was a DB SQL error, that resulted in failure in store.",
"name": "Outage on plum product listing",
"timestamp": "Jun \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e00:00\u003c/var\u003e - \u003cvar data-var='time'\u003e00:00\u003c/var\u003e UTC"
},
{
"code": "z2j1xry2hs53",
"impact": "none",
"message": "Due to issues with the DB runtime errors when there is a spike in traffic we had impact on order processing. Some orders were failing during this period. Team got alerts and provisioned additional database servers to balance the load.",
"name": "Partial incident due to database load related runtime errors",
"timestamp": "Jun \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e09:00\u003c/var\u003e - \u003cvar data-var='time'\u003e09:00\u003c/var\u003e UTC"
},
{
"code": "n6vhbyxgfpdq",
"impact": "minor",
"message": "Users accessing the Empuls platform via browser and MS Teams integration received HTTP 503 errors. The backend APIs (GraphQL, auth, integrations) remained fully operational throughout the incident. Only the frontend UI layer was affected.\nA deployment rollout of the empulsui service updated the pod template, generating a new rollouts-pod-template-hash (c4b9f4d6). The Kubernetes Service object empulsuiv2, which routes ingress traffic to the frontend pods, retained the old pod selector hash (69ddc85474) from the previous deployment.\nAs a result, the service had zero healthy endpoints — all incoming HTTP requests from the NGINX ingress controller were returned as 503s with the message no active Endpoint.\nThe underlying pods were healthy and running throughout; the issue was purely a selector mismatch between the Service and the active pods.",
"name": "Empuls Platform — Intermittent 503 Errors (Frontend Unavailable)",
"timestamp": "Jun \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e03:30\u003c/var\u003e - \u003cvar data-var='time'\u003e03:30\u003c/var\u003e UTC"
}
],
"name": "June",
"year": 2026
},
{
"incidents": [],
"name": "May",
"year": 2026
}
],
"start_time": "2026-05-01T00:00:00Z",
"time_zone": "UTC"
},
{
"end_time": "2026-04-30T23:59:59Z",
"months": [
{
"incidents": [
{
"code": "69qx6f4jpjl7",
"impact": "major",
"message": "Root cause: \nThe Dockerfile contained npm run build | true instead of RUN npm run build. The | true pattern pipes the build command's output, which has the side effect of always returning exit code 0 regardless of whether the build succeeded or failed. When the build failed (due to a code issue), Docker saw no error, happily continued building the image, and produced a container with no compiled static assets. nginx started up, found nothing in its web root, and fell back to its default page.\n\nA syntax error in the frontend code was shipped to production. The browser failed to parse or execute the malformed code, causing components to not render correctly — resulting in a broken or blank UI for users.\n\nImmediate fix: Replace npm run build | true with RUN npm run build. Optionally add a follow-up RUN step to assert the output directory is non-empty, so even a silent partial build would be caught.",
"name": "Storefront serving the nginx default page instead of the application.",
"timestamp": "Apr \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e14:10\u003c/var\u003e - \u003cvar data-var='time'\u003e14:10\u003c/var\u003e UTC"
},
{
"code": "w5v7r1p7sr35",
"impact": "major",
"message": "Timeline of diagnosis:\n\nInitial suspicion was the application layer — product listing pods were checked first and found healthy.\nInvestigation moved to the database, where slow-running MySQL queries were identified as the bottleneck.\nThose queries were killed manually, DB connections were freed, and the storefront recovered immediately.\n\nRoot cause: One or more long-running MySQL queries held database connections, exhausting the connection pool. The application pods were alive and healthy, but they couldn't complete DB calls to fetch product data — so the listing page appeared broken from the user's perspective.",
"name": "Storefront product listing page not rendering for users.",
"timestamp": "Apr \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e06:00\u003c/var\u003e - \u003cvar data-var='time'\u003e06:00\u003c/var\u003e UTC"
}
],
"name": "April",
"year": 2026
},
{
"incidents": [
{
"code": "n4dhqn3fycrz",
"impact": "major",
"message": "The Secret Vault backend service experienced a temporary outage when Kubernetes pods failed to start and entered an ImagePullBackOff state due to the required container image being unavailable.\n\nThe issue was traced to an automated image cleanup policy that removed the container image used by the running deployment, preventing Kubernetes from pulling the image.\n\nThe backend image was rebuilt and restored to the container registry, and the affected Kubernetes pods were restarted. Service recovery was monitored to ensure stability.\n\nThe service was successfully restored. Image retention policies have been updated to ensure images used by active deployments are preserved to prevent similar issues in the future.",
"name": "Issue in Xoxo-Points and Xoxo-Codes (XoxoVouchers) delivery",
"timestamp": "Mar \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e13:45\u003c/var\u003e - \u003cvar data-var='time'\u003e13:45\u003c/var\u003e UTC"
},
{
"code": "k044zsrfx6g7",
"impact": "major",
"message": "This incident has been resolved.",
"name": "Plum Storefront and Plum Pro api were down",
"timestamp": "Mar \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e12:38\u003c/var\u003e - \u003cvar data-var='time'\u003e12:39\u003c/var\u003e UTC"
},
{
"code": "8dj8mqg8lzkh",
"impact": "major",
"message": "T",
"name": "Service Degradation in UAE Region",
"timestamp": "Mar \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e08:52\u003c/var\u003e - Mar \u003cvar data-var='date'\u003e8\u003c/var\u003e, \u003cvar data-var='time'\u003e04:34\u003c/var\u003e UTC"
}
],
"name": "March",
"year": 2026
},
{
"incidents": [
{
"code": "kc2c3j9vv2n2",
"impact": "maintenance",
"message": "The scheduled maintenance has been completed.",
"name": "[Scheduled] Scheduled Maintenance: Empuls UAE Cluster",
"timestamp": "Feb \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e00:30\u003c/var\u003e - \u003cvar data-var='time'\u003e02:00\u003c/var\u003e UTC"
}
],
"name": "February",
"year": 2026
}
],
"start_time": "2026-02-01T00:00:00Z",
"time_zone": "UTC"
},
{
"end_time": "2026-01-31T23:59:59Z",
"months": [
{
"incidents": [
{
"code": "5dgbtztcrkwr",
"impact": "minor",
"message": "We experienced a brief increase in CPU utilization on our RDS instance, which led to a minor disruption in the Redemption service. The issue was promptly identified and mitigated, and normal service operations were restored within 30 minutes.There was no data loss, and the service is currently operating normally.\nWe are monitoring the systems closely and taking preventive measures to avoid recurrence.",
"name": "Service Impacted: Redemption Service",
"timestamp": "Jan \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e12:00\u003c/var\u003e - \u003cvar data-var='time'\u003e12:00\u003c/var\u003e UTC"
}
],
"name": "January",
"year": 2026
},
{
"incidents": [
{
"code": "9f6fx12yc8bl",
"impact": "major",
"message": "Earlier today, our services experienced a brief disruption due to a global outage impacting Cloudflare, which affected connectivity across multiple regions.Our systems remained fully functional, but external network dependencies caused temporary access issues.\nThe impact lasted for approximately 10 minutes, after which normal operations were restored as Cloudflare mitigated the incident.",
"name": "Resolved – Cloudflare Global Outage Impact",
"timestamp": "Dec \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e09:47\u003c/var\u003e - \u003cvar data-var='time'\u003e09:47\u003c/var\u003e UTC"
}
],
"name": "December",
"year": 2025
},
{
"incidents": [],
"name": "November",
"year": 2025
}
],
"start_time": "2025-11-01T00:00:00Z",
"time_zone": "UTC"
}
]
}