Snapshot 46253
Normalized text
Scripts and page chrome removed; this is what change detection compares.
{
"incidents": [
{
"created_at": "2026-07-26T08:19:22Z",
"id": "01KYER3RMA3X0W37TQP9AXNFCA",
"impact": "major",
"incident_updates": [
{
"body": "The issue has now been fully resolved. The root cause was scheduled maintenance from our Redis Provider failed, causing our instance to be in a degraded state. \n\nThe full write up is shared above.",
"created_at": "2026-07-26T08:48:05Z",
"display_at": "2026-07-26T08:40:00Z",
"id": "01KYESRARJ1KPFEHN51NRCZA6B",
"incident_id": "01KYER3RMA3X0W37TQP9AXNFCA",
"status": "resolved",
"updated_at": "2026-07-26T13:16:25Z"
},
{
"body": "The issue has been identified and we are working on resolving it.",
"created_at": "2026-07-26T08:19:22Z",
"display_at": "2026-07-26T08:25:00Z",
"id": "01KYER3RMABEY65T9T4VH9WG34",
"incident_id": "01KYER3RMA3X0W37TQP9AXNFCA",
"status": "identified",
"updated_at": "2026-07-26T09:20:53Z"
}
],
"name": "Redis Provider Issue",
"page_id": "01KKV0MW8BTP4TM4ZM2FC4T65F",
"resolved_at": "2026-07-26T08:40:00Z",
"status": "resolved",
"updated_at": "2026-07-26T13:32:01Z"
},
{
"created_at": "2026-07-23T09:19:48Z",
"id": "01KY74C7R7JMA75JG08N7E76W1",
"impact": "major",
"incident_updates": [
{
"body": "The issue has now been fully resolved.",
"created_at": "2026-07-23T09:42:18Z",
"display_at": "2026-07-23T09:35:00Z",
"id": "01KY75NEWMMP6JV5AZC78MX2DZ",
"incident_id": "01KY74C7R7JMA75JG08N7E76W1",
"status": "resolved",
"updated_at": "2026-07-26T09:22:33Z"
},
{
"body": "The issue has been identified and we are working on resolving it.",
"created_at": "2026-07-23T09:19:48Z",
"display_at": "2026-07-23T09:19:47Z",
"id": "01KY74C7R7RM940NVN76D2YCV9",
"incident_id": "01KY74C7R7JMA75JG08N7E76W1",
"status": "identified",
"updated_at": "2026-07-23T09:19:48Z"
}
],
"name": "Elevated 500 errors on customer methods",
"page_id": "01KKV0MW8BTP4TM4ZM2FC4T65F",
"resolved_at": "2026-07-23T09:35:00Z",
"status": "resolved",
"updated_at": "2026-07-26T13:09:02Z"
},
{
"created_at": "2026-07-23T00:23:25Z",
"id": "01KY65P3JJCX2V3SBK4T6PHTCV",
"impact": "major",
"incident_updates": [
{
"body": "The issue has now been fully resolved.\n\n**_--- POST MORTEM ---_**\n\nAutumn had three periods of partial outage overnight on July 23rd (UTC): 12:00–12:27AM, 4:10–4:20AM, and 9:00–9:26AM. During these windows, requests to our API failed with 503s - around 3% of requests in the first window, 12% in the second (peaking at 30% for a few minutes), and 5% in the third. In total, about 655k requests failed. Our database was overloaded by our own background jobs.\n\nNo events were lost, and we are replaying all customer / entity creation events. \n\n**What happened?**\n\nOur API is backed by a Postgres database, and two background systems write to it heavily: one that syncs usage balances, and one that resets entitlements at the end of billing periods. Neither had a limit on how much database work it could create at once.\n\nThree different triggers hit that same weakness:\n\n12:00AM (27 minutes). Several large customers send us traffic spikes at the top of every hour. At midnight, the spike landed while a legacy reset job was already consuming most of the database's capacity, and the balance sync work generated tens of thousands of tiny concurrent transactions. The database hit 100% CPU and the connection pool was exhausted. \n\n4:10AM (10 minutes). A long running script began executing. Each call unintentionally queued background reset work - roughly 244,000 jobs in 10 minutes.\n\n9:00AM (26 minutes). The same hourly traffic spike arrived, and this time the first transient failures triggered a retry loop from one client - 200,000 retries in 24 minutes. The retries kept the database pinned long after the original spike had passed. \n\n**What we're doing**\n\nThe root cause across all three incidents is the same: background database work had no global limit, so any burst could take down the database for everyone.\n\n- We've shipped a rebuilt entitlement reset job: it processes small bounded batches, can never overlap itself, and has a kill switch we can flip instantly.\n\n- We've shipped the ability to move heavy customers' usage tracking to async processing, which takes it off the critical database path.\n\n- We're adding hard concurrency limits and back-pressure to all background database work, so a burst queues instead of overwhelming the database.\n\n- We're adding admission control so a retry storm from one client can't extend an outage for everyone else.\n\nWe are happy to help and address any questions or concerns you may have \n\n- Ayush",
"created_at": "2026-07-23T00:31:10Z",
"display_at": "2026-07-23T00:31:10Z",
"id": "01KY6649GQ8WW6Q9PMBCQCKKNB",
"incident_id": "01KY65P3JJCX2V3SBK4T6PHTCV",
"status": "resolved",
"updated_at": "2026-07-23T11:58:35Z"
},
{
"body": "The issue has been identified and we are working on resolving it.",
"created_at": "2026-07-23T00:23:25Z",
"display_at": "2026-07-23T00:23:25Z",
"id": "01KY65P3JJ8C3WY27YF0N346VG",
"incident_id": "01KY65P3JJCX2V3SBK4T6PHTCV",
"status": "identified",
"updated_at": "2026-07-23T11:58:27Z"
}
],
"name": "Degraded Performance",
"page_id": "01KKV0MW8BTP4TM4ZM2FC4T65F",
"resolved_at": "2026-07-23T00:31:10Z",
"status": "resolved",
"updated_at": "2026-07-23T00:31:10Z"
},
{
"created_at": "2026-06-08T17:55:33Z",
"id": "01KTM6094FEXYPD2W7JW1VN2HM",
"impact": "major",
"incident_updates": [
{
"body": "Resolved",
"created_at": "2026-06-08T18:09:20Z",
"display_at": "2026-06-08T18:09:20Z",
"id": "01KTM6SH0PQFRHV5GCGRP145WT",
"incident_id": "01KTM6094FEXYPD2W7JW1VN2HM",
"status": "resolved",
"updated_at": "2026-06-08T18:09:20Z"
},
{
"body": "We are aware ~1% of requests are receiving 500 responses - we are actively working with our provider to resolve this.",
"created_at": "2026-06-08T17:55:33Z",
"display_at": "2026-06-08T17:55:33Z",
"id": "01KTM6094FCC7Y3H1NV61Z9EGM",
"incident_id": "01KTM6094FEXYPD2W7JW1VN2HM",
"status": "identified",
"updated_at": "2026-06-08T17:55:33Z"
}
],
"name": "Elevated 500 errors on customer methods",
"page_id": "01KKV0MW8BTP4TM4ZM2FC4T65F",
"resolved_at": "2026-06-08T18:09:20Z",
"status": "resolved",
"updated_at": "2026-07-26T13:14:14Z"
},
{
"created_at": "2026-04-17T12:12:45Z",
"id": "01KPDNQ7B2PN4XK62YE6CWXK35",
"impact": "minor",
"incident_updates": [
{
"body": "The issue has now been fully resolved.",
"created_at": "2026-04-17T12:12:45Z",
"display_at": "2026-04-17T04:11:00Z",
"id": "01KPDNQ7B2Q7N0KJFBD68KYA8H",
"incident_id": "01KPDNQ7B2PN4XK62YE6CWXK35",
"status": "resolved",
"updated_at": "2026-04-17T12:12:45Z"
},
{
"body": "We have been notified of the issue and are actively investigating. We will provide updates as soon as possible.",
"created_at": "2026-04-17T12:12:45Z",
"display_at": "2026-04-17T03:52:00Z",
"id": "01KPDNQ7B2AS4S2MNKYRJ4YZTV",
"incident_id": "01KPDNQ7B2PN4XK62YE6CWXK35",
"status": "investigating",
"updated_at": "2026-04-17T12:12:45Z"
}
],
"name": "Degraded Performance",
"page_id": "01KKV0MW8BTP4TM4ZM2FC4T65F",
"resolved_at": "2026-04-17T04:11:00Z",
"status": "resolved",
"updated_at": "2026-04-17T12:12:45Z"
},
{
"created_at": "2026-04-10T00:24:14Z",
"id": "01KNTCCV6SY11VJKDH8BGTCCZV",
"impact": "minor",
"incident_updates": [
{
"body": "The issue has now been fully resolved. ",
"created_at": "2026-04-10T00:24:14Z",
"display_at": "2026-04-09T19:25:00Z",
"id": "01KNTCCV6S165DW0T2Z172H4ZW",
"incident_id": "01KNTCCV6SY11VJKDH8BGTCCZV",
"status": "resolved",
"updated_at": "2026-04-10T00:24:14Z"
},
{
"body": "We've identified an issue where some of our API requests are taking longer than usual.",
"created_at": "2026-04-10T00:24:14Z",
"display_at": "2026-04-09T19:15:00Z",
"id": "01KNTCCV6SSG41GE1WFSTNFWC6",
"incident_id": "01KNTCCV6SY11VJKDH8BGTCCZV",
"status": "investigating",
"updated_at": "2026-04-10T00:24:14Z"
}
],
"name": "Degraded Service",
"page_id": "01KKV0MW8BTP4TM4ZM2FC4T65F",
"resolved_at": "2026-04-09T19:25:00Z",
"status": "resolved",
"updated_at": "2026-04-10T00:24:14Z"
},
{
"created_at": "2026-03-24T21:21:55Z",
"id": "01KMGVKGGTTWJ2GBBQ6H8PAAM7",
"impact": "major",
"incident_updates": [
{
"body": "The issue has now been fully resolved. Thank you for your patience and we will be publishing a post mortem on this soon. \n\n**_--- POST MORTEM ---_**\n\nAutumn had an outage between **8:15PM and 8:50PM UTC**. This affected **~30%** of requests to our API. One of our downstream providers, Redis Cloud, went down. We had a single point of failure with them.\n\n**What happened?**\n\nWe were investigating a minor Redis issue on our end. While we were investigating, we decided to upgrade our cluster.\n\nAn issue on their end caused our entire setup to fail. We were unable to restore it or deploy new instances.\n\n \n\n**What we’re doing**\n\nWe take full responsibility and recognise that we are at the point where a single point of failure is unacceptable:\n\n• We are making failovers and redundancy our p0. \n\n• We’re defining a formal process for touching any critical infra. \n\n• We have just hired a infra/dev-ops expert to scale our services reliably\n\n• We’ll be prioritising SDK safety measures (eg, fail-open by default)\n\n \n\n**What you can do**\nAn outage by us does not need to take your app down. This incident disproportionately affected customers that blocked usage on an Autumn error. This should never be the case.\n\n• If Autumn errors, allow requests to go through. The worst case is that users temporarily get additional usage. We can work with you to re-sync it after.\n\n• For extra safety, use our single webhook to replicate customer state into your own DB. If Autumn is down, you can fallback to that.\n\n• You can simulate errors from the Autumn API by removing your API key from your .env\n\nAny questions you have on the incident, or resolution, we're always here.",
"created_at": "2026-03-24T22:34:51Z",
"display_at": "2026-03-24T22:00:00Z",
"id": "01KMGZS2DDX648ZZF6CGE4N2D9",
"incident_id": "01KMGVKGGTTWJ2GBBQ6H8PAAM7",
"status": "resolved",
"updated_at": "2026-06-22T22:44:26Z"
},
{
"body": "The provider has been notified of the issue and are actively investigating (one cluster restored). We will provide updates as soon as possible.",
"created_at": "2026-03-24T21:21:55Z",
"display_at": "2026-03-24T21:21:55Z",
"id": "01KMGVKGGT8N2NH270560VK526",
"incident_id": "01KMGVKGGTTWJ2GBBQ6H8PAAM7",
"status": "investigating",
"updated_at": "2026-03-24T21:22:23Z"
}
],
"name": "Redis Provider Outage",
"page_id": "01KKV0MW8BTP4TM4ZM2FC4T65F",
"resolved_at": "2026-03-24T22:00:00Z",
"status": "resolved",
"updated_at": "2026-03-24T22:34:51Z"
},
{
"created_at": "2026-03-16T10:29:45Z",
"id": "01KKV33KVAP5FX8E651D19V71A",
"impact": "major",
"incident_updates": [
{
"body": "The issue has now been fully resolved.",
"created_at": "2026-03-16T10:29:45Z",
"display_at": "2026-03-15T15:45:00Z",
"id": "01KKV33KVAKZYCSTRV4RG49XAD",
"incident_id": "01KKV33KVAP5FX8E651D19V71A",
"status": "resolved",
"updated_at": "2026-06-22T22:37:25Z"
},
{
"body": "We have been notified of the issue and are actively investigating. We will provide updates as soon as possible.",
"created_at": "2026-03-16T10:29:45Z",
"display_at": "2026-03-15T15:00:00Z",
"id": "01KKV33KVAEHXGF2H6ZTV8DH3K",
"incident_id": "01KKV33KVAP5FX8E651D19V71A",
"status": "investigating",
"updated_at": "2026-03-16T10:29:45Z"
}
],
"name": "Partial Outage",
"page_id": "01KKV0MW8BTP4TM4ZM2FC4T65F",
"resolved_at": "2026-03-15T15:45:00Z",
"status": "resolved",
"updated_at": "2026-03-16T10:29:45Z"
},
{
"created_at": "2026-03-25T09:44:03Z",
"id": "01KMJ62DGSVMMRP3YX04K0X502",
"impact": "major",
"incident_updates": [
{
"body": "The issue has now been fully resolved.",
"created_at": "2026-03-25T09:44:03Z",
"display_at": "2025-11-17T07:45:00Z",
"id": "01KMJ62DGS692TC9B5FVYTPAYW",
"incident_id": "01KMJ62DGSVMMRP3YX04K0X502",
"status": "resolved",
"updated_at": "2026-03-25T09:44:03Z"
},
{
"body": "We have been notified of the issue and are actively investigating. We will provide updates as soon as possible.",
"created_at": "2026-03-25T09:44:03Z",
"display_at": "2025-11-17T07:05:00Z",
"id": "01KMJ62DGSF6NN1EGG5PEB4WZT",
"incident_id": "01KMJ62DGSVMMRP3YX04K0X502",
"status": "investigating",
"updated_at": "2026-03-25T09:44:03Z"
}
],
"name": "Infra Changes (DB overload)",
"page_id": "01KKV0MW8BTP4TM4ZM2FC4T65F",
"resolved_at": "2025-11-17T07:45:00Z",
"status": "resolved",
"updated_at": "2026-03-25T09:44:03Z"
}
],
"page": {
"id": "01KKV0MW8BTP4TM4ZM2FC4T65F",
"name": "Autumn",
"updated_at": "2026-07-26T08:40:00Z",
"url": "https://status.useautumn.com/"
}
}