Snapshot 21641
Normalized text
Scripts and page chrome removed; this is what change detection compares.
{
"incidents": [
{
"created_at": "2026-09-28T11:18:10Z",
"id": "01M3KVS5312YZX9F9YHSPT24KD",
"impact": "major",
"incident_updates": [
{
"body": "**_Monitoring_** _– Scheduled calculations have resumed and data points are updating as expected. Our team is actively monitoring the service and verifying data flow stability with affected users before marking the incident as fully resolved._\n",
"created_at": "2026-09-30T09:47:24Z",
"display_at": "2026-09-30T09:47:24Z",
"id": "01M3RVCCN9HQ6DRV7R3N3037Z1",
"incident_id": "01M3KVS5312YZX9F9YHSPT24KD",
"status": "monitoring",
"updated_at": "2026-09-30T09:47:24Z"
},
{
"body": "**Monitoring** - Fixes to scale worker capacity and improve service autoscaling have been deployed. We are observing scheduled calculations resuming execution. Due to the backlog accumulated during the degradation, jobs are currently catching up and may experience processing delays over the next few hours. We are actively monitoring the service until all scheduled calculations are fully up to date.",
"created_at": "2026-09-28T15:57:16Z",
"display_at": "2026-09-28T15:57:16Z",
"id": "01M3MBR6AMGZC7Y2AY20MERED5",
"incident_id": "01M3KVS5312YZX9F9YHSPT24KD",
"status": "monitoring",
"updated_at": "2026-09-28T15:57:16Z"
},
{
"body": "**Current Status Update — INC-3602 (Scheduled Calculations in Charts)**\n\n**Impact:**\nScheduled calculation executions in Charts are failing or significantly delayed, preventing new data points from being written. The issue currently appears isolated to cluster `az-tyo-gp-001` (verified working normally on `aw-tyo-001`).\n\n**Current Findings:**\n\n• High CPU utilization (100–1000%) observed on backend workers, causing jobs to fall hours behind schedule.\n\n• Extensive timeout errors and ~236 failed jobs recorded across the cluster over the weekend.\n\n• Newly created and existing scheduled calculations are affected.\n\n**Current Actions \u0026 Next Steps:**\n\n• Engineering is investigating the root cause of the CPU spike and timeout failures.\n\n• Evaluating remediation options to relieve backend load and process the backlog of calculations.\n\n ",
"created_at": "2026-09-28T12:28:19Z",
"display_at": "2026-09-28T12:28:19Z",
"id": "01M3KZSK9PEMP5WXMZMKVY43M5",
"incident_id": "01M3KVS5312YZX9F9YHSPT24KD",
"status": "investigating",
"updated_at": "2026-09-28T12:28:19Z"
},
{
"body": "_We are currently investigating an issue where scheduled calculations created in Charts are delayed or failing to write new data points. Our engineering team is actively investigating the root cause. Further updates will be provided as soon as they become available._",
"created_at": "2026-09-28T11:18:10Z",
"display_at": "2026-09-28T11:18:10Z",
"id": "01M3KVS531AB0ZR3QJ98TZT2FS",
"incident_id": "01M3KVS5312YZX9F9YHSPT24KD",
"status": "investigating",
"updated_at": "2026-09-28T11:18:10Z"
}
],
"monitoring_at": "2026-09-28T15:57:16Z",
"name": "Degraded Performance - Scheduled Calculations in Charts",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"status": "monitoring",
"updated_at": "2026-09-30T09:47:24Z"
},
{
"created_at": "2026-09-29T15:52:20Z",
"id": "01M3PXVWQQQWSRKE7429SY5D96",
"impact": "minor",
"incident_updates": [
{
"body": "Microsoft Azure has reported that the underlying platform issue in the `swedencentral` region has now been mitigated.\n\nOur internal metrics indicate that service health and model response times are recovering. We are actively observing system performance and endpoint latencies across all affected clusters to verify sustained stability.\n\nFurther updates will be shared as soon as full recovery is confirmed.",
"created_at": "2026-09-29T19:37:14Z",
"display_at": "2026-09-29T19:00:00Z",
"id": "01M3QAQNQ02XEQEDE0ZGYK0TKY",
"incident_id": "01M3PXVWQQQWSRKE7429SY5D96",
"status": "monitoring",
"updated_at": "2026-09-29T19:38:07Z"
},
{
"body": "We are currently tracking an external incident with Microsoft Azure affecting the Azure OpenAI Service in `swedencentral`, which impacts models deployed in the EU Datazone.\nDue to this underlying third-party provider issue outside of Cognite’s direct control, customers in the affected clusters may experience slower response times, increased LLM latencies, or intermittent request failures when using:\n\n• Atlas AI agents\n\n• Chat completion requests\n\n• Document summary features in Canvas\n\n\nIn some cases, switching to a non-OpenAI model may help mitigate the issue, however, tool calls may still require interaction with an OpenAI model and could continue to experience delays.\n\nOur team is actively monitoring the situation alongside Microsoft’s response. Further updates will be shared as soon as Azure services are restored. You can also view the direct status of the underlying Azure platform at \u003chttps://azure.status.microsoft/en-us/status\u003e",
"created_at": "2026-09-29T15:52:21Z",
"display_at": "2026-09-29T15:52:20Z",
"id": "01M3PXVWQQ1DJC7M2TZK8STN8C",
"incident_id": "01M3PXVWQQQWSRKE7429SY5D96",
"status": "investigating",
"updated_at": "2026-09-29T16:07:45Z"
}
],
"monitoring_at": "2026-09-29T19:00:00Z",
"name": "Slow response time on OpenAI Azure models EU",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"status": "monitoring",
"updated_at": "2026-09-29T19:37:14Z"
},
{
"created_at": "2026-09-22T10:08:52Z",
"id": "01M349DY9VV44Z7MDRJGGCTYJ7",
"impact": "minor",
"incident_updates": [
{
"body": "The performance degradation affecting Data Modeling search and indexing in **westeurope-1** has been fully resolved.\n\nSyncer capacity has been restored, cluster health metrics have returned to normal, and all indexing backlogs have cleared. Search, aggregation, and data modeling queries are functioning as expected.\n\nWe are continuing to monitor the services to ensure ongoing stability.\n\n\n",
"created_at": "2026-09-24T12:16:35Z",
"display_at": "2026-09-24T12:16:35Z",
"id": "01M39NH7HGP6W17825FQBF1QH7",
"incident_id": "01M349DY9VV44Z7MDRJGGCTYJ7",
"status": "resolved",
"updated_at": "2026-09-24T12:16:35Z"
},
{
"body": "The underlying cluster in `westeurope-1` has stabilized, and search and aggregate queries are functioning normally. Indexing pipelines are currently processing the remaining backlog, so some users may still see slight delays in newly ingested data appearing. We are monitoring the recovery closely.",
"created_at": "2026-09-22T10:31:57Z",
"display_at": "2026-09-22T10:31:57Z",
"id": "01M34AR6RPHDCBJ1SXWRWGM21Y",
"incident_id": "01M349DY9VV44Z7MDRJGGCTYJ7",
"status": "monitoring",
"updated_at": "2026-09-22T10:31:57Z"
},
{
"body": "We have identified an issue affecting search and data modeling indexing in the `westeurope-1` region. Users and automated pipelines in this cluster may experience delays in newly ingested data or schema updates appearing in search and data modeling views. Querying existing indexed data and other CDF services remain operational.\n\nOur engineering team has implemented traffic-shedding and resource adjustments to stabilize the underlying cluster, and indexing throughput is beginning to recover. We are closely monitoring cluster health and will provide our next update soon,",
"created_at": "2026-09-22T10:08:52Z",
"display_at": "2026-09-22T10:08:52Z",
"id": "01M349DY9V2RJZKCXF9FDHK4JZ",
"incident_id": "01M349DY9VV44Z7MDRJGGCTYJ7",
"status": "investigating",
"updated_at": "2026-09-22T10:08:52Z"
}
],
"name": " Degraded Data Modeling Indexing and Search in westeurope-1",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-09-24T12:16:35Z",
"status": "resolved",
"updated_at": "2026-09-24T12:16:35Z"
},
{
"created_at": "2026-09-23T17:45:19Z",
"id": "01M37NYE7ZM08MSQB1ST0KN815",
"impact": "minor",
"incident_updates": [
{
"body": "The performance degradation affecting Data Modeling search and indexing in **westeurope-1** has been fully resolved.\n\nSyncer capacity has been restored, cluster health metrics have returned to normal, and all indexing backlogs have cleared. Search, aggregation, and data modeling queries are functioning as expected.\n\nWe are continuing to monitor the services to ensure ongoing stability.",
"created_at": "2026-09-24T12:08:01Z",
"display_at": "2026-09-24T12:08:01Z",
"id": "01M39N1H5E3AZPT1CAMWC20D0J",
"incident_id": "01M37NYE7ZM08MSQB1ST0KN815",
"status": "resolved",
"updated_at": "2026-09-24T12:08:01Z"
},
{
"body": "• **Problem:** The Propertygraph Elasticsearch cluster in **westeurope-1** experienced degradation, temporarily entering yellow and red states.\n\n• **Impact:** Customers with projects in `westeurope-1` may experience delays in newly created instances appearing in search results, along with elevated latency for search and aggregate queries affecting data modeling indexing (`datamodelstorage`) Core data storage is unaffected.\n\n• **Root Cause:** A syncer update triggered concurrent restarts across syncer deployments simultaneously, overwhelming the search master nodes with high-volume reads. In addition, an search pod encountered a PVC mount failure during rolling updates.\n\n• **Status \u0026 Mitigation:** Mitigation steps have been applied. Our engineering team scaled down the syncers to allow the cluster to recover to a healthy green state, subsequently re-enabled the synchronizers in stages, and resolved the PVC volume mount issue. The cluster is stabilizing, and our team is actively monitoring system performance and recovery.",
"created_at": "2026-09-23T19:17:38Z",
"display_at": "2026-09-23T19:10:00Z",
"id": "01M37V7FCN1QB3ERRQR19XP56B",
"incident_id": "01M37NYE7ZM08MSQB1ST0KN815",
"status": "monitoring",
"updated_at": "2026-09-23T19:18:00Z"
},
{
"body": "We are currently experiencing performance degradation affecting search and data modeling indexing services in the **westeurope-1** region.\n**Impact:** Search and aggregate queries may experience elevated latency, and newly created instances or index updates may take longer to appear in search results. Core data storage is unaffected.\nOur engineering team is actively investigating the issue and working on mitigation steps to stabilize the service. We will provide updates as more information becomes available",
"created_at": "2026-09-23T17:45:19Z",
"display_at": "2026-09-23T17:45:19Z",
"id": "01M37NYE7ZTC0B4F1ESWK0VFE4",
"incident_id": "01M37NYE7ZM08MSQB1ST0KN815",
"status": "investigating",
"updated_at": "2026-09-23T17:45:19Z"
}
],
"name": "Degraded search and data modeling indexing performance",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-09-24T12:08:01Z",
"status": "resolved",
"updated_at": "2026-09-24T12:08:01Z"
},
{
"created_at": "2026-09-22T21:06:27Z",
"id": "01M35F20AB8EVKK1F303HZ26DK",
"impact": "minor",
"incident_updates": [
{
"body": "Service has been restored. The Elasticsearch cluster in westeurope-1 has stabilized, and search, aggregation, and data-modeling indexing are operating normally. ",
"created_at": "2026-09-23T09:05:28Z",
"display_at": "2026-09-23T09:05:28Z",
"id": "01M36R6HWZ5YQPGP8VD7ER2PYX",
"incident_id": "01M35F20AB8EVKK1F303HZ26DK",
"status": "resolved",
"updated_at": "2026-09-23T09:05:28Z"
},
{
"body": "The Elasticsearch cluster in **westeurope-1** became overloaded and entered a red state. This degraded search and data-modeling indexing. search and aggregate requests could be slow or fail, while new instances and index changes could take longer to appear.\n\n**Root cause:** A rollout restart of **pg3** overloaded the Elasticsearch cluster.\n\n**Fixing status:** The incident is now in **Monitoring**. Engineering team paused some background synchronizers to reduce load, allowed Elasticsearch to stabilize, and then started re-enabling the synchronizers. The fix was marked as applied.",
"created_at": "2026-09-23T00:07:29Z",
"display_at": "2026-09-22T22:43:00Z",
"id": "01M35SDFSKC1AA1PK25K3G6ER1",
"incident_id": "01M35F20AB8EVKK1F303HZ26DK",
"status": "monitoring",
"updated_at": "2026-09-23T00:08:09Z"
},
{
"body": "We are currently investigating performance degradation and high latency affecting search and data modeling services in the **westeurope-1** region.\n\n**Impact:** Search and aggregate requests may experience high latency or temporary failures. Newly created instances or index updates may take longer than expected to appear.\n\n**Status:** Our engineering team is currently scaling down background syncing processes to reduce load and restore cluster stability. Further updates will be posted as soon as they are available.",
"created_at": "2026-09-22T21:06:27Z",
"display_at": "2026-09-22T19:57:00Z",
"id": "01M35F20AB6SHYZD54JET8MFA8",
"incident_id": "01M35F20AB8EVKK1F303HZ26DK",
"status": "investigating",
"updated_at": "2026-09-22T21:10:57Z"
}
],
"name": "Indexing in data modeling is degraded",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-09-23T09:05:28Z",
"status": "resolved",
"updated_at": "2026-09-23T09:05:28Z"
},
{
"created_at": "2026-09-11T07:20:32Z",
"id": "01M27NDSVKCS0JGAMJHPHH6RCM",
"impact": "minor",
"incident_updates": [
{
"body": "Increased error rates were found to be confined to a particular client configuration with excessive retries, and were mitigated. Overall cluster availability remained nominal.",
"created_at": "2026-09-18T19:18:43Z",
"display_at": "2026-09-18T19:18:43Z",
"id": "01M2TZ9VK7ZXYXZZG1RWF6491D",
"incident_id": "01M27NDSVKCS0JGAMJHPHH6RCM",
"status": "resolved",
"updated_at": "2026-09-18T19:18:43Z"
},
{
"body": "**We are currently investigating an issue impacting the Data Modeling service (**`datamodelstorage`**) in the** `az-arn-001` **cluster.**\n\n**A large spike in traffic against Data Modeling endpoints has led to an elevated rate of 503 errors (affecting roughly 3% of requests) in this cluster. Users and services interacting with Data Modeling endpoints may experience intermittent request failures.**\n\n**Our engineering team is actively investigating the cause of the errors and working on mitigation steps. Further updates will be provided as more information becomes available**",
"created_at": "2026-09-11T07:20:32Z",
"display_at": "2026-09-11T07:20:32Z",
"id": "01M27NDSVKBWQQFPHHD483XK13",
"incident_id": "01M27NDSVKCS0JGAMJHPHH6RCM",
"status": "investigating",
"updated_at": "2026-09-11T07:20:32Z"
}
],
"name": "Degraded Data Modeling Availability in az-arn-001",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-09-18T19:18:43Z",
"status": "resolved",
"updated_at": "2026-09-18T19:18:43Z"
},
{
"created_at": "2026-09-11T02:44:36Z",
"id": "01M275MJA0QEH7HH6DE0Z740GW",
"impact": "minor",
"incident_updates": [
{
"body": "**The engineering team has allocated additional memory resources to the affected services, and normal performance has been restored with error rates dropping to zero. We are continuing to closely monitor the system to ensure ongoing stability. Further updates will be provided once the incident is fully resolved.**",
"created_at": "2026-09-11T03:09:14Z",
"display_at": "2026-09-11T03:09:14Z",
"id": "01M2771MSQA4JRCGQ10MM6ZCCN",
"incident_id": "01M275MJA0QEH7HH6DE0Z740GW",
"status": "monitoring",
"updated_at": "2026-09-11T03:09:14Z"
},
{
"body": "**We are currently investigating an issue causing elevated 503 error rates and service degradation for Data Modeling services within the az-tyo-gp-001 cluster. Users and applications in this cluster may experience intermittent errors or failures when accessing data modeling endpoints and related features.**\n\n**Our engineering team is actively investigating the root cause to restore normal operations. Further updates will be shared as more information becomes available.**",
"created_at": "2026-09-11T02:44:36Z",
"display_at": "2026-09-11T02:44:36Z",
"id": "01M275MJA0NR18943NA1JPD9RE",
"incident_id": "01M275MJA0QEH7HH6DE0Z740GW",
"status": "investigating",
"updated_at": "2026-09-11T02:44:36Z"
}
],
"monitoring_at": "2026-09-11T03:09:14Z",
"name": "Data Modeling service degradation in az-tyo-gp-001",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"status": "monitoring",
"updated_at": "2026-09-11T03:09:14Z"
},
{
"created_at": "2026-09-06T06:17:20Z",
"id": "01M1TNTFDEZYYCMPHHA79MZPF1",
"impact": "none",
"incident_updates": [
{
"body": "Services are fully recovered and operating normally. The elevated 5XX error rates affecting the **Timeseries** and **Sequences** APIs on cluster **az-pnq-gp-001** between approximately **07:18 and 07:45 CEST** on **September 6, 2026** were resolved following auto-healing of the FoundationDB operator pod instability. Monitoring confirms all deployments remain healthy and stable, with no data loss incurred.",
"created_at": "2026-09-06T14:54:43Z",
"display_at": "2026-09-06T14:54:43Z",
"id": "01M1VKDTTEWE00K5V8D5M6JGRA",
"incident_id": "01M1TNTFDEZYYCMPHHA79MZPF1",
"status": "resolved",
"updated_at": "2026-09-06T14:54:43Z"
},
{
"body": "**We experienced elevated 5XX error rates affecting the Timeseries and Sequences APIs on cluster** az-pnq-gp-001 **between approximately 07:18 and 07:45 CEST. Services have recovered and all deployments are healthy, with no indication of data loss. The engineering team is actively monitoring the service to ensure continued stability while completing the investigation into the underlying cause.**",
"created_at": "2026-09-06T06:17:20Z",
"display_at": "2026-09-06T05:18:00Z",
"id": "01M1TNTFDEDCH5F1K3PDZRE854",
"incident_id": "01M1TNTFDEZYYCMPHHA79MZPF1",
"status": "investigating",
"updated_at": "2026-09-06T06:21:46Z"
}
],
"name": "Elevated error rates on Timeseries and Sequences APIs",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-09-06T14:54:43Z",
"status": "resolved",
"updated_at": "2026-09-06T14:54:43Z"
},
{
"created_at": "2026-09-06T11:34:36Z",
"id": "01M1V7ZDPNGS835SN3RV3K8NMR",
"impact": "major",
"incident_updates": [
{
"body": "This incident has been resolved. The issue causing HTTP 500/503 errors when searching or aggregating Data Models in `westeurope-1` has been mitigated, and service performance has fully recovered. All systems are operating normally.",
"created_at": "2026-09-06T14:51:56Z",
"display_at": "2026-09-06T14:51:56Z",
"id": "01M1VK8R2XDGXFT3QY44M6QCGS",
"incident_id": "01M1V7ZDPNGS835SN3RV3K8NMR",
"status": "resolved",
"updated_at": "2026-09-06T14:51:56Z"
},
{
"body": "This incident has been resolved. The issue causing HTTP 500/503 errors when searching or aggregating Data Models in `westeurope-1` has been mitigated, and service performance has fully recovered. All systems are operating normally.",
"created_at": "2026-09-06T14:51:37Z",
"display_at": "2026-09-06T14:51:37Z",
"id": "01M1VK855WCV2NQKYW08V2S3R1",
"incident_id": "01M1V7ZDPNGS835SN3RV3K8NMR",
"status": "investigating",
"updated_at": "2026-09-06T14:51:37Z"
},
{
"body": "The engineering team is investigating and fixing issues with the Data Modelling search and aggregate APIs in the westeurope-1 cluster. Customer impact includes increased error rates on DM APIs and in the data exploration app in Fusion, and delays before new data is available in search results.",
"created_at": "2026-09-06T11:34:36Z",
"display_at": "2026-09-06T11:34:36Z",
"id": "01M1V7ZDPNXC3AHQ6M28B7650Z",
"incident_id": "01M1V7ZDPNGS835SN3RV3K8NMR",
"status": "investigating",
"updated_at": "2026-09-06T11:34:36Z"
}
],
"name": "Users experiencing HTTP 500/503 errors when searching or aggregating Data Models",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-09-06T14:51:56Z",
"status": "resolved",
"updated_at": "2026-09-06T14:51:56Z"
},
{
"created_at": "2026-09-04T08:17:35Z",
"id": "01M1NQX83ATXG612QGCQTAFA91",
"impact": "major",
"incident_updates": [
{
"body": "• **Problem:** The production `propertygraph-search` Elasticsearch cluster in `westeurope-1` went red for roughly 10 minutes.\n\n• **Root cause:** Propertygraph syncer activity and cluster topology created excessive master-coordination overhead. PVC mount failures were also observed on Elasticsearch master-data pods, but were not confirmed as the original cause.\n\n• **Resolution:** Engineering Team scaled down the propertygraph syncers to recover cluster health, then restored every `propertygraph-sync-v3` deployment to one replica in stages. They also reduced node/master-coordination overhead and reconfigured all 15 master-data nodes with 19 GB heaps; the cluster reached zero unassigned shards and was considered stable.",
"created_at": "2026-09-04T17:58:17Z",
"display_at": "2026-09-04T17:58:17Z",
"id": "01M1PS4H5E1VRAX8MR6KY0PBQA",
"incident_id": "01M1NQX83ATXG612QGCQTAFA91",
"status": "resolved",
"updated_at": "2026-09-04T17:58:17Z"
},
{
"body": "The engineering team is investigating and fixing issues with the Data Modelling search and aggregate APIs in the westeurope-1 cluster. Customer impact includes increased error rates on DM APIs and in the data exploration app in Fusion, and delays before new data is available in search results.",
"created_at": "2026-09-04T10:29:30Z",
"display_at": "2026-09-04T10:29:30Z",
"id": "01M1NZERWYRFAZD72QW7ZDSK3D",
"incident_id": "01M1NQX83ATXG612QGCQTAFA91",
"status": "monitoring",
"updated_at": "2026-09-04T10:29:30Z"
},
{
"body": "The engineering team is investigating and fixing issues with the Data Modelling search and aggregate APIs in the westeurope-1 cluster. Customer impact includes increased error rates on DM APIs and in the data exploration app in Fusion, and delays before new data is available in search results.",
"created_at": "2026-09-04T08:18:27Z",
"display_at": "2026-09-04T08:18:27Z",
"id": "01M1NQYTWZWDD09KVYMKXKQ6GB",
"incident_id": "01M1NQX83ATXG612QGCQTAFA91",
"status": "monitoring",
"updated_at": "2026-09-04T08:18:27Z"
},
{
"body": "The engineering team is investigating and fixing issues with the Data Modelling search and aggregate APIs in the westeurope-1 cluster. Customer impact includes increased error rates on DM APIs and in the data exploration app in Fusion, and delays before new data is available in search results.",
"created_at": "2026-09-04T08:17:35Z",
"display_at": "2026-09-04T08:17:35Z",
"id": "01M1NQX83A56G8ER30Z8MW5552",
"incident_id": "01M1NQX83ATXG612QGCQTAFA91",
"status": "monitoring",
"updated_at": "2026-09-04T08:17:35Z"
}
],
"name": "Users experiencing HTTP 500/503 errors when searching or aggregating Data Models",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-09-04T17:58:17Z",
"status": "resolved",
"updated_at": "2026-09-04T17:58:17Z"
},
{
"created_at": "2026-09-03T14:51:01Z",
"id": "01M1KW0XV8G3RTHHGQ8EERF7R1",
"impact": "major",
"incident_updates": [
{
"body": "The issue causing login failures across Cognite Data Fusion, InField, and Maintain has been resolved. The underlying deployment change was rolled back, and normal service has been fully restored across all clusters.",
"created_at": "2026-09-03T15:05:12Z",
"display_at": "2026-09-03T15:05:12Z",
"id": "01M1KWTWEV3TF88A1G5R2QWHZ8",
"incident_id": "01M1KW0XV8G3RTHHGQ8EERF7R1",
"status": "resolved",
"updated_at": "2026-09-03T15:05:12Z"
},
{
"body": "**Identified** - We are investigating an issue affecting user logins across Cognite Data Fusion (Fusion), Infield, and Maintain. Some users may experience login failures displaying the error: `\"An HTTP line is larger than 4096 bytes\"`.\n\nA recent identity service update has been identified as the likely cause, and our engineering team is actively deploying a rollback to restore full login functionality.\n\nWe will provide another update soon as the rollback is completed.",
"created_at": "2026-09-03T14:51:01Z",
"display_at": "2026-09-03T14:51:01Z",
"id": "01M1KW0XV8RNJ3H3VTJ9FQ79C7",
"incident_id": "01M1KW0XV8G3RTHHGQ8EERF7R1",
"status": "monitoring",
"updated_at": "2026-09-03T14:51:01Z"
}
],
"name": "Login Issues in Cognite Applications (Fusion, Infield, Maintain)",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-09-03T15:05:12Z",
"status": "resolved",
"updated_at": "2026-09-03T15:05:12Z"
},
{
"created_at": "2026-09-01T12:05:28Z",
"id": "01M1EDRAZ2T9GJTFQYCTW1DQVM",
"impact": "major",
"incident_updates": [
{
"body": "**Summary:** We experienced an issue affecting the Property Graph search service in the West Europe region. The Elasticsearch cluster became unstable after receiving unusually large synchronization requests, which caused circuit-breaker errors.\n\n**Steps taken:** We stabilized the Elasticsearch cluster, then re-enabled background synchronization in controlled stages while monitoring cluster health, CPU usage, and indexing rates.\n\n**Resolved status:** The service has stabilized and normal functionality has been restored. Background synchronization has resumed, and we will continue monitoring to ensure it remains healthy.",
"created_at": "2026-09-01T18:46:49Z",
"display_at": "2026-09-01T18:32:00Z",
"id": "01M1F4Q7DHDYS9FNBQ0BJNWW15",
"incident_id": "01M1EDRAZ2T9GJTFQYCTW1DQVM",
"status": "resolved",
"updated_at": "2026-09-01T18:47:23Z"
},
{
"body": "We have implemented mitigation steps for the issue impacting Data Model (DM) Search and aggregation queries in the westeurope-1 region. The service has stabilized, and we are currently monitoring the system to ensure full recovery.\n\nSearch and aggregation requests for data models in westeurope-1 should now complete successfully without HTTP 503 errors. Customers may still see slight delays in data freshness as background sync operations catch up, but normal functionality has been largely restored. Core data storage and other regions continue to operate normally.\n\nOur engineering team has successfully stabilized the underlying Elasticsearch cluster and carefully re-enabled background sync operations in controlled stages. We will continue to closely monitor cluster health and performance metrics to ensure the service remains stable.",
"created_at": "2026-09-01T18:40:09Z",
"display_at": "2026-09-01T18:30:00Z",
"id": "01M1F4B1N238EJ9EZFTQFPY7PQ",
"incident_id": "01M1EDRAZ2T9GJTFQYCTW1DQVM",
"status": "monitoring",
"updated_at": "2026-09-01T18:43:02Z"
},
{
"body": "**Summary:**\nWe are currently investigating an issue impacting Data Model (DM) Search and aggregation queries in the `westeurope-1` region.\n\n**Customer Impact:**\nRequests to search or aggregate data models in `westeurope-1` may fail with HTTP 503 errors or return incomplete/delayed results. Core data storage and other regions are operating normally.\n\n**Current Actions:**\nOur engineering team is actively working on cluster recovery and load mitigation, including throttling background sync operations to restore search performance.\n\n**Next Update:**\nWe are monitoring recovery progress and will provide another update within 60 minutes or as soon as significant updates are available.",
"created_at": "2026-09-01T12:05:28Z",
"display_at": "2026-09-01T12:05:28Z",
"id": "01M1EDRAZ2ZPFTKYSXRC59WR33",
"incident_id": "01M1EDRAZ2T9GJTFQYCTW1DQVM",
"status": "investigating",
"updated_at": "2026-09-01T12:05:28Z"
}
],
"name": "Degraded Data Model Search \u0026 Aggregations in westeurope-1",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-09-01T18:32:00Z",
"status": "resolved",
"updated_at": "2026-09-01T18:46:49Z"
},
{
"created_at": "2026-08-24T13:52:34Z",
"id": "01M0T0PPQ604AA0CXBAWPVYMGH",
"impact": "none",
"incident_updates": [
{
"body": "The issue got fixed on 27th August. All Apps are now up and running.",
"created_at": "2026-08-31T12:07:03Z",
"display_at": "2026-08-31T12:07:03Z",
"id": "01M1BVEGM9QMYSPFKMC285JF4H",
"incident_id": "01M0T0PPQ604AA0CXBAWPVYMGH",
"status": "resolved",
"updated_at": "2026-08-31T12:07:03Z"
},
{
"body": "We have successfully resolved the underlying deployment issue. Our engineering team corrected the production deployment configuration, ensuring it now points to the correct Firebase project and hosting site. The fix has been verified across multiple production clusters and confirmed to be working as expected.\nIf you had previously applied the temporary workaround of adding `ssl` to your application's **INSTALLED PACKAGES**, no further action is needed this entry can remain in place and will not interfere with the platform fix.",
"created_at": "2026-08-27T11:06:43Z",
"display_at": "2026-08-27T11:06:43Z",
"id": "01M11ED67TFVY205796CMRYJPR",
"incident_id": "01M0T0PPQ604AA0CXBAWPVYMGH",
"status": "monitoring",
"updated_at": "2026-08-27T11:06:43Z"
},
{
"body": "The issue leading to the `ModuleNotFoundError: No module named 'ssl'` error in Streamlit apps has been identified as a change in a third party library, specifically the Microsoft Authentication Library for Python (MSAL). MSAL is commonly used in conjuction with the Cognite SDK to handle authentication.\n\nMSAL version 1.38, released on 2026-08-24, references the `ssl` module without declaring a dependency on it. This works on ordinary CPython installations because that environment includes the `ssl` module by default.\n\nHowever, Streamlit in the browser uses the Pyodide runtime environment. Pyodide does not include the `ssl` module by default. This leads to failures when the latest MSAL library is used in Streamlit apps.\n\nThe engineering team is currently adding a workaround to the Cognite Streamlit runtime environment setup to ensure that the `ssl` module is always present.",
"created_at": "2026-08-25T07:44:39Z",
"display_at": "2026-08-25T07:44:39Z",
"id": "01M0VY1QQ14FKZ97K1XAF9SCEX",
"incident_id": "01M0T0PPQ604AA0CXBAWPVYMGH",
"status": "identified",
"updated_at": "2026-08-25T07:44:39Z"
},
{
"body": "We are investigating issues with CDF-hosted Streamlit apps, where apps get the error **ModuleNotFoundError: No module named 'ssl'**\nWe are investigating the issue. The issue can affect all users of Streamlit apps regardless of which CDF clusters they are hosted in. In the meantime, explicitly adding the \"ssl\" module to your app's Python requirements can fix the immediate problem and is worth trying as a workaround.\n\nNote: CDF-hosted Streamlit apps are in [public preview](https://docs.cognite.com/cdf/streamlit/index#streamlit-apps) and are not yet in [production](https://docs.cognite.com/cdf/product_feature_status#public-preview).",
"created_at": "2026-08-24T13:52:34Z",
"display_at": "2026-08-24T13:52:34Z",
"id": "01M0T0PPQ6JD26AEEQXPHRSVGC",
"incident_id": "01M0T0PPQ604AA0CXBAWPVYMGH",
"status": "investigating",
"updated_at": "2026-08-24T13:52:34Z"
}
],
"name": "CDF hosted Streamlit apps failing due to missing \"ssl\" dependency",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-08-31T12:07:03Z",
"status": "resolved",
"updated_at": "2026-08-31T12:07:03Z"
},
{
"created_at": "2026-08-19T17:37:18Z",
"id": "01M0DHJKAG6DC2N9BCFVGV4V30",
"impact": "critical",
"incident_updates": [
{
"body": "Issue is fixed, and no further degradations were observed.",
"created_at": "2026-08-20T09:06:51Z",
"display_at": "2026-08-20T09:06:51Z",
"id": "01M0F6RNR7324HY6WKPYS6C8DJ",
"incident_id": "01M0DHJKAG6DC2N9BCFVGV4V30",
"status": "resolved",
"updated_at": "2026-08-20T09:06:51Z"
},
{
"body": "The database issue has been resolved and services are recovering. We are monitoring the situation.",
"created_at": "2026-08-19T18:02:45Z",
"display_at": "2026-08-19T17:46:00Z",
"id": "01M0DK16D70XDJ96GFGG5KR5A5",
"incident_id": "01M0DHJKAG6DC2N9BCFVGV4V30",
"status": "monitoring",
"updated_at": "2026-08-19T18:03:14Z"
},
{
"body": "We are aware of a service outage affecting the CDF cluster westeurope-1. The issue has severe impact on some but not all of the projects in cluster and are caused by a database server outage after a routine maintenance operation. Cognite's engineering team is actively working with our cloud provider to resolve the issue.",
"created_at": "2026-08-19T17:37:18Z",
"display_at": "2026-08-19T16:52:00Z",
"id": "01M0DHJKAGTNYZSG3N8XR17NR8",
"incident_id": "01M0DHJKAG6DC2N9BCFVGV4V30",
"status": "investigating",
"updated_at": "2026-08-19T17:39:46Z"
}
],
"name": "Partial service outage affecting CDF westeurope-1",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-08-20T09:06:51Z",
"status": "resolved",
"updated_at": "2026-08-20T09:06:51Z"
},
{
"created_at": "2026-08-13T19:33:28Z",
"id": "01KZY9V0M8HV20XE8FTXJR0AZY",
"impact": "minor",
"incident_updates": [
{
"body": "This incident has been resolved. It is operating normally. We apologize for any inconvenience and thank you for your patience.",
"created_at": "2026-08-14T12:44:39Z",
"display_at": "2026-08-14T12:44:39Z",
"id": "01M004V595J3YJK7ZPFFNTNQ2R",
"incident_id": "01KZY9V0M8HV20XE8FTXJR0AZY",
"status": "resolved",
"updated_at": "2026-08-14T12:44:39Z"
},
{
"body": "Functions requests on az-dsm-psaas-001 returned 500s for a temporary period between 19:25 and 19:40 CEST. The immediate impact was resolved at 19:40 CET / 17:40 UTC, and we are now monitoring the clusters to ensure continued stability.",
"created_at": "2026-08-13T19:34:50Z",
"display_at": "2026-08-13T17:40:00Z",
"id": "01KZY9XGG5YFT1JQZ1N3E59PY6",
"incident_id": "01KZY9V0M8HV20XE8FTXJR0AZY",
"status": "monitoring",
"updated_at": "2026-08-13T19:37:22Z"
},
{
"body": "We are currently investigating an elevated rate of 5xx errors affecting the Functions API. The issue appears to be caused by an upstream dependency disruption (Azure Key Vault RBAC loading timeout) in the Central US region, rather than a direct cluster-level outage. The connection rates are already stabilizing, and we are monitoring the recovery.",
"created_at": "2026-08-13T19:33:28Z",
"display_at": "2026-08-13T17:25:00Z",
"id": "01KZY9V0M89YZHQ60BXXCMQTS3",
"incident_id": "01KZY9V0M8HV20XE8FTXJR0AZY",
"status": "investigating",
"updated_at": "2026-08-13T19:37:02Z"
}
],
"name": "Functions API - Degraded Performance",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-08-14T12:44:39Z",
"status": "resolved",
"updated_at": "2026-08-14T12:44:39Z"
},
{
"created_at": "2026-08-11T10:14:35Z",
"id": "01KZR527988ZK33KSH0PYMXSEF",
"impact": "critical",
"incident_updates": [
{
"body": "This incident has been resolved. Cogshop service is up and operating normally. We apologize for any inconvenience and thank you for your patience.",
"created_at": "2026-08-11T13:41:28Z",
"display_at": "2026-08-11T13:41:28Z",
"id": "01KZRGX0RR5TSR6Z82WA1JXHWV",
"incident_id": "01KZR527988ZK33KSH0PYMXSEF",
"status": "resolved",
"updated_at": "2026-08-11T13:41:28Z"
},
{
"body": "We identified the issue as a missing TLS certificate in the `wildcard-cognite-ai-tls-patch` secret. Access to the `powerops-certmgr-vault` key vault was updated to allow regeneration/import of the certificate, and we can now see the certificate has been recreated.\n\nThe Cogshop endpoint is responding again: [shop-api.az-inso-powerops.cognite.ai/docs](https://shop-api.az-inso-powerops.cognite.ai/docs). We are continuing to monitor the service closely to confirm it remains stable before marking this incident as resolved.",
"created_at": "2026-08-11T13:19:15Z",
"display_at": "2026-08-11T13:19:15Z",
"id": "01KZRFMBPP81N7DC2AZ8BZT9B7",
"incident_id": "01KZR527988ZK33KSH0PYMXSEF",
"status": "monitoring",
"updated_at": "2026-08-11T13:19:15Z"
},
{
"body": "We are currently investigating an issue causing the Cogshop service to be unavailable. Our engineering team has identified the root cause and is actively working on a fix. We will provide further updates as soon as more information is available.",
"created_at": "2026-08-11T10:14:35Z",
"display_at": "2026-08-11T10:14:35Z",
"id": "01KZR527988XWDSJ98YHXH4WGA",
"incident_id": "01KZR527988ZK33KSH0PYMXSEF",
"status": "investigating",
"updated_at": "2026-08-11T10:14:35Z"
}
],
"name": "Cogshop Service Disruption",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-08-11T13:41:28Z",
"status": "resolved",
"updated_at": "2026-08-11T13:41:28Z"
},
{
"created_at": "2026-07-22T14:42:30Z",
"id": "01KY54EDYJ6BDWVTKZ82TERR4Z",
"impact": "minor",
"incident_updates": [
{
"body": "The license key configuration for **Cognite InField** has been updated, and the missing license key watermark is now fully resolved across all affected environments.\n\nAll services are operating normally. Thank you for your patience!",
"created_at": "2026-07-22T15:17:14Z",
"display_at": "2026-07-22T15:17:14Z",
"id": "01KY56E0M33S7EFVQR91AN930H",
"incident_id": "01KY54EDYJ6BDWVTKZ82TERR4Z",
"status": "resolved",
"updated_at": "2026-07-22T15:17:14Z"
},
{
"body": "We are currently investigating an issue affecting the **InField** application where a UI license watermark (\"MUI X Missing License Key\") is displayed.\n\n**Impact:** This is a minor visual UI issue and does not impact functionality, data integrity, or core user workflows. Our engineering teams are actively investigating the root cause and working on a resolution.",
"created_at": "2026-07-22T14:42:31Z",
"display_at": "2026-07-22T14:42:30Z",
"id": "01KY54EDYJ4B9SM3C88HA5N9EA",
"incident_id": "01KY54EDYJ6BDWVTKZ82TERR4Z",
"status": "investigating",
"updated_at": "2026-07-22T14:42:31Z"
}
],
"name": "InField — UI License Key Notice(INC-3158)",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-07-22T15:17:14Z",
"status": "resolved",
"updated_at": "2026-07-22T15:17:14Z"
},
{
"created_at": "2026-07-13T11:47:59Z",
"id": "01KXDMWCEWB0QBQQJ2VRZ8Q03P",
"impact": "minor",
"incident_updates": [
{
"body": "We identified and fixed an issue affecting Cognite Maintain for OMV Petrom Romania (Asset Oil) where users with edit access could not update activity dates, duplicate activities, or run deferment calculations in production. Read-only access was not affected. The fix has been applied and the service is operating normally.\n**Root cause:** Expired application client secrets on the backend authentication path caused write operations to return `401 Unauthorized` in production only.\n\n**Resolution:** New client secrets were generated and applied. The service is operating normally as of Jul 15, 2026, 1:10pm UTC.",
"created_at": "2026-07-20T14:24:42Z",
"display_at": "2026-07-20T14:24:42Z",
"id": "01KXZYMCFXVY8XH3AATPP68YF8",
"incident_id": "01KXDMWCEWB0QBQQJ2VRZ8Q03P",
"status": "resolved",
"updated_at": "2026-07-20T14:24:42Z"
},
{
"body": "We are currently investigating an issue affecting Cognite Maintain for OMV Romania Asset Oil in Production. Users with edit rights are unable to update an activity date, create duplicates, or run deferment calculations. Browsing and read actions continue to work normally. All other customers and clusters are unaffected. Our engineering team is actively investigating the root cause and will provide further updates as soon as possible",
"created_at": "2026-07-13T11:47:59Z",
"display_at": "2026-07-13T11:47:59Z",
"id": "01KXDMWCEWRE4RR3JDHX7VQMWY",
"incident_id": "01KXDMWCEWB0QBQQJ2VRZ8Q03P",
"status": "investigating",
"updated_at": "2026-07-13T11:47:59Z"
}
],
"name": "Cognite Maintain – Activity Edit, Duplicate and Deferment Calculation Failing",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-07-20T14:24:42Z",
"status": "resolved",
"updated_at": "2026-07-20T14:24:42Z"
},
{
"created_at": "2026-07-08T15:53:37Z",
"id": "01KX16YJJMXCF651K826K54SNJ",
"impact": "minor",
"incident_updates": [
{
"body": "The MUI X Missing license key warning is no longer showing in production, and Search/Data Exploration is behaving normally.",
"created_at": "2026-07-09T15:43:36Z",
"display_at": "2026-07-09T15:43:36Z",
"id": "01KX3RRYQEXFPPPR09XM06G26R",
"incident_id": "01KX16YJJMXCF651K826K54SNJ",
"status": "resolved",
"updated_at": "2026-07-09T15:43:36Z"
},
{
"body": "The team has deployed the revert for the change linked to the MUI X Missing license key overlay, with that the incident is now contained where we will be monitoring the applied changes to see the stability of the fix.",
"created_at": "2026-07-08T17:31:26Z",
"display_at": "2026-07-08T17:31:26Z",
"id": "01KX1CHNJBDDSJ7Z776HN9SECR",
"incident_id": "01KX16YJJMXCF651K826K54SNJ",
"status": "monitoring",
"updated_at": "2026-07-08T17:31:26Z"
},
{
"body": "The team has identified the issue and a fix is currently on the way to contain the issue.",
"created_at": "2026-07-08T16:15:09Z",
"display_at": "2026-07-08T16:15:09Z",
"id": "01KX185ZR46TBSB4W3NSSZ4XEM",
"incident_id": "01KX16YJJMXCF651K826K54SNJ",
"status": "identified",
"updated_at": "2026-07-08T16:15:09Z"
},
{
"body": "We have identified a UI interuption affecting a subset of projects in the cluster that is related to a MUI X Missing license key overlay in the Search UI. The team is currently investigating and a fix is on the way.",
"created_at": "2026-07-08T15:53:37Z",
"display_at": "2026-07-08T15:53:37Z",
"id": "01KX16YJJM3TFK13SHBJCXF073",
"incident_id": "01KX16YJJMXCF651K826K54SNJ",
"status": "investigating",
"updated_at": "2026-07-08T15:53:37Z"
}
],
"name": "Missing key warning error in search UI",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-07-09T15:43:36Z",
"status": "resolved",
"updated_at": "2026-07-09T15:43:36Z"
},
{
"created_at": "2026-06-25T11:10:48Z",
"id": "01KVZ7KC5F40BNH29KAQX1KET0",
"impact": "minor",
"incident_updates": [
{
"body": "The engineering team has confirmed that the Custom Apps page is fully operational for all users following the fix. The incident has been resolved and is now closed.",
"created_at": "2026-06-25T16:25:21Z",
"display_at": "2026-06-25T16:25:21Z",
"id": "01KVZSKB50EQN7GKC1HVRCR83V",
"incident_id": "01KVZ7KC5F40BNH29KAQX1KET0",
"status": "resolved",
"updated_at": "2026-06-25T16:25:21Z"
},
{
"body": "The engineering team has identified the root cause and deployed a fix. The Custom Apps view has been confirmed to be functioning normally. The service is being monitored to ensure continued stability.",
"created_at": "2026-06-25T11:14:57Z",
"display_at": "2026-06-25T11:14:57Z",
"id": "01KVZ7TZ9BYW7NW3MPZQM6NDJJ",
"incident_id": "01KVZ7KC5F40BNH29KAQX1KET0",
"status": "monitoring",
"updated_at": "2026-06-25T11:14:57Z"
},
{
"body": "An issue has been identified affecting Custom Apps for some users, where the page may display **\"No apps available\"** despite apps being present. The root cause is currently under investigation, and updates will be provided as more information becomes available.",
"created_at": "2026-06-25T11:10:48Z",
"display_at": "2026-06-25T10:51:00Z",
"id": "01KVZ7KC5FVW4Y05RD3HFB1BDS",
"incident_id": "01KVZ7KC5F40BNH29KAQX1KET0",
"status": "investigating",
"updated_at": "2026-06-25T11:11:33Z"
}
],
"name": "Degraded Experience for Custom Apps",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-06-25T16:25:21Z",
"status": "resolved",
"updated_at": "2026-06-25T16:25:21Z"
},
{
"created_at": "2026-06-04T18:34:07Z",
"id": "01KT9YM0EAPH2CV7CTKNQ93Y87",
"impact": "minor",
"incident_updates": [
{
"body": "The issue has now been fully contained and marked as resolved.",
"created_at": "2026-06-18T10:19:11Z",
"display_at": "2026-06-18T10:19:11Z",
"id": "01KVD3VTD8ZW31Z59KHXFKG50Q",
"incident_id": "01KT9YM0EAPH2CV7CTKNQ93Y87",
"status": "resolved",
"updated_at": "2026-06-18T10:19:11Z"
},
{
"body": "The fix has been applied and indexing has resumed. We are currently monitoring the state of the fix at this stage.",
"created_at": "2026-06-11T11:55:24Z",
"display_at": "2026-06-11T11:55:24Z",
"id": "01KTV8JZWZDNPAEDV7XBX7JH0Z",
"incident_id": "01KT9YM0EAPH2CV7CTKNQ93Y87",
"status": "monitoring",
"updated_at": "2026-06-11T11:55:24Z"
},
{
"body": "**The fix for the Elasticsearch mapping bug has been deployed to staging and initial results look good. Pending successful staging validation, rollout to production is planned throughout next week. The workaround remains to avoid using JSON keys with shared prefixes and dotted variants.**",
"created_at": "2026-06-05T13:02:55Z",
"display_at": "2026-06-05T13:02:55Z",
"id": "01KTBY29YXPHT5SN8PKFNCAZ87",
"incident_id": "01KT9YM0EAPH2CV7CTKNQ93Y87",
"status": "identified",
"updated_at": "2026-06-05T13:02:55Z"
},
{
"body": "A bug in the Elasticsearch mapping refresh is causing search and aggregate results to not update for projects using JSON properties where a plain key and a dotted key share a prefix (e.g. `\"min\"` and `\"min.A\"` in the same JSON object). A first fix was insufficient; a second fix (PR #1908) is now in pipeline and rollout to `az-eastus-1` is expected within ~1 week, pending testing.\n\n**Example of triggering data pattern:**\n\n```\n{\n \"min\": \"value\",\n \"min.A\": \"nested_value\"\n}\n```\n\n\nAny JSON property where a plain key (`min`) and a dotted key (`min.A`) share the same prefix will trigger the naming collision.\nCurrently fixing is in progress.\n",
"created_at": "2026-06-05T07:29:41Z",
"display_at": "2026-06-05T07:29:41Z",
"id": "01KTBB0455Z2E5N9YR4GJRCEYG",
"incident_id": "01KT9YM0EAPH2CV7CTKNQ93Y87",
"status": "identified",
"updated_at": "2026-06-05T07:29:41Z"
},
{
"body": "**We have identified the cause of an issue affecting Data Modeling search and aggregate results on the az-eastus-1 cluster. This issue is limited to a specific data pattern and does not affect all customers only projects ingesting JSON data where one key is used as a prefix before a dot (.) in another key within the same JSON object are impacted.**\n\n**Workaround: To avoid triggering this issue, please ensure JSON objects do not contain keys where one key shares a prefix with another key before a dot (.) or end of string. For example, a structure where one key appears as a prefix of another dot-separated key should be avoided.**\n\n**Our engineering team is actively working on a permanent fix. We will provide further updates as the situation develops.**",
"created_at": "2026-06-04T19:10:16Z",
"display_at": "2026-06-04T19:10:16Z",
"id": "01KTA0P6VKER4NJY45BXK5XQ83",
"incident_id": "01KT9YM0EAPH2CV7CTKNQ93Y87",
"status": "identified",
"updated_at": "2026-06-04T19:10:16Z"
},
{
"body": "We are currently investigating an issue where Data Modeling search and aggregate results are not refreshing as expected on the az-eastus-1 cluster. Users may experience incorrect or zero counts in search results, and newly ingested data may not appear in the CDF Search UI. Our engineering team has identified the root cause and a fix is currently in progress. We will provide further updates as the situation develops.",
"created_at": "2026-06-04T18:34:07Z",
"display_at": "2026-06-04T18:34:07Z",
"id": "01KT9YM0EAYK1YGZ1YD8CFEJ2P",
"incident_id": "01KT9YM0EAPH2CV7CTKNQ93Y87",
"status": "investigating",
"updated_at": "2026-06-04T18:34:07Z"
}
],
"name": "Data Modeling Search/Aggregates Not Refreshing",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-06-18T10:19:11Z",
"status": "resolved",
"updated_at": "2026-06-18T10:19:11Z"
},
{
"created_at": "2026-06-09T12:30:09Z",
"id": "01KTP5S573C928Y3S1BS8R4DD5",
"impact": "minor",
"incident_updates": [
{
"body": "File upload propagation delays (`isUploaded` flag) in the `az-eastus-1` cluster have been fully resolved. The root cause was an Azure Event Grid outage in the East US region (05:00 UTC June 9 – 06:30 UTC June 10, 2026), which prevented blob storage notifications from reaching our internal file processing pipeline. Propagation lag has now stabilized at normal, near-zero levels. All affected workflows, including CDF Function deployments and file-dependent processes, should be fully operational. We will continue to monitor system metrics to ensure stability.",
"created_at": "2026-06-11T06:59:41Z",
"display_at": "2026-06-11T06:59:41Z",
"id": "01KTTQNGCP77409DYM2MG2QHC3",
"incident_id": "01KTP5S573C928Y3S1BS8R4DD5",
"status": "resolved",
"updated_at": "2026-06-11T06:59:41Z"
},
{
"body": "We have identified a significant delay in the propagation of the `isUploaded` flag for files uploaded in the az-eastus-1 cluster. Files may take over 2 hours to be marked as uploaded after a successful upload.\n\nRoot cause identified: At least one project in the az-eastus-1 cluster is being rate limited by the DMS (Data Modeling Service), which is causing the file storage subscriber to back up and delay processing of upload notifications.\n\nOur engineering team has applied a fix to increase the DMS write concurrency limit for the affected project, and the change has been deployed. We are actively monitoring the situation to confirm recovery.\n\nWe will provide further updates as the situation progresses.",
"created_at": "2026-06-09T12:30:09Z",
"display_at": "2026-06-09T12:30:09Z",
"id": "01KTP5S573G4WQYCK8S8BBT8Y9",
"incident_id": "01KTP5S573C928Y3S1BS8R4DD5",
"status": "identified",
"updated_at": "2026-06-09T12:30:09Z"
}
],
"name": "files_is_uploaded_propagated_lag",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-06-11T06:59:41Z",
"status": "resolved",
"updated_at": "2026-06-11T06:59:41Z"
},
{
"created_at": "2026-06-04T14:52:05Z",
"id": "01KT9HXEGM0KFBM7CWZVD37PEF",
"impact": "minor",
"incident_updates": [
{
"body": "Data Workflow availability has been fully restored in cluster az-tyo-gp-001. Validation and monitoring confirm normal operation, and the incident has been resolved and closed.",
"created_at": "2026-06-04T14:53:48Z",
"display_at": "2026-06-04T14:20:00Z",
"id": "01KT9J0KDDPKQHMJDCVBSB5XM5",
"incident_id": "01KT9HXEGM0KFBM7CWZVD37PEF",
"status": "resolved",
"updated_at": "2026-06-04T14:56:47Z"
},
{
"body": "Data Workflow is unavailable in cluster **az-tyo-gp-001**, impacting workflow execution and related operations. Investigation and recovery efforts are ongoing.",
"created_at": "2026-06-04T14:52:05Z",
"display_at": "2026-06-04T13:30:00Z",
"id": "01KT9HXEGMZ5RH03JNC3FSHX14",
"incident_id": "01KT9HXEGM0KFBM7CWZVD37PEF",
"status": "investigating",
"updated_at": "2026-06-04T14:56:38Z"
}
],
"name": "Data Workflow Unavailable in Cluster az-tyo-gp-001",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-06-04T14:20:00Z",
"status": "resolved",
"updated_at": "2026-06-04T14:53:48Z"
},
{
"created_at": "2026-06-03T13:08:54Z",
"id": "01KT6SKSNATPN7A9WQACV4RR4P",
"impact": "major",
"incident_updates": [
{
"body": "The incident is now resolved as the fix was deployed to mitigate the UI rendering issue in the agent chat.",
"created_at": "2026-06-03T14:52:47Z",
"display_at": "2026-06-03T14:52:47Z",
"id": "01KT6ZJ0ZNMCFEA74TP9SPGEZ9",
"incident_id": "01KT6SKSNATPN7A9WQACV4RR4P",
"status": "resolved",
"updated_at": "2026-06-03T14:52:47Z"
},
{
"body": "We have identified a potential issue surrounding AI Chat agent, where the agent selector dropdown is rendering with heavily overlapped/duplicated text, making it effectively unusable. The issue is a frontend rendering regression, internally detected at 14:35 CEST today.",
"created_at": "2026-06-03T13:13:02Z",
"display_at": "2026-06-03T13:13:02Z",
"id": "01KT6SVC2JH1YX9KAV5HGPD2VS",
"incident_id": "01KT6SKSNATPN7A9WQACV4RR4P",
"status": "identified",
"updated_at": "2026-06-03T13:13:02Z"
},
{
"body": "We have identified a potential issue surrounding AI Chat agent, where the agent selector dropdown is rendering with heavily overlapped/duplicated text, making it effectively unusable. The issue is a frontend rendering regression, internally detected at 14:35 CEST today.",
"created_at": "2026-06-03T13:08:54Z",
"display_at": "2026-06-03T13:08:54Z",
"id": "01KT6SKSNA61QSFBYAWG69JJ6M",
"incident_id": "01KT6SKSNATPN7A9WQACV4RR4P",
"status": "investigating",
"updated_at": "2026-06-03T13:08:54Z"
}
],
"name": "Rendering issue identified in the Agent selector in Atlas AI",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-06-03T14:52:47Z",
"status": "resolved",
"updated_at": "2026-06-03T14:52:47Z"
},
{
"created_at": "2026-06-01T07:35:27Z",
"id": "01KT11QT9SXEREF6NTH8HNP8S5",
"impact": "major",
"incident_updates": [
{
"body": "The team has identified the root cause of the issue and successfully deployed a fix. As a result, the popup is no longer appearing, and normal functionality has been restored.\n\nThis incident has been fully resolved.",
"created_at": "2026-06-01T08:08:21Z",
"display_at": "2026-06-01T08:08:21Z",
"id": "01KT13M1R787CXSX8K2V5ZZVMS",
"incident_id": "01KT11QT9SXEREF6NTH8HNP8S5",
"status": "resolved",
"updated_at": "2026-06-01T08:08:21Z"
},
{
"body": "We have identified that `Polyfill.io` is asking for credentials on page load for hub.cognite.com the team is still investigating this issue.",
"created_at": "2026-06-01T07:35:27Z",
"display_at": "2026-06-01T07:35:27Z",
"id": "01KT11QT9SKBZB2S56A5PMK2C3",
"incident_id": "01KT11QT9SXEREF6NTH8HNP8S5",
"status": "investigating",
"updated_at": "2026-06-01T07:35:27Z"
}
],
"name": "Unwanted popup appearing in Cognite hub page",
"page_id": "01H16FDGE21SXKCBJCZJVXEHSX",
"resolved_at": "2026-06-01T08:08:21Z",
"status": "resolved",
"updated_at": "2026-06-01T08:08:21Z"
}
],
"page": {
"id": "01H16FDGE21SXKCBJCZJVXEHSX",
"name": "Cognite Service",
"updated_at": "2026-09-30T09:47:24Z",
"url": "https://status.cognite.com/"
}
}