<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Artie Status - Incident history</title>
    <link>https://artie.instatus.com</link>
    <description>Artie</description>
    <pubDate>Mon, 27 Jul 2026 23:00:00 +0000</pubDate>
    
<item>
  <title>Infrastructure upgrade for Artie Dashboard</title>
  <description>
    Type: Maintenance
    Duration: 2 hours

    Affected Components: Artie Dashboard
    Jul 22, 22:56:39 GMT+0 - Identified - We are planning for a scheduled maintenance during that time. Jul 27, 23:00:01 GMT+0 - Identified - Maintenance is now in progress Jul 28, 01:00:00 GMT+0 - Completed - Maintenance has completed successfully 
  </description>
  <content:encoded>
    <![CDATA[<p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 2 hours</p>
    <p><strong>Affected Components:</strong> </p>
    &lt;p&gt;&lt;small&gt;Jul &lt;var data-var=&#039;date&#039;&gt; 22&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;22:56:39&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Identified&lt;/strong&gt; -
  We are planning for a scheduled maintenance during that time..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jul &lt;var data-var=&#039;date&#039;&gt; 27&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;23:00:01&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Identified&lt;/strong&gt; -
  Maintenance is now in progress.&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jul &lt;var data-var=&#039;date&#039;&gt; 28&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;01:00:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Completed&lt;/strong&gt; -
  Maintenance has completed successfully.&lt;/p&gt;
]]>
  </content:encoded>
  <pubDate>Mon, 27 Jul 2026 23:00:00 +0000</pubDate>
  <link>https://artie.instatus.com/maintenance/cmrwom8w604i80klll561fkrv</link>
  <guid>https://artie.instatus.com/maintenance/cmrwom8w604i80klll561fkrv</guid>
</item>

<item>
  <title>Spark cluster upgrade in us-east-1</title>
  <description>
    Type: Maintenance
    Duration: 1 hour

    Affected Components: AWS US-East-1 Data Plane
    Jun 26, 19:38:47 GMT+0 - Identified - This upgrade only involves pipelines with Iceberg destinations. No impact is expected. Jun 29, 22:00:00 GMT+0 - Completed - Maintenance has completed successfully. Jun 29, 21:00:01 GMT+0 - Identified - Maintenance is now in progress 
  </description>
  <content:encoded>
    <![CDATA[<p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 1 hour</p>
    <p><strong>Affected Components:</strong> </p>
    &lt;p&gt;&lt;small&gt;Jun &lt;var data-var=&#039;date&#039;&gt; 26&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;19:38:47&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Identified&lt;/strong&gt; -
  This upgrade only involves pipelines with Iceberg destinations. No impact is expected..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jun &lt;var data-var=&#039;date&#039;&gt; 29&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;22:00:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Completed&lt;/strong&gt; -
  Maintenance has completed successfully..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jun &lt;var data-var=&#039;date&#039;&gt; 29&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;21:00:01&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Identified&lt;/strong&gt; -
  Maintenance is now in progress.&lt;/p&gt;
]]>
  </content:encoded>
  <pubDate>Mon, 29 Jun 2026 21:00:00 +0000</pubDate>
  <link>https://artie.instatus.com/maintenance/cmqvc3ndj05xf2poapcgh0mtg</link>
  <guid>https://artie.instatus.com/maintenance/cmqvc3ndj05xf2poapcgh0mtg</guid>
</item>

<item>
  <title>Scheduled database maintenance</title>
  <description>
    Type: Maintenance
    

    Affected Components: Artie Dashboard
    Jun 26, 00:27:12 GMT+0 - Completed - We&#039;re performing planned maintenance to improve the reliability of the Artie Dashboard. During this window the dashboard may be briefly unavailable or slower to load. Your data pipelines are not affected and will continue running normally. No data is impacted. 
  </description>
  <content:encoded>
    <![CDATA[<p><strong>Type:</strong> Maintenance</p>
    
    <p><strong>Affected Components:</strong> </p>
    &lt;p&gt;&lt;small&gt;Jun &lt;var data-var=&#039;date&#039;&gt; 26&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;00:27:12&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Completed&lt;/strong&gt; -
  We&#039;re performing planned maintenance to improve the reliability of the Artie Dashboard. During this window the dashboard may be briefly unavailable or slower to load. Your data pipelines are not affected and will continue running normally. No data is impacted..&lt;/p&gt;
]]>
  </content:encoded>
  <pubDate>Fri, 26 Jun 2026 00:27:12 +0000</pubDate>
  <link>https://artie.instatus.com/maintenance/cmqu6yoro018p2jqpct37b7nu</link>
  <guid>https://artie.instatus.com/maintenance/cmqu6yoro018p2jqpct37b7nu</guid>
</item>

<item>
  <title>Networking upgrade for Artie Dashboard</title>
  <description>
    Type: Maintenance
    Duration: 30 minutes

    Affected Components: Artie Dashboard
    Jun 22, 18:19:51 GMT+0 - Identified -  Jun 25, 00:00:01 GMT+0 - Identified - Maintenance is now in progress Jun 25, 00:30:00 GMT+0 - Completed - Maintenance has completed successfully 
  </description>
  <content:encoded>
    <![CDATA[<p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 30 minutes</p>
    <p><strong>Affected Components:</strong> </p>
    &lt;p&gt;&lt;small&gt;Jun &lt;var data-var=&#039;date&#039;&gt; 22&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;18:19:51&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Identified&lt;/strong&gt; -
  .&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jun &lt;var data-var=&#039;date&#039;&gt; 25&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;00:00:01&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Identified&lt;/strong&gt; -
  Maintenance is now in progress.&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jun &lt;var data-var=&#039;date&#039;&gt; 25&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;00:30:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Completed&lt;/strong&gt; -
  Maintenance has completed successfully.&lt;/p&gt;
]]>
  </content:encoded>
  <pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate>
  <link>https://artie.instatus.com/maintenance/cmqpjiq5j00is2kpijbyhooc7</link>
  <guid>https://artie.instatus.com/maintenance/cmqpjiq5j00is2kpijbyhooc7</guid>
</item>

<item>
  <title>Database upgrade</title>
  <description>
    Type: Maintenance
    Duration: 10 hours

    Affected Components: Artie Dashboard
    May 19, 01:14:00 GMT+0 - Completed - Database upgrade is complete. May 19, 01:04:00 GMT+0 - Identified - We are running scheduled maintenance to upgrade our database. 
  </description>
  <content:encoded>
    <![CDATA[<p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 10 hours</p>
    <p><strong>Affected Components:</strong> </p>
    &lt;p&gt;&lt;small&gt;May &lt;var data-var=&#039;date&#039;&gt; 19&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;01:14:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Completed&lt;/strong&gt; -
  Database upgrade is complete..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;May &lt;var data-var=&#039;date&#039;&gt; 19&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;01:04:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Identified&lt;/strong&gt; -
  We are running scheduled maintenance to upgrade our database..&lt;/p&gt;
]]>
  </content:encoded>
  <pubDate>Tue, 19 May 2026 01:04:00 +0000</pubDate>
  <link>https://artie.instatus.com/maintenance/cmpbxuc7s000zqull0ehw6g83</link>
  <guid>https://artie.instatus.com/maintenance/cmpbxuc7s000zqull0ehw6g83</guid>
</item>

<item>
  <title>Dashboard API timing out</title>
  <description>
    Type: Incident
    Duration: 1 hour

    Affected Components: Artie Dashboard
    Apr 7, 18:00:00 GMT+0 - Investigating - We are currently investigating this incident. Apr 7, 19:00:27 GMT+0 - Resolved - This incident has been resolved. We will prepare a post-mortem for this incident shortly. 
  </description>
  <content:encoded>
    <![CDATA[<p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 1 hour</p>
    <p><strong>Affected Components:</strong> </p>
    &lt;p&gt;&lt;small&gt;Apr &lt;var data-var=&#039;date&#039;&gt; 7&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;18:00:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Investigating&lt;/strong&gt; -
  We are currently investigating this incident..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Apr &lt;var data-var=&#039;date&#039;&gt; 7&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;19:00:27&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Resolved&lt;/strong&gt; -
  This incident has been resolved. We will prepare a post-mortem for this incident shortly..&lt;/p&gt;
]]>
  </content:encoded>
  <pubDate>Tue, 7 Apr 2026 18:00:00 +0000</pubDate>
  <link>https://artie.instatus.com/incident/cmnoy1ndj054az9eh3mdnab88</link>
  <guid>https://artie.instatus.com/incident/cmnoy1ndj054az9eh3mdnab88</guid>
</item>

<item>
  <title>NGINX Upgrade</title>
  <description>
    Type: Maintenance
    Duration: 1 hour

    Affected Components: Artie Dashboard
    Mar 25, 23:00:01 GMT+0 - Identified - Maintenance is now in progress Mar 26, 00:00:00 GMT+0 - Completed - Maintenance has completed successfully Mar 25, 23:00:00 GMT+0 - Identified - We are planning to do some maintenance upgrades to our Artie dashboard. Our dashboard and dashboard API will be temporarily degraded during this period. 
  </description>
  <content:encoded>
    <![CDATA[<p><strong>Type:</strong> Maintenance</p>
    <p><strong>Duration:</strong> 1 hour</p>
    <p><strong>Affected Components:</strong> </p>
    &lt;p&gt;&lt;small&gt;Mar &lt;var data-var=&#039;date&#039;&gt; 25&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;23:00:01&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Identified&lt;/strong&gt; -
  Maintenance is now in progress.&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Mar &lt;var data-var=&#039;date&#039;&gt; 26&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;00:00:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Completed&lt;/strong&gt; -
  Maintenance has completed successfully.&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Mar &lt;var data-var=&#039;date&#039;&gt; 25&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;23:00:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Identified&lt;/strong&gt; -
  We are planning to do some maintenance upgrades to our Artie dashboard. Our dashboard and dashboard API will be temporarily degraded during this period..&lt;/p&gt;
]]>
  </content:encoded>
  <pubDate>Wed, 25 Mar 2026 23:00:00 +0000</pubDate>
  <link>https://artie.instatus.com/maintenance/cmn2frzdp0urifbud6dqtlw6c</link>
  <guid>https://artie.instatus.com/maintenance/cmn2frzdp0urifbud6dqtlw6c</guid>
</item>

<item>
  <title>Dashboard API is down</title>
  <description>
    Type: Incident
    Duration: 51 minutes

    Affected Components: Artie Dashboard
    Feb 20, 02:10:00 GMT+0 - Resolved - Root cause: our NGINX controller became unresponsive, which caused the Dashboard API to be unavailable. As a result, customers were unable to log into the UI or make pipeline changes.

Service has now been fully restored and systems are operating normally.

We’ll continue to monitor closely to ensure stability. Please reach out to our team if you experience any lingering issues. Feb 20, 01:22:00 GMT+0 - Investigating - We are currently investigating this incident. 
  </description>
  <content:encoded>
    <![CDATA[<p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 51 minutes</p>
    <p><strong>Affected Components:</strong> </p>
    &lt;p&gt;&lt;small&gt;Feb &lt;var data-var=&#039;date&#039;&gt; 20&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;02:10:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Resolved&lt;/strong&gt; -
  Root cause: our NGINX controller became unresponsive, which caused the Dashboard API to be unavailable. As a result, customers were unable to log into the UI or make pipeline changes.

Service has now been fully restored and systems are operating normally.

We’ll continue to monitor closely to ensure stability. Please reach out to our team if you experience any lingering issues..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Feb &lt;var data-var=&#039;date&#039;&gt; 20&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;01:22:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Investigating&lt;/strong&gt; -
  We are currently investigating this incident..&lt;/p&gt;
]]>
  </content:encoded>
  <pubDate>Fri, 20 Feb 2026 01:22:00 +0000</pubDate>
  <link>https://artie.instatus.com/incident/cmlu8y5u20a5cwndff1ziqgpg</link>
  <guid>https://artie.instatus.com/incident/cmlu8y5u20a5cwndff1ziqgpg</guid>
</item>

<item>
  <title>Networking issues in use1</title>
  <description>
    Type: Incident
    Duration: 7 hours and 54 minutes

    Affected Components: Artie Dashboard, AWS US-East-1 Data Plane
    Jan 31, 00:30:00 GMT+0 - Identified - A config change has affected networking for pipelines in the us-east-1 data plane. We are continuing to work on a fix for this incident. Jan 31, 03:26:57 GMT+0 - Monitoring - There&#039;s a CNI error. We are seeing some clients recover and are still working towards a full recovery. Jan 31, 04:39:17 GMT+0 - Monitoring - We’ve identified intermittent DNS issues on a specific node and are actively working to resolve them. We’re working to restore full reliability. Jan 31, 05:14:24 GMT+0 - Monitoring - We are now working with AWS to diagnose the partial DNS failure. Jan 31, 07:42:31 GMT+0 - Monitoring - We are continuing to see intermittent DNS errors. The issue has been escalated with AWS, and an AWS engineer is joining the investigation shortly. We’ll share another update as we learn more. Jan 31, 08:24:08 GMT+0 - Resolved - We’ve identified the root cause of the DNS instability: an out-of-order VPC CNI upgrade led to an undocumented version mismatch, which caused conflicts in our DNS routing rules. The issue has now been resolved, and we are closely monitoring all pipelines to ensure continued stability. Feb 10, 01:34:05 GMT+0 - Postmortem - ### Incident Overview

* **Date:** January 30, 2025
* **Affected AWS region:** us-east-1
* **Duration:** partial outage 6.5 hours, total outage 1 hour

### Summary

On January 30, a routine infrastructure upgrade to our Kubernetes networking layer caused a service disruption in one of our primary cloud regions. What began as a partial outage escalated into a full outage affecting data pipelines for customers in the us-east-1 region. We sincerely apologize for the impact this had on your operations.

Service was fully restored the same evening after our team identified and resolved the underlying issue. It was a version mismatch between critical networking components in our cluster.

### Impact

We understand how important reliability is to your business, and we are deeply sorry for the disruption. During the outage:

* **Service was unavailable** in the us-east-1 region for 7.5 hours.
* A small number of affected customers experienced **data replication interruptions**, which our team worked to remediate promptly.

We recognize that any downtime is unacceptable, and we take full responsibility for this incident.

### What happened

Our engineering team was performing a planned upgrade to a networking add-on (vpc-cni) to increase capacity in our Kubernetes infrastructure. During the upgrade:

1. The change was applied successfully in a non-production environment but encountered compatibility issues in the production cluster.
2. A rollback was not possible because the previous component version was no longer supported by the underlying platform.
3. The issue was compounded by version mismatches across several interdependent networking components, which made diagnosis more complex than expected.

Our team identified the root cause as an outdated networking component (kube-proxy) that was incompatible with the rest of the upgraded stack. Once the fix was applied, service was restored.

### What we&#039;re sorry about

We want to be transparent; this incident was avoidable. We moved forward with a change that, in hindsight, required more thorough validation and coordination. Specifically:

* The production environment differed from the test environment in ways we did not fully account for.
* We did not have sufficient safeguards in place to catch the version incompatibility before it caused an outage.
* Our incident response process was slower than it should have been, and we did not escalate to our cloud provider&#039;s engineering team quickly enough.

We are sorry for the disruption and for falling short of the reliability you expect from us.

### Our plan going forward

We are actively carrying out a comprehensive set of improvements to prevent incidents like this from happening again:

**Stronger change controls**

* We have strengthened our change management process so that all planned infrastructure migrations require formal approval and a minimum of two engineers working together to apply changes.
* Non-reversible changes are restricted to business hours (Monday through Thursday) to ensure full team availability and vendor support coverage.

**Improved testing and validation**

* We are auditing and standardizing component versions across all of our clusters to eliminate hidden drift between environments.
* Production and non-production clusters are being aligned to ensure test results are representative.

**Faster incident response**

* We are formalizing a war room procedure with clearly defined roles: a primary engineer making changes, a secondary approving them, an incident commander coordinating, and a communications lead keeping customers informed.
* We are establishing clear severity definitions and escalation timelines so that vendor support is engaged earlier when needed.

**Infrastructure resilience**

* We are working toward running multiple clusters per region so that workloads can be migrated if a single cluster is compromised.
* We are building the capability to move customer pipelines between clusters with minimal disruption.
* We are proactively engaging our cloud provider to review and approve infrastructure upgrade plans before they are executed.

---

We value your trust and are committed to earning it back through these concrete actions. If you have any questions or concerns about this incident or our remediation plan, please do not hesitate to reach out to your account contact or our support team. 
  </description>
  <content:encoded>
    <![CDATA[<p><strong>Type:</strong> Incident</p>
    <p><strong>Duration:</strong> 7 hours and 54 minutes</p>
    <p><strong>Affected Components:</strong> , </p>
    &lt;p&gt;&lt;small&gt;Jan &lt;var data-var=&#039;date&#039;&gt; 31&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;00:30:00&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Identified&lt;/strong&gt; -
  A config change has affected networking for pipelines in the us-east-1 data plane. We are continuing to work on a fix for this incident..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jan &lt;var data-var=&#039;date&#039;&gt; 31&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;03:26:57&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Monitoring&lt;/strong&gt; -
  There&#039;s a CNI error. We are seeing some clients recover and are still working towards a full recovery..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jan &lt;var data-var=&#039;date&#039;&gt; 31&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;04:39:17&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Monitoring&lt;/strong&gt; -
  We’ve identified intermittent DNS issues on a specific node and are actively working to resolve them. We’re working to restore full reliability..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jan &lt;var data-var=&#039;date&#039;&gt; 31&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;05:14:24&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Monitoring&lt;/strong&gt; -
  We are now working with AWS to diagnose the partial DNS failure..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jan &lt;var data-var=&#039;date&#039;&gt; 31&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;07:42:31&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Monitoring&lt;/strong&gt; -
  We are continuing to see intermittent DNS errors. The issue has been escalated with AWS, and an AWS engineer is joining the investigation shortly. We’ll share another update as we learn more..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jan &lt;var data-var=&#039;date&#039;&gt; 31&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;08:24:08&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Resolved&lt;/strong&gt; -
  We’ve identified the root cause of the DNS instability: an out-of-order VPC CNI upgrade led to an undocumented version mismatch, which caused conflicts in our DNS routing rules. The issue has now been resolved, and we are closely monitoring all pipelines to ensure continued stability..&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Feb &lt;var data-var=&#039;date&#039;&gt; 10&lt;/var&gt;, &lt;var data-var=&#039;time&#039;&gt;01:34:05&lt;/var&gt; GMT+0&lt;/small&gt;&lt;br&gt;&lt;strong&gt;Postmortem&lt;/strong&gt; -
  ### Incident Overview

* **Date:** January 30, 2025
* **Affected AWS region:** us-east-1
* **Duration:** partial outage 6.5 hours, total outage 1 hour

### Summary

On January 30, a routine infrastructure upgrade to our Kubernetes networking layer caused a service disruption in one of our primary cloud regions. What began as a partial outage escalated into a full outage affecting data pipelines for customers in the us-east-1 region. We sincerely apologize for the impact this had on your operations.

Service was fully restored the same evening after our team identified and resolved the underlying issue. It was a version mismatch between critical networking components in our cluster.

### Impact

We understand how important reliability is to your business, and we are deeply sorry for the disruption. During the outage:

* **Service was unavailable** in the us-east-1 region for 7.5 hours.
* A small number of affected customers experienced **data replication interruptions**, which our team worked to remediate promptly.

We recognize that any downtime is unacceptable, and we take full responsibility for this incident.

### What happened

Our engineering team was performing a planned upgrade to a networking add-on (vpc-cni) to increase capacity in our Kubernetes infrastructure. During the upgrade:

1. The change was applied successfully in a non-production environment but encountered compatibility issues in the production cluster.
2. A rollback was not possible because the previous component version was no longer supported by the underlying platform.
3. The issue was compounded by version mismatches across several interdependent networking components, which made diagnosis more complex than expected.

Our team identified the root cause as an outdated networking component (kube-proxy) that was incompatible with the rest of the upgraded stack. Once the fix was applied, service was restored.

### What we&#039;re sorry about

We want to be transparent; this incident was avoidable. We moved forward with a change that, in hindsight, required more thorough validation and coordination. Specifically:

* The production environment differed from the test environment in ways we did not fully account for.
* We did not have sufficient safeguards in place to catch the version incompatibility before it caused an outage.
* Our incident response process was slower than it should have been, and we did not escalate to our cloud provider&#039;s engineering team quickly enough.

We are sorry for the disruption and for falling short of the reliability you expect from us.

### Our plan going forward

We are actively carrying out a comprehensive set of improvements to prevent incidents like this from happening again:

**Stronger change controls**

* We have strengthened our change management process so that all planned infrastructure migrations require formal approval and a minimum of two engineers working together to apply changes.
* Non-reversible changes are restricted to business hours (Monday through Thursday) to ensure full team availability and vendor support coverage.

**Improved testing and validation**

* We are auditing and standardizing component versions across all of our clusters to eliminate hidden drift between environments.
* Production and non-production clusters are being aligned to ensure test results are representative.

**Faster incident response**

* We are formalizing a war room procedure with clearly defined roles: a primary engineer making changes, a secondary approving them, an incident commander coordinating, and a communications lead keeping customers informed.
* We are establishing clear severity definitions and escalation timelines so that vendor support is engaged earlier when needed.

**Infrastructure resilience**

* We are working toward running multiple clusters per region so that workloads can be migrated if a single cluster is compromised.
* We are building the capability to move customer pipelines between clusters with minimal disruption.
* We are proactively engaging our cloud provider to review and approve infrastructure upgrade plans before they are executed.

---

We value your trust and are committed to earning it back through these concrete actions. If you have any questions or concerns about this incident or our remediation plan, please do not hesitate to reach out to your account contact or our support team..&lt;/p&gt;
]]>
  </content:encoded>
  <pubDate>Sat, 31 Jan 2026 00:30:00 +0000</pubDate>
  <link>https://artie.instatus.com/incident/cml1qnb9z01f9jyy58zozjs8q</link>
  <guid>https://artie.instatus.com/incident/cml1qnb9z01f9jyy58zozjs8q</guid>
</item>

  </channel>
  </rss>