Incident Report: Data loss due to storage failure on the zeze node

Top / News / Incident Report: Data loss due to storage failure on the zeze node

Posted: October 8, 2026

We deeply regret to inform you that a critical outage occurred in our storage service (zeze.teracloud.jp, hereinafter referred to as zeze), resulting in the loss of customer data. 

We sincerely apologize for the significant inconvenience and concern this incident has caused our customers. Below, this outline will cover what happened, the technical background of the incident, its causes in the clearest way possible, and our future actions.

Target nodezeze node (zeze1 storage pool)
Affected usersUsers who had data stored on the above node
Date of incidentOctober 1, 2026 (JST) ~ Ongoing

This outage is limited to the zeze node. Customers using other nodes ad services are not affected and can use them as usual.

You can check if you are affected by the issue in the address bar of your file browser or via the WebDAV connection URL.

Further details of the situation can be found in the following notice page.

1. What occurred?

On October 1, 2026, one hard disc (hereinafter referred to as a “disc”) failed on the zeze node where customer data was stored, resulting in data corruption.

The storage system was configured with 11 discs in a ZFS RAIDZ3 configuration. RAIDZ3 uses triple parity (redundancy information for data recovery), allowing data to be recovered even if up to three discs fail simultaneously.

The failed disc was disconnected from the storage system, and the system continued operating with the remaining 10 discs. Up to this point, the situation was within the expected operating limits of RAIDZ3.

However, around the time of the failure, other discs connected to the same control card (HBA, the component that connects the discs to the server) as the failed disc also became unresponsive to commands. In addition, some data was found to be corrupted beyond recovery using RAIDZ3. RAIDZ3 cannot recover data when four or more locations within the same data set are corrupted simultaneously.

As of October 1, 61 files had become unreadable. Because some of these files were system files required to operate the service, the service became unavailable.

It is important to note that, at this point, all 10 remaining discs were reported as “DEGRADED.” This does not mean that all 10 dics had failed simultaneously. When ZFS attempts to read data that cannot be recovered, it cannot determine which disc contains the corrupted data, so it records errors against all discs. Once the number of these error records exceeded a certain threshold, all discs were classified as “DEGRADED.”

The following day, October 2, three additional discs connected to the same control card became unusable one after another. Including the initial failure, a total of four discs had failed, exceeding the three-disc failure tolerance of RAIDZ3 and rendering the entire storage system unreadable.

2. Progression of the Incident and Response

Date and Time(JST)Details
10/1 01:43A disc connected to the control card (HBA) stopped responding to commands.
10/1 01:45One disc failed and was disconnected from the storage system. Operation continued with the remaining 10 discs.
10/1 01:53The monitoring system detected an anomaly. Our engineers confirmed that one disc had failed. However, no data read errors had been detected at this point, and it was determined that the service had not been affected.
10/1 06:36Troubleshooting began. We found that system files required to operate the service were unreadable and confirmed that the service was experiencing an issue.
10/1 07:06The service environment was shut down to prevent further damage.
10/1 08:15We confirmed that transferring all data at once was not possible. We decided to prepare a new node and recover the data individually, based on what could be read.
10/1 08:15The notice “Notice of Outage in zeze” was published.
10/1 14:03Analysis of the logs indicated a high probability that the control card was faulty.
10/1 16:48〜19:06Work was carried out at the data center to address the control card issue and update its firmware.
10/1 20:35A follow-up update was added to the outage notice.
10/2Two additional discs stopped responding, bringing the total number of failed discs to three. This exhausted the three-disc fault tolerance provided by RAIDZ3, and the storage system entered a suspended state (SUSPENDED).
10/2 13:51We decided to provide customers using the zeze node with a new data area on a new node.
10/2 14:40A follow-up update was added to the outage notice.
10/2 17:17Replacement discs were installed at the data center.
10/2 19:21A fourth disc became unavailable, rendering the entire storage system unreadable (UNAVAIL).
10/2 21:03We determined that data recovery was virtually impossible and shifted to a response plan based on the assumption of total data loss.
10/3 00:16A follow-up update was added to the outage notice.
10/8This page was published, summarizing the technical background, cause, extent of the damage, and future response measures related to the incident.

3. Scope of impact

As of October 8, the storage system on the zeze node remains completely inaccessible, and the data stored there cannot be accessed.

We are proceeding on the assumption that all data stored on the zeze node has been lost.

4. Regarding the cause

The investigation into the cause of the incident is ongoing.

Below, we distinguish between facts confirmed from the logs and inferences drawn from those facts. This information may be updated as the investigation progresses.

Triggering factors

We believe that an internal failure in one disc triggered simultaneous instability in multiple discs connected to the same control card. When a disc becomes unresponsive, the control card performs recovery operations, such as aborting commands or resetting the disc. We believe these recovery operations affected the processing of other discs connected to the same control card.

Predisposing factors (underlying factors)

The following two factors are believed to have contributed to the incident.

The first is the arrangement of the disc connections. While RAIDZ3 can tolerate the simultaneous failure of up to three discs, six of the eleven discs were connected to a single control card. As a result, a malfunction in that card could affect more discs than RAIDZ3 can tolerate.

The second is the control card's firmware. We are investigating whether a flaw in the error recovery process of an older firmware version may have contributed to the spread of the problem.

As of this announcement, the cause of the failure of the three discs that became unusable on October 2 remains uncertain. The issue may have originated in either the control card or the discs themselves. However, we are prioritizing the control card as a potential cause because it was shared by multiple affected discs, and are considering measures to prevent similar incidents in the future.

5. Recurrence prevention measures

To prevent similar issues from recurring, the following steps will be taken:

Inspection of all nodes using the same type of control card

We are checking other nodes to see if they are using the same type of control card and firmware, and if so, we will perform emergency maintenance, including firmware updates and component replacements. This check is being carried out as a top priority to prevent further damage.

Review of disc connection configuration

We will review the disc connection configuration for each node to ensure that a malfunction in a single control card or cable does not simultaneously affect more than three discs, which is the capacity of RAIDZ3.

Transition to a Ceph-based architecture

The core issue with this failure was that redundancy was confined to a single device and a single pool, meaning that a complex failure within that system could result in total data loss. Even RAIDZ3's triple parity could not recover the data when the damaged areas overlapped.

While this configuration was the best option available when InfiniCLOUD's (formerly TeraCLOUD) service first launched, we plan to migrate our storage infrastructure to a virtualized infrastructure using Ceph and replicate data across multiple chassis. This will eliminate the dependence on a single device and a single storage pool.

6. Future actions

As mentioned above, we are proceeding with the assumption that all data has been lost. Please understand in advance that there is a high possibility that we will not be able to recover or retrieve your data.

We sincerely apologize again for causing this situation, especially considering our responsibility to handle your valuable data.

 

For Inquiries

For inquiries regarding this matter, please use the link below.