Posted: October 8, 2026
We deeply regret to inform you that a critical outage occurred in our storage service (zeze.teracloud.jp, hereinafter referred to as zeze), resulting in the loss of customer data.
We sincerely apologize for the significant inconvenience and concern this incident has caused our customers. Below, this outline will cover what happened, the technical background of the incident, its causes in the clearest way possible, and our future actions.
| Target node | zeze node (zeze1 storage pool) |
|---|---|
| Affected users | Users who had data stored on the above node |
| Date of incident | October 1, 2026 (JST) ~ Ongoing |
This outage is limited to the zeze node. Customers using other nodes ad services are not affected and can use them as usual.
You can check if you are affected by the issue in the address bar of your file browser or via the WebDAV connection URL.
Further details of the situation can be found in the following notice page.
1. What occurred?
On October 1, 2026, one hard disc (hereinafter referred to as a “disc”) failed on the zeze node where customer data was stored, resulting in data corruption.
The storage system was configured with 11 discs in a ZFS RAIDZ3 configuration. RAIDZ3 uses triple parity (redundancy information for data recovery), allowing data to be recovered even if up to three discs fail simultaneously.
The failed disc was disconnected from the storage system, and the system continued operating with the remaining 10 discs. Up to this point, the situation was within the expected operating limits of RAIDZ3.
However, around the time of the failure, other discs connected to the same control card (HBA, the component that connects the discs to the server) as the failed disc also became unresponsive to commands. In addition, some data was found to be corrupted beyond recovery using RAIDZ3. RAIDZ3 cannot recover data when four or more locations within the same data set are corrupted simultaneously.
As of October 1, 61 files had become unreadable. Because some of these files were system files required to operate the service, the service became unavailable.
It is important to note that, at this point, all 10 remaining discs were reported as “DEGRADED.” This does not mean that all 10 dics had failed simultaneously. When ZFS attempts to read data that cannot be recovered, it cannot determine which disc contains the corrupted data, so it records errors against all discs. Once the number of these error records exceeded a certain threshold, all discs were classified as “DEGRADED.”
The following day, October 2, three additional discs connected to the same control card became unusable one after another. Including the initial failure, a total of four discs had failed, exceeding the three-disc failure tolerance of RAIDZ3 and rendering the entire storage system unreadable.
2. Progression of the Incident and Response
| Date and Time(JST) | Details |
|---|---|
| 10/1 01:43 | A disc connected to the control card (HBA) stopped responding to commands. |
| 10/1 01:45 | One disc failed and was disconnected from the storage system. Operation continued with the remaining 10 discs. |
| 10/1 01:53 | The monitoring system detected an anomaly. Our engineers confirmed that one disc had failed. However, no data read errors had been detected at this point, and it was determined that the service had not been affected. |
| 10/1 06:36 | Troubleshooting began. We found that system files required to operate the service were unreadable and confirmed that the service was experiencing an issue. |
| 10/1 07:06 | The service environment was shut down to prevent further damage. |
| 10/1 08:15 | We confirmed that transferring all data at once was not possible. We decided to prepare a new node and recover the data individually, based on what could be read. |
| 10/1 08:15 | The notice “Notice of Outage in zeze” was published. |
| 10/1 14:03 | Analysis of the logs indicated a high probability that the control card was faulty. |
| 10/1 16:48〜19:06 | Work was carried out at the data center to address the control card issue and update its firmware. |
| 10/1 20:35 | A follow-up update was added to the outage notice. |
| 10/2 | Two additional discs stopped responding, bringing the total number of failed discs to three. This exhausted the three-disc fault tolerance provided by RAIDZ3, and the storage system entered a suspended state (SUSPENDED). |
| 10/2 13:51 | We decided to provide customers using the zeze node with a new data area on a new node. |
| 10/2 14:40 | A follow-up update was added to the outage notice. |
| 10/2 17:17 | Replacement discs were installed at the data center. |
| 10/2 19:21 | A fourth disc became unavailable, rendering the entire storage system unreadable (UNAVAIL). |
| 10/2 21:03 | We determined that data recovery was virtually impossible and shifted to a response plan based on the assumption of total data loss. |
| 10/3 00:16 | A follow-up update was added to the outage notice. |
| 10/8 | This page was published, summarizing the technical background, cause, extent of the damage, and future response measures related to the incident. |
- SUSPENDED refers to a state in which the storage system has stopped reading and writing data. UNAVAIL refers to a state in which the entire storage system is inaccessible because the required number of discs is unavailable.
3. Scope of impact
As of October 8, the storage system on the zeze node remains completely inaccessible, and the data stored there cannot be accessed.
We are proceeding on the assumption that all data stored on the zeze node has been lost.
4. Regarding the cause
The investigation into the cause of the incident is ongoing.
Below, we distinguish between facts confirmed from the logs and inferences drawn from those facts. This information may be updated as the investigation progresses.
- The initial anomaly was an internal failure of one disc. The disc returned responses indicating an internal failure in response to every command.
- Immediately before the failure, processing delays were recorded on the control card connected to the disc. Other discs connected to the same control card also experienced issues, including failing to respond to commands or being unrecognised by the system.
- Six of the eleven discs in the storage system were connected to this single control card.
- No communication errors were recorded in the signal path (cables and connections) between the control card and the disc.
- The control card was running an older firmware version for which the manufacturer had released an update.
- The three discs that became unusable on October 2 were all connected to the same control card.
Triggering factors
We believe that an internal failure in one disc triggered simultaneous instability in multiple discs connected to the same control card. When a disc becomes unresponsive, the control card performs recovery operations, such as aborting commands or resetting the disc. We believe these recovery operations affected the processing of other discs connected to the same control card.
Predisposing factors (underlying factors)
The following two factors are believed to have contributed to the incident.
The first is the arrangement of the disc connections. While RAIDZ3 can tolerate the simultaneous failure of up to three discs, six of the eleven discs were connected to a single control card. As a result, a malfunction in that card could affect more discs than RAIDZ3 can tolerate.
The second is the control card's firmware. We are investigating whether a flaw in the error recovery process of an older firmware version may have contributed to the spread of the problem.
As of this announcement, the cause of the failure of the three discs that became unusable on October 2 remains uncertain. The issue may have originated in either the control card or the discs themselves. However, we are prioritizing the control card as a potential cause because it was shared by multiple affected discs, and are considering measures to prevent similar incidents in the future.
5. Recurrence prevention measures
To prevent similar issues from recurring, the following steps will be taken:
Inspection of all nodes using the same type of control card
We are checking other nodes to see if they are using the same type of control card and firmware, and if so, we will perform emergency maintenance, including firmware updates and component replacements. This check is being carried out as a top priority to prevent further damage.
Review of disc connection configuration
We will review the disc connection configuration for each node to ensure that a malfunction in a single control card or cable does not simultaneously affect more than three discs, which is the capacity of RAIDZ3.
Transition to a Ceph-based architecture
The core issue with this failure was that redundancy was confined to a single device and a single pool, meaning that a complex failure within that system could result in total data loss. Even RAIDZ3's triple parity could not recover the data when the damaged areas overlapped.
While this configuration was the best option available when InfiniCLOUD's (formerly TeraCLOUD) service first launched, we plan to migrate our storage infrastructure to a virtualized infrastructure using Ceph and replicate data across multiple chassis. This will eliminate the dependence on a single device and a single storage pool.
6. Future actions
- We will prepare a new data area on the new node and allocate it to customers using the zeze node. We will inform you of the start date for using the new data area as soon as it is decided.
- We will continue our attempts to get the unusable disc recognized again. However, the chances of success are extremely low.
- Data that can be read will be copied sequentially to the new data area.
As mentioned above, we are proceeding with the assumption that all data has been lost. Please understand in advance that there is a high possibility that we will not be able to recover or retrieve your data.
We sincerely apologize again for causing this situation, especially considering our responsibility to handle your valuable data.