the maintenance has completed. Services should be back online at this time. If your VM is still offline, please reach out to support@lunanode.com.
Progress
2 public update(s)
Remediation: the maintenance has completed. Services should be back online at this time. If your VM is still offline, please reach out to support@lunanode.com.
There will be up to 6 hours of network downtime (although much shorter downtime is expected) due to maintenance being performed by our datacenter (Cogent Toronto). They say "The purpose of this work is to apply and activate software module upgrades on code running on the edge routers per the manufacturer's recommendations."
Cause: There will be up to 6 hours of network downtime (although much shorter downtime is expected) due to maintenance being performed by our datacenter (Cogent Toronto). They say "The purpose of this work is to apply and activate software module upgrades on code running on the edge routers per the manufacturer's recommendati…
Cause: We are currently aware of a networking issue causing packet loss in our Toronto datacenter. Our upstream provider has identified the root cause and is actively working on a resolution. We will provide further updates as more i nformation becomes available.
Cause: We are currently aware of a networking issue causing packet loss in our Toronto datacenter. Our upstream provider has identified the root cause and is actively working on a resolution. We will provide further updates as more information becomes available.
There will be up to 3 hours of network downtime (although much shorter downtime is expected) due to maintenance being performed by our datacenter (Cogent Toronto). They say "The purpose of this work is to apply and activate software module upgrades on code running on the edge routers per the manufacturer's recommendations."
Cause: There will be up to 3 hours of network downtime (although much shorter downtime is expected) due to maintenance being performed by our datacenter (Cogent Toronto). They say "The purpose of this work is to apply and activate software module upgrades on code running on the edge routers per the manufacturer's recommendati…
the filesystem repair is complete and virtual machines have been restored. Data on four virtual machines was lost due to disk corruption. We will compensate affected users with three months of credit for downtime and twelve months of credit for data loss -- please open a ticket to request the credit be added to your account.
Progress
12 public update(s)
Data impact
Data loss confirmed
Compensation & remedy
Compensation or remedy offered
Cause: the filesystem repair is complete and virtual machines have been restored. Data on four virtual machines was lost due to disk corruption. We will compensate affected users with three months of credit for downtime and twelve months of credit for data loss -- please open a ticket to request the credit be added to your ac…
Remediation: the filesystem repair is complete and virtual machines have been restored. Data on four virtual machines was lost due to disk corruption. We will compensate affected users with three months of credit for downtime and twelve months of credit for data loss -- please open a ticket to request the credit be added to your ac…
services are back online at this time. The downtime occurred due to a hardware issue with one of the SSDs on the hypevisor; the SSD needed to be replaced but due to the failure mode (where it did not immediately fail) it also caused data corruption to the filesystem. The filesystem is now repaired but at least five virtual machines have corrupted disks. If y…
Progress
8 public update(s)
Data impact
Data loss reported
Compensation & remedy
Compensation or remedy offered
Cause: services are back online at this time. The downtime occurred due to a hardware issue with one of the SSDs on the hypevisor; the SSD needed to be replaced but due to the failure mode (where it did not immediately fail) it also caused data corruption to the filesystem. The filesystem is now repaired but at least five vir…
Remediation: services are back online at this time. The downtime occurred due to a hardware issue with one of the SSDs on the hypevisor; the SSD needed to be replaced but due to the failure mode (where it did not immediately fail) it also caused data corruption to the filesystem. The filesystem is now repaired but at least five vir…
We rebooted this hypervisor due to abnormal system errors. We will continue monitoring the stability of the hypervisor in case the reboot does not address the problems.
Cause: We rebooted this hypervisor due to abnormal system errors. We will continue monitoring the stability of the hypervisor in case the reboot does not address the problems.
Cause: There will be up to one hour of network downtime due to maintenance being performed by our datacenter (Cogent Toronto). They say "the purpose of this maintenance is Network Maintenance hardware and software upgrades reboot required."
Toronto service interrupting upstream network maintenance
4 Jul 2024, 13:00
Our upstream provider, Cogent, will be performing maintenance on network equipment. There is expected two outages of 15-30 minutes each as the maintenance is carried out.
We rebooted this hypervisor due to abnormal system errors. We will continue monitoring the stability of the hypervisor in case the reboot does not address the problems.
Cause: We rebooted this hypervisor due to abnormal system errors. We will continue monitoring the stability of the hypervisor in case the reboot does not address the problems.
Cause: NameSilo has re-activated lndyn.com. Services should be restored at this time. We have not received additional information about the original reason for the suspension yet, just notification that it was re-activated. We are in the process of migrating all of our domains off of NameSilo due to the 24 hours of downtime t…
Remediation: NameSilo has re-activated lndyn.com. Services should be restored at this time. We have not received additional information about the original reason for the suspension yet, just notification that it was re-activated. We are in the process of migrating all of our domains off of NameSilo due to the 24 hours of downtime t…
a prolonged outage occurred today due to incorrect identification of affected hypervisor while diagnosing the issue. Once we identified the correct hypervisor, the hypervisor was brought back online by replacing a faulty disk. Network connectivity to VMs that were using this hypervisor as network node was affected as well.
Progress
2 public update(s)
Cause: a prolonged outage occurred today due to incorrect identification of affected hypervisor while diagnosing the issue. Once we identified the correct hypervisor, the hypervisor was brought back online by replacing a faulty disk. Network connectivity to VMs that were using this hypervisor as network node was affected as w…
Remediation: a prolonged outage occurred today due to incorrect identification of affected hypervisor while diagnosing the issue. Once we identified the correct hypervisor, the hypervisor was brought back online by replacing a faulty disk. Network connectivity to VMs that were using this hypervisor as network node was affected as w…
we have heard from the datacenter that the power issue was due to breaker tripping incorrectly. Cogent says they have swapped the breaker with a new breaker and we should not experience any further issues with that unit.
Progress
5 public update(s)
Cause: we have heard from the datacenter that the power issue was due to breaker tripping incorrectly. Cogent says they have swapped the breaker with a new breaker and we should not experience any further issues with that unit.
Remediation: most services are restored at this time. Remaining services should be online shortly.
Cause: packet loss has come up again. It is due to DDoS attack targeting multiple IPs making it difficult to null route. We are working again to restore connectivity.
One Toronto hypervisor (a4120f5bedfa) offline for 20-30 minutes due to hardware failure (disk plane). Services are restored after swapping the disks to a spare physical server.
Cause: One Toronto hypervisor (a4120f5bedfa) offline for 20-30 minutes due to hardware failure (disk plane). Services are restored after swapping the disks to a spare physical server.
Remediation: One Toronto hypervisor (a4120f5bedfa) offline for 20-30 minutes due to hardware failure (disk plane). Services are restored after swapping the disks to a spare physical server.
Cause: One Toronto hypervisor (1ee830f71342) has one power supply failed. Due to mismatch in firmware we were unable to swap the failed power supply with one in stock.
We observed two brief network interruptions lasting a couple minutes each; network is now stable.
Progress
1 public update(s)
Remediation: Our upstream provider in Toronto is performing a software update on the core router resulting in intermitten network issues. We expect the update to be completed within the next 30 minutes..
we have not observed further issues following the configuration change.
Progress
2 public update(s)
Cause: The network issues due to router malfunction similar to "19 October 2022" issue have recurred. The latest network downtime was for two minutes at 11:00 EST. We upgraded router firmware last night as change list indicated it may solve the problem but it did not. We are continuing to look for alternative solutions as pre…
b4fa814b28f1 is online now. The outage started during routine maintenance and checkup of power bar in one rack. However, power utilization surged seemingly due to faulty network switch and tripped two breakers. After power was restored, several hypervisors had issues booting, most of them booted after removing disks not associated with root filesystem. But b…
Progress
4 public update(s)
Cause: b4fa814b28f1 is online now. The outage started during routine maintenance and checkup of power bar in one rack. However, power utilization surged seemingly due to faulty network switch and tripped two breakers. After power was restored, several hypervisors had issues booting, most of them booted after removing disks no…
Remediation: b4fa814b28f1 is online now. The outage started during routine maintenance and checkup of power bar in one rack. However, power utilization surged seemingly due to faulty network switch and tripped two breakers. After power was restored, several hypervisors had issues booting, most of them booted after removing disks no…
we have not seen more outages, so it appears that the sfp module replacement on 21 October 2022 has resolved the issue despite the one final outage after the replacement.
Progress
5 public update(s)
Cause: there was another brief outage due to router issue today morning, tonight at around 21:00 EDT we will swap sfp modules to see if it is issue with an sfp module.
Remediation: we have not seen more outages, so it appears that the sfp module replacement on 21 October 2022 has resolved the issue despite the one final outage after the replacement.
Our upstream provider, Cogent Communications, will be performing network maintenance in the Toronto datacenter starting midnight of 29 July 2022. The maintenance will cause network downtime in our Toronto region for up to 60 minutes between 30 July 2022 00:00 EST and 05:00 EST.
Progress
1 public update(s)
Cause: Our upstream provider, Cogent Communications, will be performing network maintenance in the Toronto datacenter starting midnight of 29 July 2022. The maintenance will cause network downtime in our Toronto region for up to 60 minutes between 30 July 2022 00:00 EST and 05:00 EST.
Cause: This appears due to datacenter outage of rack 46-D06 ( http://vms.status-ovhcloud.com/index_rbx6.html ). We are waiting for reply from the datacenter (OVH).
We continue to see temperature returning to normal at the datacenter. At this time all services in toronto region should be online, if you continue to experience difficulties please open a ticket or email support@lunanode.com. We apologize for the inconvenience this incident have caused.
Progress
3 public update(s)
Cause: the issues in Toronto today are caused by failure in the datacenter HVAC system due to a recent storm, meaning that a wide range of equipment is overheating. We will post updates as we receive more information from the datacenter (Cogent Toronto).
Remediation: HVAC units are being restored and temperature alerts on the servers are clearing
7887b099ad9f complete. This concludes the scheduled maintenance. If your virtual machine does not have network please log into dynamic.lunanode.com and power cycle the virtual machine; if the problem persists, please contact us at support@lunanode.com
The networking upgrade has been completed, virtual machines had to be rebooted in the region for the new configuration to take effect. We apologize for the inconvinience.
Remediation: The networking upgrade has been completed, virtual machines had to be rebooted in the region for the new configuration to take effect. We apologize for the inconvinience.
the second disk that was giving read errors has now also been replaced, and the RAID1 building has finished. At this time both disks have been replaced and the RAID1 array is in good status.
Progress
10 public update(s)
Data impact
Data loss confirmed
Cause: the second disk that was giving read errors has now also been replaced, and the RAID1 building has finished. At this time both disks have been replaced and the RAID1 array is in good status.
Remediation: services should be restored at this time. Please open ticket if your VM is still offline.
Cause: We are investigating network downtime in Roubaix. Services appear to be offline due to issue with vrack infrastructure from the datacenter that we use in Roubaix (OVH).
services remain online but we continue to monitor the situation. We still do not have details regarding the vrack infrastructure outage from the datacenter (OVH), which caused the network outage. An additional issue with volume storage was identified that arose because of our interventions while investigating and attempting to resolve the network outage; spe…
Progress
4 public update(s)
Remediation: Services have been restored. It is unclear at this time whether outage is caused by datacenter infrastructure issue or our network node server having kernel lockup. We will continue to investigate.
all services were successfully migrated off the hypervisor. However after further investigation, we do not find any disk issues on the hypervisor, so we have brought it back into service.
Cause: Due to a detected potential disk issue on Roubaix hypervisor e3ba7defacb5, we will perform emergency maintenance involving migrating all virtual machines off the hypervisor to other hypervisors. This will involve a reboot of each VM with 2-10 minutes downtime depending on the disk size of the VM. The maintenance will b…
we have switched over from primary router to backup router, since it seems to be router crashing problem. We will investigate further if the backup router has the same problem.
the maintenance was performed at 23:45 instead due to some issues. But it is done now with up to 2 minute network interruption to most services.
Progress
1 public update(s)
Cause: the maintenance was performed at 23:45 instead due to some issues. But it is done now with up to 2 minute network interruption to most services.
RFO: our monitoring system detected that one disk in RAID10 group failed today morning. We use service from OVH in Montreal and Roubaix and requested disk replacement. It seems they needed to turn server off to replace the disk instead of hot-swapping it. But now the RAID10 group is restored to normal. Note that in our main location Toronto we can always hot…
Progress
2 public update(s)
Cause: One Montreal SSD hypervisor (de6e4fa83fdc) is offline for emergency maintenance due to disk issue. We expect 10-20 minutes downtime.
Remediation: RFO: our monitoring system detected that one disk in RAID10 group failed today morning. We use service from OVH in Montreal and Roubaix and requested disk replacement. It seems they needed to turn server off to replace the disk instead of hot-swapping it. But now the RAID10 group is restored to normal. Note that in our…
the downtime was due to kernel panic. We will check further to determine if the current kernel is sufficient to avoid future recurrence of this downtime incident.
Progress
2 public update(s)
Cause: the downtime was due to kernel panic. We will check further to determine if the current kernel is sufficient to avoid future recurrence of this downtime incident.
Cause: We perform emergency maintenance on hypervisor 7fd1830faf11 in Toronto due to memory bugs associated with old kernel. Downtime ten minutes needed to update to new kernel to resolve these issues.
Cause: Hypervisor 1ee830f71342 is offline due to power cable tension issue and technician error during installation/migration of equipment for/to new 20A circuit. Services should be back online in five to ten minutes.
RFO Update (06 July 2020 21:00 EDT) : Outage summary: at 06 July 2020 10:05 EDT there was power surge at Cogent's Toronto 245 Consumers Rd datacenter which led to server restarts and one PDU failure in one of our racks. (Other Cogent customers were also impacted and throughout the day the datacenter was quite crowded.) Due to the PDU failure, servers in that…
Progress
13 public update(s)
Cause: RFO Update (06 July 2020 21:00 EDT) : Outage summary: at 06 July 2020 10:05 EDT there was power surge at Cogent's Toronto 245 Consumers Rd datacenter which led to server restarts and one PDU failure in one of our racks. (Other Cogent customers were also impacted and throughout the day the datacenter was quite crowded.)…
Remediation: RFO Update (06 July 2020 21:00 EDT) : Outage summary: at 06 July 2020 10:05 EDT there was power surge at Cogent's Toronto 245 Consumers Rd datacenter which led to server restarts and one PDU failure in one of our racks. (Other Cogent customers were also impacted and throughout the day the datacenter was quite crowded.)…
the reboot last night did not solve the latency problem, but we apply additional tuning steps today and confirm the issue is resolved and latency is low and stable.
Remediation: the reboot last night did not solve the latency problem, but we apply additional tuning steps today and confirm the issue is resolved and latency is low and stable.
the network appears to be stable now and we are closing this issue.
Progress
3 public update(s)
Cause: finally the datacenter has issue open about this incident: http://travaux.ovh.net/?do=details&id=45319& . Also at this time the network connectivity is fully offline. We observe heavy packet loss in Roubaix starting at 14:40 EDT due to datacenter issue. We are contacting the datacenter to investigate this problem.
Remediation: the network is back online now but it may still be very unstable. We have received absolutely no update from the datacenter.
we do not see more issues recently. We believe reboot and updated kernel solves the problem.
Progress
3 public update(s)
Cause: One Roubaix SSD hypervisor is offline for ten minutes from 16:45 to 16:55 EDT due to emergency maintenance to correct kernel issue causing high latency and dropped packets.
our team cannot find issue last night, switching to backup router and replacing fiber module does not help. Now the packet loss is gone. It must have been datacenter issue. Big waste of time.
Progress
2 public update(s)
Cause: we still see 1% packet loss in Toronto. It does not appear to be datacenter-wide issue. Our team is on-site and still investigating. We may switch to backup router or perform other similar actions that may cause brief network disruptions (less than one minute).
we do not see any further issues. Update 3 (09 April 2020 13:25 EDT) : services are back online at this time. We continue to monitor the hypervisor, but we believe the issues should be resolved.
Progress
3 public update(s)
Remediation: we do not see any further issues. Update 3 (09 April 2020 13:25 EDT) : services are back online at this time. We continue to monitor the hypervisor, but we believe the issues should be resolved.
services are back online at this time. Here is OVH issue.
Progress
2 public update(s)
Cause: Internal and external network for VMs on one Montreal hypervisor is offline due to datacenter issue. We are in communication with OVH to resolve the problem.
Remediation: services are back online at this time. Here is OVH issue.
Due to substantial cost increases imposed by the datacenter we use in Montreal and Roubaix, we are sunsetting those locations on 31 January 2023.
Compensation & remedy
Compensation or remedy offered
Cause: Due to substantial cost increases imposed by the datacenter we use in Montreal and Roubaix, we are sunsetting those locations on 31 January 2023.
Compensation/remedy: Make sure to migrate your VMs before 31 January 2023. You will receive three months of free credit for VMs migrated before 31 January 2023.