Insight Cloud Resilience – Comparison of solutions for data protection and cyber recovery in the cloud

Philip Röder — 09. Dec 2025
Reading Time: 5:05 minutes

Cloud Resilience - Vergleich von Lösungen für Datensicherung und Cyber Recovery in der Cloud

Hyperscalers also offer backup solutions. However, simply backing up data in the cloud is not the same as ensuring data security in the cloud. Shared responsibility models make a thorough needs analysis for cloud resilience and recovery essential. Ultimately, the responsibility for data security always rests with the company itself.

Why cloud resilience is becoming increasingly important

The importance of resilience in general has increased significantly in recent years. This is due not only to legal requirements—such as NIS2 and DORA—but also to the significantly increased risk of cyberattacks. According to Bitkom, approximately 87 percent of companies were affected by cyberattacks this year. Many companies are already working with partners to develop concepts for making their on-premises infrastructure as resilient as possible against cyber threats. This shift has also been observed on the vendor side in recent years. Particularly in the disaster recovery environment, the scope of functionality has shifted considerably towards cyber resilience. A key issue that most vendors—such as Rubrik and Veeam—are trying to address in this area is the fastest and most successful recovery possible after a cyberattack.

This topic is also gaining increasing importance in the cloud environment. Customers are continuously modernizing their infrastructure, and many are relying entirely or at least partially on hyperscalers to comprehensively meet their requirements. As a result, many cloud, hybrid, and multi-cloud environments are emerging. Bitkom predicts that every company will be using cloud services within five years.

Even though the simplicity of the cloud typically suggests that one no longer needs to deal with traditional IT infrastructure issues, according to the providers' shared responsibility models, companies themselves are responsible for the security of their data. Therefore, if a cyberattack occurs on a company's cloud environment, the company must be able to independently restore its critical business processes.

Using the cloud does not automatically mean that one is resilient; rather, it creates another area within the infrastructure where one is responsible for comprehensive data backup and recovery capabilities. In practice, it is often evident that this topic is still new territory for many cloud teams. In addition to technological safeguards, gaps are also frequently found in the link to general Business Continuity Management (BCM) and to the emergency strategy with regard to the cloud.

What is cloud resilience – and what is resilience in the cloud?

Cloud Computing Insider defines cloud resilience as the "ability to withstand disruptions" and problems such as hardware defects, software errors, cyberattacks, network outages, or natural disasters. Resilience encompasses all cloud components and systems, such as servers, storage systems, networks, and applications." It is emphasized that "cloud resilience [...] is about predicting potential IT disruptions in a company. This includes business continuity planning and the question of how essential cloud services can be maintained or restored as quickly as possible after an outage." (Source: Cloud Computing Insider, December 4, 2025)

In practice, from our perspective, implementation can be divided into two different strategic approaches:

  • Companies that are not yet in the cloud or that host a large part of their infrastructure on-premises want to improve their overall resilience using the cloud. Examples include automated recovery tests to the cloud, additional data backup to the cloud, and failover scenarios from on-premises to the cloud. Here, the cloud is used as part of a comprehensive BCM/DR concept.

In this article, we primarily focus on cloud resilience for companies in a cloud or hybrid infrastructure:

  • How can I use cloud-native tools and additional tools to build a cloud architecture that ensures maximum resilience for my critical business processes? Key principles include data redundancy and distributed systems, automation, monitoring, failover/multi-region, zero trust, and backup & disaster recovery.

Below, we examine the best practices of cloud resilience in the backup and recovery environment, as this area of ​​resilience, in a crisis, is the only way to restore critical data as the "last line of defense."

Data loss in the cloud can be caused by more than just cyberattacks. Misconfigurations, accidental deletion, deliberate deletion by employees, IAM errors, and incorrectly configured storage policies can all lead to significant and business-critical data loss. Although very rare, outages at the cloud provider itself can also result in temporary inaccessibility of business-critical data.

What to consider for end-to-end cloud resilience

There are several reasons why cloud resilience is often highly complex. One is the inherent technical complexity.

Hyperscaler architectures and functionalities

Despite their commonalities, different hyperscalers have distinct architectures and functionalities. It's crucial to understand how the standard configurations work and their respective advantages and disadvantages. Additionally, all hyperscalers offer a multitude of interdependent services, further increasing complexity. Distributing data across multiple regions can make maintaining consistency across these services and functionalities challenging. Complex permission models can lead to users receiving permissions they shouldn't have. For example, development teams are often able to copy data from production environments and populate their test environments with it. If the same security guidelines are not followed in these test environments as in the production environment, this presents a potential attack surface.

Hybrid infrastructures

Another layer of complexity arises from hybrid infrastructures. Most companies still have an on-premises infrastructure, and many also rely on multiple hyperscalers. The different data sources, application architectures, and dependencies can significantly complicate a clean disaster recovery process.

Workload security

In addition to these more technical aspects, cloud computing also brings organizational complexity. Due to various teams with different responsibilities (cloud, security, network, DevOps) and different tools and processes, coherent documentation and clear ownership structures for specific workloads are often lacking. Frequently, there is no overall organizational concept that considers the entire infrastructure in terms of its resilience.

External and internal threats to cloud resilience

In addition to the threat of cyberattacks, misconfigurations, accidental or malicious deletion, and technical risks, vendor lock-in and data portability, as well as audit, compliance, and legal hold of relevant processes, play a significant role in why such a comprehensive concept is essential.

A holistic recovery strategy is essential

In summary, a clear and holistic recovery strategy is essential, especially in a cloud or hybrid environment, to ensure business success.

Which resilience and recovery concept is right for my infrastructure?

At the outset, it's crucial to conduct a process analysis and subsequent gap analysis to precisely identify the needs and all systems and requirements that are essential for the company's resilience.

Step 1: Process analysis from data classification to recovery plan

Before making technical decisions and implementing the plan, it's essential, both in the area of ​​cloud resilience and cyber resilience in general, to first conduct a structured process analysis:

  • to define potential gaps with the BCM strategy and regulatory requirements, and
  • to classify business-critical applications and data.
  • Important: The rest of the infrastructure and any SaaS applications used should also be considered.

Therefore, all existing documents related to Business Continuity Management (BCM) should be reviewed, and existing regulatory requirements should be systematically addressed. Based on this, various categories are developed to help classify the criticality of individual business processes. Each of these categories is then assigned SLAs to define how quickly these processes must be operational again in an emergency to ensure business success. Following this analysis, it must be identified which systems support the respective processes and how critical each system is to ensuring the continuity of the respective process.

Once all critical processes and their associated systems have been documented, the objects and data belonging to these systems are identified. Often, there are also overarching systems such as Identity and Access Management (IAM), which is crucial for the operational capability of all systems. After this categorization, a concrete plan can be created outlining the sequence in which which technical objects must be restored in the event of data loss.

After assigning the individual technical objects to specific recovery processes and timeframes in accordance with the KPIs (RTO/RPO) defined in the SLAs, various recovery tests are conducted based on the currently available technical capabilities. These tests differ primarily according to the trigger for a recovery (data encryption, data loss, hardware failure, failure of individual systems).

Step 2: Gap analysis and decision-making

Using these test scenarios, a gap analysis is created, which is then used to identify the technical solutions that will need to be implemented in the future.

The identified gaps help define the scope of a new solution and establish clear requirements for the technologies to be backed up. In addition to traditional backup, topics such as disaster recovery, failover and failback, malware scanning, immutability, air-gapping and threat hunting, as well as high availability for specific systems, play a crucial role.

The existing skills within the company, the usability of the tool, and automation possibilities are also relevant. Of course, data retention and availability requirements also play a crucial role in the decision-making process. (In-cloud, cross-region, on-premises, data sovereignty)

Last but not least, the costs and licensing models of the various providers naturally determine the most suitable solution.

Once all factors have been identified, a well-informed decision can be made regarding a future-proof cloud resilience concept that seamlessly integrates into the existing strategy, processes, and infrastructure.

Which cloud resilience technology providers are available on the market?

Below, we compare and highlight what to consider when selecting cloud resilience vendors, which providers offer advantages based on our experience, and the role played by solutions from well-known hyperscalers.

Do all hyperscalers offer backup solutions? – Yes, but their standard solutions aren't comprehensive enough.

All of the most well-known hyperscalers (AWS Cloud, Microsoft Azure, and Google Cloud Platform (GCP)) offer their own solutions to increase resilience. This includes both backup solutions (AWS Backup, Azure Backup, Google Cloud Backup & DR) and solutions for faster data recovery, such as snapshot-based backups of primary data. In most cases, these solutions also perform a standard backup of existing data to a certain extent.

However, if you want to use these solutions as part of a more complex infrastructure, or if you have multiple applications, data sources, a large number of technical objects, or even sub-organizations within your hyperscaler, the use of data backup and disaster recovery tools becomes a complex task. SLAs must be manually assigned to all objects to be backed up.

Data Backup in the cloud – rarely the responsibility of cloud or development teams

This task usually falls to cloud or development teams, rather than infrastructure teams, who rarely deal with data backup. Recovery scenarios and recovery tests must be configured and set up manually, and cloud-native tools don't offer the ability to scan existing backups for malware to minimize recovery time in the event of a cyberattack. Furthermore, the built-in tools can only adequately back up cloud-native workloads, which means that in a hybrid or multi-cloud infrastructure, different disaster recovery tools must be used – further increasing complexity.

Incorrect settings can also lead to high, unexpected costs. To implement a robust overall technical concept using the built-in tools, it is recommended to consult with specialized experts about the available options.

Cloud backup architectures with Veeam, Rubrik, Cohesity, and Commvault

Most major backup vendors have now expanded into the cloud market, offering disaster recovery solutions in the cyber resilience sector. The best-known vendors offering cloud resilience solutions are Veeam, Rubrik, Cohesity, and Commvault. Each vendor has its strengths and weaknesses and focuses on different areas. Therefore, choosing the right vendor is always a matter of personal preference and should not be made without the thorough process and gap analysis described above.

Veeam

Veeam has been a long-established backup provider, serving a large customer base, primarily in the mid-market, but also in the enterprise sector. Veeam's backup solutions are therefore highly reliable and have a broad range of applications. Companies already using Veeam to back up their on-premises systems often consider Veeam again when it comes to cloud-based data backup. Over the past decade, Veeam has positioned itself as one of the leading backup providers and is continuously developing its expertise in cyber resilience. With its multi-cloud approach, Veeam offers the ability to back up data within a single platform, even in infrastructures that utilize multiple cloud providers. The Veeam Data Platform allows for the backup of AWS, Azure, and GCP. Veeam's cloud-specific backup solutions are available, as usual, in the marketplaces of the hyperscalers.

One drawback of the Veeam solution is its relatively complex architecture, as Veeam combines various solutions into a single platform. You have to familiarize yourself with each interface individually, and it's a lot of work to set up backup policies for all the necessary workloads.

Rubrik Security Cloud

From the outset, Rubrik positioned itself not as a "simple backup provider," but clearly as a leader in the cyber resilience field. Its goal is to minimize its customers' cyber recovery time and ensure they can resume operations as quickly as possible after a cyberattack. Rubrik also offers the ability to back up data from all major hyperscalers via its Rubrik Security Cloud platform.

The ease of use is particularly noteworthy. SLA-based backup policies, which can be applied to all technical objects after a one-time setup, allow for the effective implementation of structured cyber resilience processes, as described in the first part of this article. Additionally, Rubrik indexes the metadata of each backup, enabling ransomware scans of backups within seconds after a cyberattack. This ensures rapid identification of the last clean recovery point and a significantly reduced recovery time.

The company is also frequently at the forefront of trending topics such as Agentic AI. Due to its strong security focus, the ecosystem is more limited than that of other providers, and the company relies heavily on its own proprietary solutions.

The potential drawback of Rubrik Security Cloud for some end customers lies in its pricing model, which is based on the functionality used.

Cohesity Data Cloud

The Cohesity Data Cloud is the comprehensive cyber resilience solution from the manufacturer Cohesity. Through the acquisition of Veritas in 2025, Cohesity significantly expanded its reach and now has access to a large customer base. Cohesity also follows a platform approach and is able to back up a wide variety of hyperscalers from a single interface. Predefined backup rules can be assigned to individual resources via predefined policies. Configuring these policies is somewhat more complex and should be done by experts to ensure optimal performance and coverage.

With Cohesity DataHawk as an additional feature, Cohesity also offers the ability to react early to potential threats and reduce the threat landscape through AI-driven anomaly detection in backups. Backups can also be scanned for attack vectors to identify malware within backups or to determine a clean restore point. Thanks to its very open ecosystem, Cohesity also allows third-party developers to offer apps on the platform. Cohesity is also among the more expensive solutions on the market and sometimes has a more complex licensing model than other vendors.

Commvault

With years of experience in the market, Commvault offers a comprehensive cyber resilience platform – the Commvault Cloud – which is primarily used by many enterprise customers. Commvault is also capable of securing the data of all major hyperscalers within a single platform. The platform has evolved over several years and is very comprehensive.

A particularly noteworthy feature of Commvault is its CleanRoom approach, which allows individual workloads to be restored within a controlled environment to verify recoverability and identify malware in backups. This makes it easier to implement and monitor holistic, automated recovery tests. Furthermore, there are extensive functionalities for automating enterprise-level recovery.

Commvault also offers the option to restore some infrastructure configurations and provides a wide range of workloads that can be secured.

The biggest disadvantage of Commvault is the complexity of the historically grown environment, or of the individual tools, which can lead to increased internal costs.

Conclusion: Vendor comparison

In summary, the individual tools share some commonalities, such as centralized control, monitoring, automation, and the reduction of workload for existing teams.

However, there are also differences that should be evaluated based on the use case, existing infrastructure, and defined strategic goals. The technical and procedural requirements in the on-premises environment should also be considered.

Summary regarding cloud resilience: Sound strategies in the interplay of technology, organization, and people

Overall, it can be said that using hyperscalers does not absolve companies of their responsibility to address resilience and cyber recovery, nor to meet regulatory and compliance requirements. With the shared responsibility models of hyperscaler solution providers, the responsibility for data security remains with the company itself.

Before proceeding with a technical evaluation of the necessary tools, the organizational foundations that can be crucial in an emergency should first be established. Clear governance and defined KPIs should be used to develop structured recovery plans in line with the company's overall business continuity management (BCM) strategy. Understanding the criticality of individual processes and related applications and data, as well as identifying gaps in KPI compliance, provides the foundation for creating a technological vision for the future.

Resilience is always an interplay between technology, processes, organization, and people. A combination of hyperscaler-native tools with tools from vendors mentioned above, coupled with cyber resilience and cloud expertise, can be crucial building blocks for successfully implementing a business continuity management strategy. However, this does not absolve companies of the responsibility to define the right strategies and implement them in future-proof architectures. In such a process, companies should work with experts to define the necessary strategies and, based on these, make a well-informed selection of the required tools.

You were interested in this, then you may also be interested in...