Cover photo

Thoughts Sharing: Data Gravity matters more than ever

The increase in Data Generation is occurring at an astonishing pace. Companies worldwide have a premium on data, allowing them to make insightful decisions. And companies are shifting toward data-driven, cloud-based, and AI-powered structures. Data has become the bottom layer of our modern society and an immaterial extension of the body: our digital fingerprint. Everything, no matter if it is online or offline, can be collected, exchanged, processed, and enriched to be able to deliver value. Data has become a critical-and-sensitive asset. Massive data production has given rise to the problem of data management and the associated question of the mobility of applications and services pertaining to massive datasets.

The surge of new data-intensive applications like Edge Computing, Analytics, AI/ML, or IoT has contributed to the exponential growth of data. A great example from Data Center Magazine highlights the issue “By 2025, connected devices alone will generate an estimated 79 zettabytes of information. […] In 2016, the entire volume of all data on earth amounted to just 18 zettabytes - or about 720bn Blu Ray copies of Blade Runner: The Final Cut”. 80% of data worldwide will reside in enterprises in 2025.

There is an interesting chart from the Data Gravity Index report explaining the relationship between Data Creation/Generation and Data Gravity. In simple words: you can create data from anything, you can aggregate enriched data to get new insights, and repeat the process over and over to get more valuable insights. As the dataset grows, it becomes difficult to move it.

Source: Data Gravity Index Report
Source: Data Gravity Index Report

Hence, Data Gravity matters more than ever. The term is quite interesting; let's shed some light on it.

What is Data Gravity?

The term owes its origin thanks to Dave McCrory. In 2010, he coined the term "Data Gravity" on his blog to explain the attraction of applications and services for massive datasets. McCrory drew an analogy based upon the laws of Gravity, which viewed a considerably large dataset as a massive body or a planet in the universe of data. As the increase in the mass of a body leads to greater gravitational pull towards the massive body, snowballing of a dataset causes the applications and services, which are smaller datasets, to be pulled towards the massive dataset. When the data accumulate enough, it is nearly impossible to move, so the services and applications are forced to converge toward the data location to maintain a certain level of performance.

The proximity to the data affords the additional applications and services greater Throughput and lesser Latency leading to more excellent reliability of the applications and services. However, moving closer to bigger datasets would force to on-premise or full-on one cloud vendor. Below there is an illustration of the concept.

Illustration of the Data Gravity Concept (Source: https://www.tigosolutions.com/feedstory/1030)
Illustration of the Data Gravity Concept (Source: https://www.tigosolutions.com/feedstory/1030)

And here is an interesting formula provided by John McCrory to calculate Data Gravity by taking into account all the important variables involved: Latency, Bandwith, and Data Mass.

Data Gravity Formula (Source: John McCrory's blog and infoq.com)
Data Gravity Formula (Source: John McCrory's blog and infoq.com)

The surge of Data Gravity

Ten years earlier, McCrory was able to anticipate the phenomenon of Data Gravity. Today the numbers tend to corroborate what he envisioned. According to a report of the Data Gravity Index, G2000 Enterprises will produce data at the breathtaking rate of 14 million gigabytes per second by 2024. It says,” Data Gravity Intensity, as measured in gigabytes per second, is expected to grow across 53 metros by a compound annual growth rate of 139% globally through 2024’’. This will warrant an additional 20,000 petabytes of storage annually. Organizations with giant footprints worldwide generate data at an overwhelming rate. Over 5 billion people interact with the data in this day and age. The number is expected to rise to 6 billion by 2025. Every person will come across data once every 18 seconds. Hence, the engagement with data will go over the roof.

Smart devices, including sensors and cameras, generate real-time operational data, which has to be processed quickly to be of value. The world is responding to the problem of data gravity by creating data centers and hybrid cloud solutions, scaling up the hyper-scale industry, and edge computing networks. The surge in Data Gravity is driven by the growing real-time insight needs, the democratization of Cloud Technology, and data generation.

The Challenge of Data Gravity

Here are some key challenges which are related to Data Gravity.

External Data and Diverse Sources

Traditionally, the data of a company was bound within the company's warehouses. The company establishment encapsulated the tools, technology, and people to make use of the data. However, the situation has changed due to the advent of external data. According to Constellation Research, over 60% of the most crucial company data will be external. Furthermore, the data sources in a modern enterprise are highly diverse: from comments on social networks to excel sheets. It is simply not possible to bring the data from various sources to a central data warehouse to be used for any value.

Size of Data and Time

The interview of Lin Nease and Denis Vilfort of Hewlett Packard Enterprise(HPE), has highlighted two significant challenges accompanying Data Gravity: the size of the data and time.

Data has become an indispensable stack for modern businesses, irrespective of their sizes. Enterprise Digitalization leads to colossal data generation for large firms with a global presence. The allocation of resources and architecture for edge computing doesn't prevent data generation.

The goal to achieve smooth operationalization through automation piles up more data. Hence, data production peaks at the operational sites of a business enterprise. The greater the size of the data, the greater the data gravity. For example, camera streams to monitor various business processes generate hundreds of megabytes per second. This large amount of data has to be processed immediately to yield value (CCTV data has value only in real-time to ensure on-site employees’ safety or measure live productivity). An effort to send the data to a central data repository to be processed later is counterproductive as it inhibits real-time decision-making and increases costs. Edge devices warrant instant communication and data processing. Time is of the essence.

Additionally, the dataset is too heavy to be moved to a core cloud. So broadly speaking, data gravity may be viewed as the interaction between the amount of data, the distance, and the processing capacity.

Analytic Workload

Data Analytics is the lifeblood of any enterprise of this decade. Data and analytic experts have to come up with the answers to essential queries which guide the growth of a company. According to an IBM-sponsored study carried out by Forrester in 2020, most decision-makers agreed that assembling data for analytics takes more time than it should. Data gathering is accomplished chiefly manually, which is work-intensive and time-consuming. Consequently, processing slows down, latency increases, and innovative effort has to give way to operational working

Furthermore, continuous shuffling of data gives rise to technical issues. The continuous moving of data in a hybrid cloud environment gives rise to unfathomable complexity on occasion. All these factors deprive companies of real-time and insightful decision-making needed for growth in a highly competitive environment.

Here are some responses to the study.

IBM & Forrest Consulting: Leverage Data where It originates to drive substantial business benefits (Oct. 2020)
IBM & Forrest Consulting: Leverage Data where It originates to drive substantial business benefits (Oct. 2020)

Impact of Data Gravity on Companies' Cloud Strategy

It follows pretty naturally that the companies have to spend a great deal more to purchase cloud services. Ever mounting expenditure puts considerable strain on the financial resources of an enterprise. Based on the Data Gravity Index report, Global 2000 Enterprises spend $2.6 trillion annually on IT Infrastructure & Networking. The data continue growing in terms of the big data set and the data pulled to the data set, and so does the need for more storage. Companies opt for Azure, Google Cloud, and AWS often to create a hybrid cloud strategy to limit the influence of gravitational forces and avoid vendor lock-in, but increases the level of complexity and chances of technical burdens.

Furthermore, a central cloud core has to give way to on-premise or hybrid cloud strategy. This is not to say that a central database is not required at all, but that edge computing has to be under more significant focus. Companies have to rethink the Cloud strategy and Enterprise Architecture with a Data Gravity-first approach to overcome the potential performance/technical issues, mitigate Data Management risk, and limit high financial costs (e.g. migration).

Final Words

To put it all together, I believe Data Gravity matters more than ever now, it could be the large challenge of the cloud industry of this decade. Enterprises have to confront the phenomenon of enormous data generation, processing, storage, and integration. The exponential increase in data creation is not going to reverse; instead, it will flourish, creating more significant challenges for the enterprises and their cloud strategy. Parallelly, the fast adoption of data-intensive applications, like Analytics, AI/ML, or Edge Computing, does not make the job easier.

Data Gravity will push forward new types of issues related to Security – Scalability – Data Velocity, and enterprise data architecture.


I hope you enjoyed reading this blog post! if you have any questions and comments, or just want to chat, please feel free to reach out to me.