Data lakes is an ubiquitous term that is commonly used by vendors and enterprises. Typically, data lakes are spoken as a solution for data silos and as a enabler for digital transformation. The key idea behind the data lake is “no data left behind” — The goal is to use all types of data that an enterprise can lay hands on and use it to build competitive advantages.
However, the semantics of what a data lake is and the value it provides varies from customer to customer. There is no one technology (or a combination of technology) or process or methodology for creating a data lake. Traditionally, data warehouses were the data lakes — Primarily they use the transnational data that an enterprise gathered and use it for strategic planning and tactical execution. This continues to be key type of data lake and will continue for a foreseeable future. Customers understand the value and products and the process that enable this type of data lake is matured and well understood.
Essentially, what the traditional data warehouses does is it provides an excellent rare view mirror to make decisions going forward. However, as companies embark on digital transformation this is not sufficient. Let me explain — Digital transformation is not just about buying new hardware / software or moving into cloud or including IoT, it is fundamental shift on how an business wants to delight its customers, build competitive advantages, attract talent and grow. Among the many threads the enterprise has to transform and weave new threads, the key thread that revolves around “data” that an enterprise gathers and ways in which it can harness this data and successfully drive transformation.
Needless to say, the scope of data lake (creating and using the data lake) is much broader than traditional data warehouses. Data lakes needs to consider the ingestion, transformation, harmonization, blending of various types of data, gaining insights, visualizing those insights, acting on those insights is a super set of traditional data warehouses.
Typically, a top down approach is essential to define how the data lake and it usage should evolve. Based on this the data lake and its usage can be designed and implemented. The choice of technologies and integrating these technologies is a broad subject and requires its own blog post.
Based on my experience speaking with customers and partners, the following points are key aspects to any data lake
-Ingestion of data into the data lake
-Operationalizing the insights gained from the data lake
In my upcoming blogs I will go deeper into these two points.




