A conversation with Prateek Kansal, Head of Engineering, India operations, Komprise.

Artificial Intelligence has shifted from an experimental tool to a boardroom-level priority. The competitive edge that comes from monetising data is now measured in direct revenue gains. Yet, most of the data enterprises hold is unstructured – scattered across formats and silos, messy to manage and difficult to feed into AI systems effectively. Traditional data approaches like extract, transform and load (ETL) fall short here.
To unpack why this is happening and what enterprises need instead, Intelligent CXO sat down with Prateek Kansal, Head of Engineering for Komprise in India. In this wide-ranging conversation, Kansal explains the financial stakes of unstructured data, the limitations of ETL and how AI-driven data pipelines are becoming essential for performance, governance and cost control.
Why are CXOs so focused on unstructured data pipelines for AI right now?
Kansal: The pace of AI development is extraordinary and executives see real money on the table. The MIT Center for Information Systems Research reported that top-performing organisations generate about 11% of their revenues from data monetisation compared with just 2% for lower performers. That’s a fivefold gap and it’s widening.
The challenge is that AI thrives on unstructured data – documents, medical images, video, audio, logs and sensor data – which makes up about 90% of what enterprises store. But this data is noisy and cluttered. Feeding it indiscriminately into AI models increases costs and degrades outcomes. CXOs are intent on building automated governed pipelines because without them, AI projects either stall or deliver poor results while competitors push ahead.
What makes feeding unstructured data to AI so difficult?
Kansal: Unstructured data is both critical and problematic. It’s essential for inferencing but riddled with quality issues. Imagine an AI model sifting through terabytes of irrelevant or outdated files. The results will be inaccurate and costs for storage and compute will spike.
Sending too much uncurated data also raises the risk of exposing sensitive information such as intellectual property, personally identifiable information or regulated healthcare records. Organisations need systematic ways to profile and curate unstructured data at scale so only the right subsets reach AI processes.
Why can’t traditional ETL tools solve this?
Kansal: ETL was designed for a different world. It excels at pulling structured data from transactional systems, spreadsheets or relational databases – places where everything is neatly organised.
Unstructured data is a different beast. It’s massive in scale, scattered across silos and hard to classify. AI workflows are iterative and non-linear, with branching paths rather than straight lines. Moving unstructured data through a rigid ETL framework doesn’t work.
AI pipelines fill the gap. They can scan across on-premises, cloud and edge environments, filter irrelevant data, enrich it with metadata and feed it into AI models in the right form at the right time. Pipelines also apply AI-specific transformations like embeddings and chunking and ensure governance and auditability.
Why is metadata so central to these pipelines?
Kansal: Metadata is the compass. Without it, you’re lost in a sea of billions of files across formats – PDFs, CT scans, tweets, MP4s, IoT logs and more.
By tagging files with attributes such as type, owner, date or semantic meaning, metadata turns raw data into something searchable and actionable. Modern pipelines continuously enrich metadata – sometimes using AI itself – and index it into a global catalog that spans silos.
This enables a healthcare researcher to instantly locate the right diagnostic scans or a retail analyst to retrieve product images for a recommendation engine. Metadata unlocks unstructured data and makes it usable for AI.
How do pipelines improve both data quality and governance?
Kansal: Quality and governance go hand in hand. Pipelines enrich files with metadata, ensure freshness and flag reliability issues. They also provide governance guardrails, classifying and quarantining sensitive or protected information before it flows into AI models.
Automated pipelines maintain audit trails, documenting what data was used in which AI process. That level of accountability builds trust in AI outputs and ensures compliance. Without it, enterprises risk reputational damage, regulatory penalties or biased outcomes.
How do pipelines help control AI’s escalating costs?
Kansal: AI is resource-hungry and therefore expensive. Running large models across petabytes of uncurated data consumes massive compute and storage resources. If the same raw data is processed repeatedly, costs multiply.
Pipelines act as a cost-control mechanism by curating only the relevant subsets of data for a given AI task. This reduces redundant processing and accelerates performance. Models train and run faster on smaller, better-targeted datasets. For enterprises under budget pressure, pipelines are essential to scale AI sustainably.
How do pipelines keep AI applications current and relevant?
Kansal: One of the most overlooked challenges in AI is keeping models updated. A chatbot trained on last quarter’s catalog risks frustrating customers with outdated information.
AI pipelines solve this by continuously feeding fresh data into AI systems. In healthcare, that might mean delivering the latest imaging files to diagnostic models. In retail, it could mean ensuring assistants reflect real-time inventory. Keeping AI applications current is a defining advantage of automated pipelines.
What obstacles do enterprises face in building pipelines for unstructured data?
Kansal: Three main challenges stand out.
1. Scale: Many industries – healthcare, media, financial services – manage tens of petabytes of unstructured data. Moving it all is slow and costly. A better approach is to build a global index and metadatabase.
2. Diversity: File types vary enormously and metadata is inconsistent or absent. Automation is needed to harvest metadata from filesystems, headers, application layers and even external systems like CRMs and ERPs.
3. Governance: Sensitive data often lurks deep within storage systems, beyond the reach of conventional tools. Pipelines must include intelligent indexing, policy-based classification and automated detection of sensitive data.
Without these specialised capabilities, enterprises risk spending heavily on AI with little to show for it.
How will AI pipelines evolve in the next few years?
Kansal: Pipelines will grow smarter, more adaptive and more strategic. We’ll see advances in intelligent indexing, local preprocessing and governance controls at massive scale. Pipelines will also learn from usage patterns, dynamically optimising how data is curated and moved.
Compliance will become embedded by default, with auditability built into workflows. Over time, pipelines won’t just support AI – they’ll enable data monetisation as a core business strategy. By unlocking the value of unstructured data, they’ll turn what was once a liability into a powerful competitive asset.
The Bottom Line
Traditional ETL cannot handle the scale and complexity of unstructured data. Automated pipelines provide the curation, governance and cost control necessary to transform this messy resource into a strategic advantage.
As Kansal notes, the organisations that master unstructured data pipelines won’t just build better AI – they’ll reshape their competitive position by monetising data in ways laggards will struggle to match.


