
Financial institutions have spent years expanding the range of data available to investment, trading and risk teams. Increasingly, however, the challenge is not obtaining more data but making sense of what they already have.
Unstructured data is estimated to account for 80-90% of the information generated withinfinancial institutions, encompassing everything from regulatory filings and earnings-call transcripts to news, social media, satellite imagery and alternative datasets. Yet without the ability to classify, contextualise and govern this information, sheer volume can become a liability rather than an advantage.
That is one of the central conclusions of the new A-Team Group white paper, From Water to Wine: Deploying AI as the Multiplier for the Value of Unstructured Data, commissioned by LSEG Data & Analytics.
The paper argues that AI is changing the economics of working with unstructured information. Tasks that once required analysts to manually sift through thousands of documents can increasingly be automated, allowing specialists to concentrate on interpretation and decision-making.
But there is an important qualification: sophisticated AI cannot compensate for weak data foundations.
From information extraction to new signals
Early applications of generative AI largely accelerated existing processes – extracting information from filings, summarising documents or identifying changes in management commentary.
The emerging opportunity goes considerably further. AI models can increasingly quantify information that was previously difficult to capture systematically: subtle differences in sentiment, changes in language over time and even characteristics of speech such as hesitation, pitch and tonal variation.
The paper cites research suggesting that this additional granularity matters. For example, simply classifying a chief executive’s comments as positive may provide little predictive value, whereas isolating more specific characteristics such as future-oriented optimism can potentially reveal stronger signals.
This points towards an important development for alternative data practitioners. AI is not simply allowing firms to process existing datasets faster; it is potentially creating entirely new classes of machine-readable signals from previously qualitative information.
Deterministic and generative AI need each other
Another useful distinction made in the paper is between deterministic and generative AI.
Generative models are particularly effective at synthesising large volumes of content, answering questions across document libraries and identifying relationships between disparate sources. But deterministic techniques remain essential for tasks where consistency and auditability matter – entity identification, tagging, classification and disambiguation, for example.
Rather than competing approaches, the paper argues that the two should be treated a layers of the same architecture. Deterministic processing can provide the structured, traceable foundation upon which generative models operate.
For institutions deploying AI into production investment or trading workflows, that
distinction is significant. An impressive model output is of limited value if users cannot establish which information produced it or reproduce the process by which the result wasreached.
The infrastructure becomes part of the alpha
This also shifts attention from models towards the underlying data platform.
Unstructured and alternative datasets arrive with radically different characteristics. Real- time news and sentiment may need to be processed within milliseconds or seconds, while satellite-derived indicators might arrive in large periodic batches. Attempting to force both through the same architecture can increase cost and complexity while potentially destroying the time value of faster-moving signals.
The white paper therefore advocates flexible platforms capable of supporting real-time and batch consumption alongside hybrid cloud deployment.
Interoperability is becoming increasingly important too, particularly as agentic AI emerges. The paper highlights Model Context Protocol (MCP) as one mechanism for standardising how AI agents interact with different datasets and analytical tools.
That potentially makes data infrastructure itself a source of competitive advantage. If a potentially valuable dataset takes months to onboard, competitors may have exploited its informational value before it reaches production. As the paper puts it: “Operational speed is not an efficiency metric — it is an alpha metric.”
Governance becomes more important, not less
Perhaps the strongest message, however, concerns governance.
As AI systems consume more data and agents begin chaining together multiple models, datasets and analytical processes, seemingly small problems can propagate rapidly. Poor provenance, licensing restrictions or extraction errors may become embedded within downstream models and ultimately influence investment decisions.
The familiar principle of “garbage in, garbage out” therefore becomes considerably more consequential when AI operates at scale.
Rights management is particularly relevant for alternative data. Licensing agreements increasingly distinguish between activities such as classification, querying and generative reproduction. Institutions consequently need provenance, usage rights and auditability embedded within their data infrastructure rather than added afterwards
.
This leads to perhaps the paper’s most useful conclusion: competitive advantage is unlikely to come simply from deploying the newest AI model.
Instead, it will come from combining AI with high-quality data, effective classification, open and interoperable infrastructure, clear licensing rights and human oversight. Increasingly, the difficult part of extracting value from unstructured data may not be building the intelligence at the top of the stack, but ensuring the foundations underneath it are strong enough to support it.
Download the white paper here.
Subscribe to our newsletter


