
By Mihir Shah, adviser to FINBOURNE Technology.
For a long time now, financial institutions have poured billions into building what has become known as the modern data stack. Data lakes, warehouses, ingestion pipelines and cloud platforms have all become the dominant blueprint for becoming a data-driven organisation. Yet the industry’s enthusiasm rests on a flawed assumption. Most of these architectures were originally designed to analyse a business rather than run one. That distinction may sound esoteric, but it has profound implications for every investment manager now racing to make the most of artificial intelligence (AI).
There is no doubt that, consolidating data into a cloud data warehouse has transformational value and every asset manager should be investing heavily in that capability. But AI agents need to make autonomous decisions in real time way before the data reaches the data lake. These agents need a sound data model for context to make real time decisions in operational workflows like Order entry, Trading, Portfolio construction, Research, Risk, Compliance and Treasury. Focusing only on the Data Lake leaves the job only half done. Operational data architecture needs equal focus and investment to harness the power of Agentic AI.Influential Paper
The prevailing vision of the modern data platform can be traced to an influential paper published several years ago by a16z, the Silicon Valley venture capital firm. That paper, and the reference architecture it set out, had two significant blind spots. First, it addressed only the analytics side of the business, warehouses, lakes, and the pipelines that feed them, and said almost nothing about the operational systems that execute trades and manage portfolios. Second, and more consequentially, it never mentioned the data model at all. It catalogued technologies while leaving out the one component that makes any of them useful, a gap that matters even more in the age of AI.
Many of the modern cloud data platforms are, in essence, very capable engines for storing and querying data, but they are not meant to run operational functions. They don’t come with a business data model built in: that still has to be built by the customer, along with the semantics that make the data meaningful. The real foundation of any investment platform is how its data is organised, structured and related. In reality, the conceptual data model of an asset manager is remarkably stable. Portfolios, positions, trades, securities and a handful of other core entities have formed the backbone of the industry for decades. In fact, there are only around ten fundamental entities that matter, and that number has barely changed in half a century. Without a coherent data model, fund managers simply create more silos.
The consequences become apparent as the business grows. One large asset manager found itself running some 400 separate applications, supported by roughly 300 databases, each holding its own version of the business. What begins as a manageable technology estate can evolve into an operational constraint, making it increasingly difficult to aggregate risk, scale investment processes or gain a consistent view of the organisation. The platform becomes an obstacle to growth rather than an enabler of it.
Bigger Challenge
The challenge becomes even greater at large, diversified financial groups. Multiple business divisions spanning brokerage, pensions and asset management often operate across separate data environments built over many years, in some cases well over one hundred separate data warehouses within a single organisation. Corporate functions such as finance, risk and audit struggle to obtain a consistent view of the organisation. At the same time, leadership teams are recognising a looming strategic reality: AI models are becoming widely available to everyone. Competitive advantage will not come from access to algorithms. It will come from access to decades of proprietary data, organised in a way that machines can actually use.
But this is where the discussion often goes wrong. The value is not simply in possessing proprietary data. Large language models require context in order to generate useful outputs. That context comes from semantics: an understanding of what data represents, how it relates to other information and how it reflects the real world. Those semantics are embedded in the enterprise data model. Without them, even the most advanced AI system is left interpreting disconnected fragments rather than understanding the business itself.
This is where many firms have their priorities backwards. Traditionally, organisations build applications first and worry about integration later. Each system creates its own version of reality, forcing firms to reconcile conflicting data downstream in warehouses. Integration becomes expensive, complex and perpetual.
Data warehouses still matter enormously. They are indispensable for historical analysis, machine learning and exploratory research. However, they are not operational systems. Warehouses are designed for unpredictable analytical queries. Businesses are run through predictable transactions that require speed, consistency and reliability.
The rush to AI is forcing firms to confront a reality they have long been able to defer. The firms that benefit most from AI will not be those with the largest data lakes or the newest models but those that understand their own business well enough to provide machines with the context needed to reason about it. The true foundation of competitive advantage is not the modern data stack. It is the enterprise data model that sits beneath it – and most firms have not built one.
Subscribe to our newsletter


