About a-team Marketing Services
The knowledge platform for the financial technology industry

A-Team Insight Blogs

Google Just Put a Price on Old Enterprise Data

Subscribe to our newsletter

By Jerry Caviston, CEO, Archive360.

In August, Google was selected as the successful bidder in a $10 million auction for enterprise data from Spirit Airlines. The proposed sale still requires court approval, but the bid has already delivered a useful signal. Historical data can retain value long after the process that created it has ended.

The collection includes about 100 million emails, 500 million Microsoft Teams items and millions of files, along with operational and commercial records. Google has said the data could help improve its products and AI models. An independent third party would de-identify it before transfer and several consumer-facing databases have been excluded.

The dataset offers something the public web cannot: a private operating history created by another company. Its potential value lies in the decisions, exceptions and relationships captured across millions of records.

Historical data can improve enterprise AI

Most organisations have treated old data as a storage cost, a compliance obligation or a migration problem. AI adds another question: could this history improve a current model, assistant or automated process?

An old ERP system may show how orders, pricing and exceptions moved through the business. Service records can ground an AI assistant in company-specific knowledge. Archived documents and communications preserve specialised language and the reasoning behind decisions. Retired applications may also provide real cases for testing AI performance.

The value may come less from one record than from patterns across many of them. Years of linked contracts, transactions, correspondence and outcomes can reveal how a business actually operates. Competitors cannot easily reproduce that history, and a general-purpose model is unlikely to reconstruct it reliably from public sources.

Not every old dataset is valuable. Some information is duplicated, inaccurate or too sensitive to use. A useful collection must contain distinctive knowledge for a defined AI initiative, with established quality, provenance and permitted use.

Most Legacy Estates Hide Rather Than Preserve Value

Many organisations cannot identify their historical data well enough to assess it. Files have been copied between repositories. Application exports have lost their original structure. Ownership is unclear, and access permissions no longer reflect the people or processes that created the information.

For AI, context is part of the data. A transaction without its source system, field definitions and related records may be misinterpreted. A document without provenance may be outdated, while a conversation separated from its participants and permissions can be misleading.

This is a common failure during application retirement. Organisations extract records from a legacy platform, place them in low-cost storage and consider the project complete. The data remains, but its schema, relationships, metadata and security model may be weakened or lost. It satisfies a storage requirement while offering little value to AI. For financial institutions, the affected systems may contain decades of payments, trading, client service, risk or operational history.

Governance Determines Whether Data Is Usable

The Spirit auction also shows why value cannot be separated from governance. The proposed transaction has drawn objections involving employee confidentiality, intellectual property and contractual rights. Removing names and personal identifiers does not resolve every issue. Data may still include privileged material or third-party content that cannot lawfully be repurposed.

Age creates other risks. Historical records can reflect obsolete policies, outdated product terms or poor data practices. Archives often contain duplicates and incomplete versions. If an AI system cannot distinguish an approved policy from a discarded draft, giving it more data can make its answers worse.

A usable AI dataset needs more than a connector to an archive. The organisation must know where the information came from, who may access it and whether the intended use is permitted. It must preserve relevant context, enforce retention and legal holds, and separate approved content from records that should be restricted or disposed of.

A Governed Archive Creates A Controlled Path To AI

A governed archive can preserve data from retired ERP systems, file shares and other legacy applications without flattening away the information needed to understand it. Indexing and cataloguing make the collection easier to assess, while retention, legal hold and access policies continue after the original system is gone.

That archive can provide a controlled path to AI. Teams can select an approved collection without exposing the entire data estate, trace its lineage and document what the AI system accessed.

This preserves future options without requiring indefinite retention. A dataset with no approved AI use today may become relevant as priorities change. If its context has been stripped away or it remains trapped in an inaccessible application, that value may be impossible to recover.

The Decision Starts Before Data Is Deleted

Google’s selected $10 million bid is not proof that every archive contains a hidden fortune. It is evidence that private operating history can have real value for AI, even after personal identifiers are removed and even when the data was generated by another company.

Before retiring an application or deleting a dataset, organisations should ask what operating knowledge it contains, whether it could support a defined AI use and what rights attach to it. They must also determine whether the information is complete and reliable enough for that purpose. Those decisions belong in AI planning, not at the end of a storage migration.

Google’s bid provides a visible price signal. It should prompt data leaders to look again at the systems they plan to retire and the archives they already hold. Organisations that can identify, govern and prepare this history will have an AI advantage that cannot be replicated with a generic public dataset.

Subscribe to our newsletter

Related content

WEBINAR

Recorded Webinar: Executing the Migration to Cloud to Enable Scalability and Innovation

Cloud-based services and processing have become essential to financial institutions as their data management demands have become more complex and expansive. Thousands of organisations have made the jump from their limited on-premises tech stacks to the near-infinite scalability opportunities of public and private clouds. They have also been motivated by the need to modernise their...

BLOG

12 Leading Vendors Operationalising AI & ML with Robust Data Pipelines

The transition of artificial intelligence and machine learning (ML) models from experimental sandboxes to production environments remains a persistent operational friction point. While quantitative researchers and data scientists can often demonstrate alpha in isolated backtesting environments, the institutionalisation of these models requires a level of data pipeline robustness, latency control and regulatory auditability that research...

EVENT

Data Management Summit London

Now in its 17th year, Data Management Summit (DMS) London returns In April 2027, to explore how to use data and AI to drive measurable business outcomes reliably, repeatedly and at scale.

GUIDE

Regulatory Data Handbook 2026 – Fourteenth Edition

Welcome to the fourteenth edition of A-Team Group’s Regulatory Data Handbook. Supervisors increasingly expect firms to demonstrate which rules apply, which data supports each obligation, who owns the control and how exceptions are identified and resolved. Policies and implementation programmes must now be supported by records that can withstand regulatory scrutiny. This edition examines material...