About a-team Marketing Services
The knowledge platform for the financial technology industry

A-Team Insight Blogs

How ExtractAlpha Builds Trust in the Age of AI Signals

Subscribe to our newsletter

When a data provider sells raw data, the buyer can inspect what they are getting. When a provider sells a signal – a ranking, a score, a forecast of forward returns – the buyer is being asked to trust a research process they cannot necessarily see. That asymmetry sits at the centre of the systematic data business, and it becomes more acute as the modelling underneath those signals moves from countable features to transformer architectures whose behaviour is harder to describe, let alone audit.

ExtractAlpha announced its Transcripts AI Model on 20 July, an upgraded natural language processing signal built on company earnings call transcripts. The model uses contextual embeddings to capture language patterns associated with future stock performance, and extends coverage beyond the US to English-language calls across the Americas, EMEA and APAC, plus a Japanese-language module built on data from SCRIPTS Asia. ExtractAlpha reports strong risk-adjusted performance across global markets and outperformance against traditional transcript sentiment approaches; those results are set out in an accompanying white paper.

Vinesh Jha, Founder and CEO of ExtractAlpha, started the firm in 2013. A quant by training, he began his career in sell-side research before joining StarMine, where he built analytics on the accuracy of sell-side estimates and recommendations, and later moved to Morgan Stanley’s process-driven trading unit, PDT, building market-neutral equity strategies. His interest in unusual datasets dates from the August 2007 quant quake, when crowded positioning across systematic funds produced a chain of forced liquidations – an experience that pointed towards data nobody else held.

How Much Should A Signal Vendor Reveal?

What a prop desk could keep entirely to itself, a vendor has to explain. Asked how much ExtractAlpha discloses about its modelling, Jha reaches for a distinction. “I’d characterise it as a grey box, somewhere between a clear box and a black box,” he tells Market & Alt Data Insight. “We have to keep some things proprietary, or someone could replicate what we’ve done. But we also need to be very transparent about our research process, and we need to gain our clients’ trust – they’re essentially outsourcing part of their alpha-generation process.”

In practice this means a detailed white paper for each dataset, setting out the modelling process, the in-sample and out-of-sample design, and the general methodology. Although clients do not receive access to source code, Jha states that they receive enough detail to form a view on whether the model is robust and whether it has been overfitted.

What clients want varies by their own sophistication. A new desk at a multi-strategy firm or a recent fund launch may be content with a derived signal that is quick to deploy. An established quant shop with substantial in-house research capacity is more likely to want the raw data, or the features built on top of it. ExtractAlpha sells across that range, sourcing data it collects itself through its crowdsourced Estimize platform, data bought from established vendors, and data obtained through partnerships with firms that hold interesting assets but lack the capital markets research capability to monetise them.

Does Selling A Signal Destroy It?

The objection that follows any signal licensing model is crowding. If a signal works, and multiple funds trade it, the alpha is competed away and the client who bought it last is the one holding the decayed version.

“The topic comes up in almost every client meeting,” says Jha. “Clients naturally ask whether the signals will become crowded, whether other firms are using them, and how we think about that risk. Those are all valid questions – maybe a little overblown, but valid.”

His response is that ExtractAlpha is a small company; clients use the data differently from one another; and limits-to-arbitrage arguments mean no fund allocates heavily to a single signal, because these are risky positions with drawdowns rather than clean arbitrages.

The company has also treated the question as a commercial risk to itself. Jha says ExtractAlpha ran a research project to construct measures of how crowded a given signal might be – examining whether volume was rising in the tails of its stock rankings, and whether stocks in those tails had begun to cluster and move together more than previously. It compared those metrics across the two years before each signal’s launch and the two years after, on the reasoning that adoption takes time. Jha says it found no significant upward shift.

Where Does AI Stop?

The Transcripts AI Model is an upgrade rather than a first attempt. ExtractAlpha released its original transcripts signal in 2021, following research begun around 2018, covering US earnings calls sourced from FactSet. That model used word embeddings – clusters of words – to forecast returns, without imposing any prior view on whether a given cluster was positive or negative. It became the firm’s most successful product on its own account.

The upgrade adds context. Unlike a dictionary-based approach, which counts words in isolation and cannot distinguish between senses of a word, a transformer model reads the surrounding language to put those words in context. Jha’s illustration is the word “bank”, which might refer to a river bank or a financial institution depending on what sits around it.

“The question you then have to ask is how the AI knows those differences. It knows because it has studied language – and part of language is earnings call transcripts. You have to make sure the model isn’t inadvertently using information that wouldn’t have been available at that point in history,” says Jha.

Look-ahead bias renders historical simulation results meaningless, and it is harder to control in a pre-trained model than in a feature set the researcher has constructed himself. A client cannot detect it from the output. Jha describes an escalating ladder of AI use inside the research process. Code generation, where the researcher knows exactly what they want implemented, is low risk. Building internal tooling – backtesting infrastructure, results interpretation – is only slightly higher. Building a signal where the design is specified in advance, as with the Transcripts AI Model, is more aggressive but still directed. Beyond that lies signal origination.

“The next, more extreme version is where you say, ‘Hey AI, come up with signals for us.’ We’re not doing that yet, but it’s an area we’re actively researching,” he says.

The Horizon Question

ExtractAlpha targets a 90-day prediction horizon on the transcripts signal, rather than the intraday response times that news sentiment products are built around.

“Over the last 10 years, simple news sentiment in the US has become really fast – if you’re not acting on it within minutes, it’s gone,” says Jha.

That suits the firm’s client base, which sits in mid-frequency territory, and Jha sees convergence towards that band from both directions: high-frequency firms extending their holding periods, slower managers speeding up, and discretionary managers becoming more quantamental. He notes that such convergence would sharpen the crowding problem rather than ease it.

Delivery is evolving alongside. ExtractAlpha distributes via S3, Snowflake, FTP and APIs, is listed on the FactSet marketplace, and is examining Model Context Protocol implementations for AI-driven consumption. Jha is sceptical about marketplaces as a distribution channel, though – the product requires a quant researcher to explain it, and the consultative sales process runs over weeks.

Feedback

For a business built on understanding how data generates returns, ExtractAlpha operates with limited visibility into what its clients do with the signals. Some are forthcoming; most are not. Feedback arrives as a purchase decision rather than a critique, which leaves the company inferring demand from which datasets attract trials.

“That’s the downside of having the smartest customers in the world – they’re pretty secretive, and they’re not going to tell you what they’re doing. We ask people, and we get nothing in response,” says Jha.

Without that feedback, clients are left relying on what the vendor says about its own performance. ExtractAlpha has retired few datasets, and continues to support earlier model versions – clients on the 2021 transcripts signal are not obliged to migrate. But Jha volunteers that one of the firm’s signals, built on conventional price, volume and fundamental data, has become genuinely crowded and no longer performs as it once did. It remains available and supported, but with few new trials.

“Some datasets naturally decay over time,” he concludes. “We believe it’s important to be transparent about that and continuously evaluate what still adds value. We’re not trying to hide anything.”

Subscribe to our newsletter

Related content

WEBINAR

Recorded Webinar: The Data Foundation for Alpha – How fragmented data is eroding hedge fund performance

Alpha depends on more than models, talent and execution. It depends on the quality, consistency and timeliness of the data behind every investment decision. Many hedge funds still operate with fragmented datasets, inconsistent identifiers and manual reconciliation processes that slow research, distort signals and increase operational risk. As firms scale across strategies, regions and asset...

BLOG

Selling the Proof, Not the Data

For most of the past decade, a vendor specialising in natural-language processing (NLP) could sell a hedge fund something it could not easily make for itself. Sentiment scored from news, filings and earnings calls, provided as a clean feed, mapped to individual stocks and timestamped for point-in-time use. The work of pulling that signal out...

EVENT

Eagle Alpha Alternative Data Conference, Fall, New York, hosted by A-Team Group

Now in its 8th year, the Eagle Alpha Alternative Data Conference managed by A-Team Group, is the premier content forum and networking event for investment firms and hedge funds.

GUIDE

Risk & Compliance

The current financial climate has meant that risk management and compliance requirements are never far from the minds of the boards of financial institutions. In order to meet the slew of regulations on the horizon, firms are being compelled to invest in their systems in order to cope with the new requirements. Data management is...