
By Arun Sundaram, head of commercial data and analytics, LMAX Group.
Best execution was once about securing the strongest possible result for clients on an individual trade. Over the past few decades, the industry has become much better at measuring that outcome through fill rates, slippage and rejection ratios. But the next frontier isn’t the trade itself. It’s the quality of the data the trade creates.
Execution was once treated as a post-trade measure, assessed after the fact and then left behind. Now it feeds forward. Systematic strategies, real-time analytics and machine-learning models all learn from historical execution data. Every order that is priced, matched and filled becomes part of the dataset that informs the next decision. An AI model is just an engine. Its quality is entirely a function of the fuel, and the fuel is data. So best execution now covers two things that used to be one: how good the trade was, and how good the record of it is.
That second part is easy to underrate. In our own systematic work, the data has mattered more to performance than anything we did with model architecture. Models amplify signals, but they also create noise, and they’re at their worst in precisely the high-frequency, cross-venue settings institutions now trade in. Feed one bad input and you don’t get a slightly worse model; you get a brittle one that routes badly.
The Challenges
Three factors matter most, and none of them are easily apparent.
Inconsistent execution logic is the first. Any discretionary intervention at the venue level doesn’t just affect the trade in front of it. It fundamentally affects the record. An unobservable variable in the execution is an unobservable variable in the training data, and the model can’t tell a change in the plumbing from a change in the market. A best execution wrinkle in the moment becomes a permanent hole in the dataset.
Aggregation is the second, and it’s the one that looks most harmless. Institutional data gets combined across sources, venues and time windows before a model ever sees it. Do that carelessly and you sand off the microstructure that carries the information. Prices that were never available at the same instant get treated as if they were. Liquidity sitting in three different places gets added up as if it were one pool. The result is clean, orderly and a description of a market that doesn’t exist.
Then, comes time stamping. At the scale these systems run at, where the events that matter happen inside a few milliseconds, timestamp accuracy stops being housekeeping. It sets the order of events, and order is causation. Misaligned packets at that granularity don’t just add noise, they flip the signal entirely. The model decides a price move came before the flow that caused it, and you’ve built a strategy on a relationship pointing the wrong way.
Chasing Alpha
The solution is to perfect the architecture. Treat every order the same way, with no discretionary hold, and the data that comes out is deterministically reproducible: same inputs, same outputs, run it again and you get the same thing. We often use that as an internal test, and it’s an unforgiving one. If you can’t reproduce a signal deterministically from the raw data and a documented set of transformations, it isn’t ready for production, no matter how clever the model looks on a slide. Integrity has to be built where the data is born. Nothing downstream puts back what was lost at the source.
It’s also why chasing alpha keeps leading back to the boring work. Our research with Macro Hive reached the same conclusion. The competitive edge wasn’t created by a more sophisticated model. It came from cleaner, venue-specific data with accurate timestamps, better outlier handling and standardised microstructure features.
Critical Oversight
None of this takes the human out, and in fact, makes human oversight much more valuable. A good market-data person spots a structure break, a venue outage, regime changes and a stale feed long before the automated checks do. That’s the split worth building around: let the system flag the anomaly, let a person work out what caused it and what to do. Automate without that and all you’ve done is scale up how fast bad data spreads.
We are experiencing a convergence. Data infrastructure and execution infrastructure are turning into the same problem, and the venues with the cleanest execution end up, almost as a by-product, holding the most useful data. The transparency the industry fought for in how trades get done now has to cover the record they leave behind. Best execution can’t stop at the fill anymore. It must reach the data layer, because that’s where the next decision is really being made.
Subscribe to our newsletter


