Can you use autoregressive diffusion to generate market data?
In the field of quantitative finance, models typically work with a sequence of market data events for a specific symbol, such as orders being added to an order book, canceled, or executed. However, generative models are less common in this domain. A generative model could not only provide a point estimate of a symbol's price but also synthesize book events, including their arrival timing on the exchange, resulting in a much more detailed prediction.
Market data has both discrete and continuous features. While actions like adding or removing resting orders are discrete, price parameters are often treated as continuous due to the high cardinality. Additionally, timing is crucial, as the arrival of new orders can exhibit spiky distributions and burst patterns, especially around whole-number times.
To explore these complexities, Kavish, a research intern, developed an autoregressive diffusion model for market data. Autoregressive diffusion and flow-matching models are effective for dealing with sequential and multimodal data in fields like image, video, robotics, and audio. Kavish's goal was to create an event-level generative model of market data using autoregressive diffusion.
The intern's research involved analyzing four years of US equities data, including timestamps, prices, event types (trades, orders, cancellations, etc.), and other relevant information. The diffusion model aimed to generate subsequent event features, which were then appended to the data stream. Kavish employed an encoder–diffuser architecture, specifically a causally masked transformer encoder to produce a latent embedding at each time step.
The latent embedding was then fed into an event-kind head, generating a 2-categorical probability distribution to determine whether the next event was a trade or a BBO update. Continuous targets, like price and time, were generated using a diffusion head conditioned on the latent embedding and the event type using AdaLN conditioning.
During inference, only the encoded latent of the last event is supplied to the event-kind head, and a sample is taken from the output distribution. The following event's continuous features are generated using the diffusion head, conditioned on the last event's latent and the sampled event. The DDPM, or denoising diffusion probabilistic model, aims to predict the noise added to a sample, ultimately deriving a clean estimate by subtracting the scaled noise prediction from the true target and dividing it by the remaining signal level.
However, Kavish discovered that flow matching outperformed DDPM in this context. Flow matching interpolates linearly between noise and data, and the network is trained to predict the velocity along this line. As a result, the sampled trajectories are nearly straight, and flow models can be integrated accurately in relatively few steps. This approach proved more stable and accurate compared to DDPM, as it reduced the overflows that occurred when using more sampling steps.
Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.