The performance decay of LLM trading strategies
A couple of months ago, I wrote about an experiment where researchers from the San Francisco Fed asked ChatGPT to forecast inflation.
Recent research conducted by Chinese scientists investigated the performance decay of AI trading strategies powered by large language models (LLMs), such as ChatGPT and GPT-4o. The study discovered that these models struggle to maintain accuracy when applied to live markets compared to their performance in backtests. This phenomenon occurred because the training data used for LLMs leaks into the models, creating misleading forecasts once the training window closes.
To address this issue, the researchers proposed a method where trading LLMs create their own trading strategies, tested not only in backtests but also in counterfactual scenarios. The essence of this approach lies in running Monte Carlo simulations to test the strategies in artificial environments that were not part of the LLM's original training data. If the trading strategies perform well in both the backtest and the counterfactual testing, there is a higher likelihood of profitability in real-world applications.
The experiment focused on five LLM-based methods and restricted them to GPT-4o, ensuring that the training cutoff was known to be October 2023. The trading LLMs were then tasked with trading in the constituents of the Nasdaq 100 for two periods: Q2 and Q3 2021 (within the training window) and Q3 and Q4 of 2024 (testing the model after the training window closed).
While both periods showed similar returns for the Nasdaq 100 averaging around 13.5%, the total return of the trading LLMs varied significantly. In the backtest setting, the models achieved total returns ranging from 30% to 44%, significantly outperforming the index. However, when applied to 2024, the total return dropped between 9% and 22%.
The researchers suggest that this method could improve the robustness of AI-driven trading strategies, as the trading agents applied to pre-training and post-training periods demonstrated better results.
Written by urgent.news from Klement on Investing's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.