AI trading models are facing challenges in live market environments, with most systems reporting losses, according to recent public trading competitions. The Alpha Arena competition, operated by tech startup Nof1, highlighted these struggles as eight advanced AI systems, including Anthropic’s Claude and OpenAI’s ChatGPT, traded U.S. technology stocks with a starting capital of $10,000 each. The competition revealed that the overall portfolio suffered a loss of about one-third, with only six out of 32 outcomes turning a profit. The competition data showed significant discrepancies in trading behavior among the AI models. For instance, Alibaba’s Qwen executed 1,418 trades in one round, while Grok 4.20 placed only 158 orders. The models also displayed varied decision-making tendencies, with Claude favoring long positions and Gemini showing a preference for shorting. Despite these challenges, some models, like ChatGPT, demonstrated potential in specific areas, achieving a 68% accuracy rate in predicting earnings forecast directions for Q4 2025. The limitations of AI trading models are attributed to their inability to effectively weigh numerous factors affecting stock prices, leading to issues like poor trade timing and excessive trading. As traditional backtesting methods prove inadequate for LLMs, live market testing remains the primary evaluation method. Nof1 plans to enhance its AI models for the next season of Alpha Arena by providing them with more data sources and capabilities, although the company focuses on offering tools for retail traders rather than deploying AI directly on trading floors.