We gave 5 LLMs $100K to trade stocks for 8 months

>>cheese+(OP)
Just one run per model? That isn't backtesting. I mean technically it is, but "testing" implies producing meaningful measures.

Also just one time interval? Something as trivial as "buy AI" could do well in one interval, and given models are going to be pumped about AI, ...

100 independent runs on each model over 10 very different market behavior time intervals would producing meaningful results. Like actually credible, meaningful means and standard deviations.

This experiment, as is, is a very expensive unbalanced uncharacterizable random number generator.

>>Neverm+57
To their credit, they say in the article that the results aren't statistically significant. It would be better if that disclaimer was more prominently displayed though.

The tone of the article is focused on the results when it should be "we know the results are garbage noise, but here is an interesting idea".

zlacker