At the dawn of wide use of AI, a much-read paper asked whether Large Language Models (LLM) can develop earnings forecasts that beat analysts’ forecasts. It provided financial anonymous statements to GPT4 (without firm names) and asked it to forecast the direction (up or down) of future earnings (but not the amount). It beat human analysts on average in situations where analysts tend to struggle and worked just as well as a trained ML model. It appears that this is due to GPT4 being able to develop business narratives, something which analysts presumably do.
Perhaps the analyst benchmark is not a good one? Analysts are known to be sometimes biased in their forecasts. Perhaps a better benchmark might be forecasts from a focused financial statement analysis. That would be interesting to see. As would the ability of AI to forecast the amount of earnings and not just the direction.
See Kim, A., M. Muhn, and V. Nilolaev. 2024. Financial Statement Analysis with Large Language Models. Chicago Booth Research Paper. At https://ssrn.com/abstract=4835311.