The Illusion of Thinking: Strengths and limitations of reasoning models [pdf]

>>amrrs+(OP)
All the environments the test (Tower of Hanoi, Checkers Jumping, River Crossing, Block World) could easily be solved perfectly by any of the LLMs if the authors had allowed it to write code.

I don't really see how this is different from "LLMs can't multiply 20 digit numbers"--which btw, most humans can't either. I tried it once (using pen and paper) and consistently made errors somewhere.

>>thomas+kb1
Well that's because all these LLMs have memorized a ton of code bases with solutions to all these problems.

zlacker