RESEARCH

Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games

ArXiv cs.AI · Tue, 02 Jun 2026 04:00:00 GMT

arXiv:2606.00103v1 Announce Type: new Abstract: We introduce a multi-turn interactive framework for reasoning evaluation that treats reasoning as active evidence acquisition and belief updating. Wherein, LLMs receive only the task rules, must issue targeted queries to a hidden en

Read original source Discuss with A.S.I.S