paper
Assessing Consciousness-Related Behaviors Using the Maze Test
A navigation test measures task performance while leaving sentience unresolved.
Rui A. Pimenta, Tim Schlippe, and Kristina Schaaff evaluate twelve models on text-described mazes requiring first-person navigation instructions. Thirty-four test mazes follow separately allocated demonstrations.[1]
Method and contribution
Complete-path accuracy scores fully correct solutions; partial-path accuracy counts consecutive correct steps before the first error. Performance varies by model and prompting condition. These measures distinguish starting a valid plan from carrying it through.[1]
Alignment relevance and limits
The results inform reliable sequential behavior and Machine Introspection discussions. The authors associate these capabilities with components of consciousness, but their ethical statement explicitly says the test does not measure consciousness itself. Neither success nor failure proves the presence or absence of phenomenal experience. Claims about persistent self-models remain an interpretation of behavioral scores. This paper differs from MazeEval, which uses interactive coordinate feedback rather than a complete text-described maze.
Sources
Pages that link here
- Machine Introspection concept
- MazeEval paper
- Rui A. Pimenta person
Last updated 2026-10-07