paper
MazeEval
A benchmark tests coordinate-based navigation in English and Icelandic.
Hafsteinn Einarsson evaluates eight models using generated mazes, coordinate feedback, distances to walls, navigation history, and a movement interface. Five fixed-seed mazes are tested per standard size.[1]
Contribution and relevance
The benchmark reports model disparities, repetitive navigation failures, and aggregate English–Icelandic performance differences. It connects Distributional Shift with the reliability of agents operating through tools.[1]
Evidence limits
A two-language comparison does not isolate training-data quantity as the causal explanation. Individual language comparisons do not survive the stated multiple-comparison correction. Larger o3 mazes receive additional exploratory testing rather than the repeated standard evaluation; success on one 30-by-30 maze is not a general perfect-performance guarantee. Novel task design also does not prove that all related training exposure is absent. The task evaluates navigation behavior, not consciousness. Compare Assessing Consciousness-Related Behaviors Using the Maze Test for a different setup and interpretation.
Sources
Pages that link here
Last updated 2026-10-07