paper

MazeEval

A benchmark tests coordinate-based navigation in English and Icelandic.

Hafsteinn Einarsson evaluates eight models using generated mazes, coordinate feedback, distances to walls, navigation history, and a movement interface. Five fixed-seed mazes are tested per standard size.[1]

Contribution and relevance

The benchmark reports model disparities, repetitive navigation failures, and aggregate English–Icelandic performance differences. It connects Distributional Shift with the reliability of agents operating through tools.[1]

Evidence limits

A two-language comparison does not isolate training-data quantity as the causal explanation. Individual language comparisons do not survive the stated multiple-comparison correction. Larger o3 mazes receive additional exploratory testing rather than the repeated standard evaluation; success on one 30-by-30 maze is not a general perfect-performance guarantee. Novel task design also does not prove that all related training exposure is absent. The task evaluates navigation behavior, not consciousness. Compare Assessing Consciousness-Related Behaviors Using the Maze Test for a different setup and interpretation.

Sources

  1. MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models · Source record src-035 · Back to claim ↑1 ↑2

Last updated 2026-10-07