Patience Cave Launches MazeBench 3D Benchmark, Top AI Models Score 12%
Patience Cave introduced MazeBench, a puzzle-based, 3D benchmark designed to evaluate artificial intelligence models on long-term planning and visual spatial reasoning. Current top-scoring systems, including GPT-5.6 Sol, Fable 5, Opus 5 and Kimi K3, achieved a maximum score of 12%.
The benchmark requires continuous learning through recurring techniques and includes camera controls that prevent agents from brute-forcing solutions. While human operators can solve the levels, the results highlight a significant performance gap for foundation models tackling extended spatial tasks. The benchmark creator is using the 80% milestone to track progress toward more reliable long-horizon reasoning.
From the sources (10 posts)
@xeophonSuper excited to release @patience_cave's MazeBench to the public! I think it is one of the most exciting OOD benchmarks out there: Ultra-long horizon, beatable by humans, while hard for agents, even with Python access.
@xeophonMazeBench has recurring techniques, which the model has to learn, making it also a continual learning benchmark. Some of the levels also require camera control to solve, making it even harder for agents to just brute force.
@grad62304977RT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…
@htihleRT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…
@teortaxestexRT @xeophon: Super excited to release @patience_cave's MazeBench to the public! I think it is one of the most exciting OOD benchmarks out…
@andrewcurran_RT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…
@scaling01RT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…
@mlstreettalkRT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…
@ofirpressRT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…
@testingcatalogPatience Cave announced MazeBench, a new puzzle-based benchmark in a 3D form. 12% is the current top score 👀 - GPT 5.6 Sol - Fable 5 - Opus 5 - Kimi K3 How long will it take AI to get above 80? P.S. Looks like a sneaky plan to delay m