Command Palette
Search for a command to run...

Patience Cave Launches MazeBench 3D Benchmark, Top AI Models Score 12%

aiai-modelingai-research-evals 10 posts · 9 accounts

Patience Cave introduced MazeBench, a puzzle-based, 3D benchmark designed to evaluate artificial intelligence models on long-term planning and visual spatial reasoning. Current top-scoring systems, including GPT-5.6 Sol, Fable 5, Opus 5 and Kimi K3, achieved a maximum score of 12%.

The benchmark requires continuous learning through recurring techniques and includes camera controls that prevent agents from brute-forcing solutions. While human operators can solve the levels, the results highlight a significant performance gap for foundation models tackling extended spatial tasks. The benchmark creator is using the 80% milestone to track progress toward more reliable long-horizon reasoning.

From the sources (10 posts)

@xeophon

Super excited to release @patience_cave's MazeBench to the public! I think it is one of the most exciting OOD benchmarks out there: Ultra-long horizon, beatable by humans, while hard for agents, even with Python access.

@xeophon

MazeBench has recurring techniques, which the model has to learn, making it also a continual learning benchmark. Some of the levels also require camera control to solve, making it even harder for agents to just brute force.

@grad62304977

RT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…

@htihle

RT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…

@teortaxestex

RT @xeophon: Super excited to release @patience_cave's MazeBench to the public! I think it is one of the most exciting OOD benchmarks out…

@andrewcurran_

RT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…

@scaling01

RT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…

@mlstreettalk

RT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…

@ofirpress

RT @patience_cave: Introducing MazeBench, a 3D open world environment that evaluates long term planning with visual spatial reasoning. It…

@testingcatalog

Patience Cave announced MazeBench, a new puzzle-based benchmark in a 3D form. 12% is the current top score 👀 - GPT 5.6 Sol - Fable 5 - Opus 5 - Kimi K3 How long will it take AI to get above 80? P.S. Looks like a sneaky plan to delay m

Preview built on a synthetic news corpus (16 weeks, Apr–Jul 2026). Impact calls are model reads, not price data.

About Archive