An Exam for Active Observers
ActiveVision tests whether models can solve visual problems that require
iterative observation, rather than a single glance. Every scene is generated by a
deterministic program then re-rendered photorealistically while preserving the structure.
Each example is one static image. The video clips show the solving process.
Benchmark results
| # | Model | Score | Accuracy |
|---|---|---|---|
| ★ | Humanreference | 81.7 / 85 | 96.1% |
| 1 | gpt-5.5 | 9 / 85 | 10.6% |
| 2 | gpt-5.5 | 8 / 85 | 9.4% |
| 3 | gpt-5.5 | 7 / 85 | 8.2% |
| 3 | gemini-3.5-flash | 7 / 85 | 8.2% |
| 3 | gemini-3.5-flash | 7 / 85 | 8.2% |
| 3 | gemini-3.5-flash | 7 / 85 | 8.2% |
| 3 | gpt-5.5 | 7 / 85 | 8.2% |
| 8 | gemini-3.5-flash | 5 / 85 | 5.9% |
| 8 | claude-fable-5 | 5 / 85 | 5.9% |
| 8 | gemini-3.1-pro-preview | 5 / 85 | 5.9% |
| 8 | gemini-3.1-pro-preview | 5 / 85 | 5.9% |
| 12 | claude-fable-5 | 4 / 85 | 4.7% |
| 12 | claude-fable-5 | 4 / 85 | 4.7% |
| 12 | claude-opus-4-7 | 4 / 85 | 4.7% |
| 15 | gemini-3.1-pro-preview | 3 / 85 | 3.5% |
| 15 | claude-fable-5 | 3 / 85 | 3.5% |
| 15 | claude-fable-5 | 3 / 85 | 3.5% |
| 18 | gpt-5.5 | 2 / 85 | 2.4% |
| 18 | claude-opus-4-8 | 2 / 85 | 2.4% |
| 18 | gpt-5.5text-only | 2 / 85 | 2.4% |
| # | Agent | Score | Accuracy |
|---|---|---|---|
| 1 | Fable 5Claude Code | 43 / 85 | 50.6% |
| 2 | GPT-5.5Codex | 32 / 85 | 37.6% |
| 3 | Opus 4.8Claude Code | 21 / 85 | 24.7% |
The tasks
Dataset construction
Every puzzle is generated by a program, so its answer is known exactly — then re-rendered as a photorealistic scene that keeps the geometry intact. No human annotation anywhere in the loop.