An Exam for Active Observers
ActiveVision tests whether models can solve visual problems that require
iterative observation, rather than a single glance. Every scene is generated by a
deterministic program then re-rendered photorealistically while preserving the structure.
Each example is one static image. The video clips show the solving process.
Benchmark results
| # | Model | Score | Accuracy |
|---|---|---|---|
| ★ | Humanreference | 81.7 / 85 | 96.1% |
| 1 | gpt-6-astra | 65 / 85 | 76.5% |
| 2 | gpt-6-astra | 60 / 85 | 70.6% |
| 3 | gpt-6-astra | 59 / 85 | 69.4% |
| 4 | gpt-6-astra | 56 / 85 | 65.9% |
| 5 | gpt-6-astra | 45 / 85 | 52.9% |
| 6 | gpt-5.6-sol | 20 / 85 | 23.5% |
| 7 | gpt-5.6-sol | 19 / 85 | 22.4% |
| 8 | gpt-5.6-sol | 17 / 85 | 20.0% |
| 9 | gpt-5.6-sol | 16 / 85 | 18.8% |
| 10 | claude-fable-5-1 | 12 / 85 | 14.1% |
| 11 | gpt-5.6-sol | 11 / 85 | 12.9% |
| 12 | claude-fable-5-1 | 10 / 85 | 11.8% |
| 13 | gpt-5.5 | 9 / 85 | 10.6% |
| 14 | gpt-5.5 | 8 / 85 | 9.4% |
| 15 | gpt-5.5 | 7 / 85 | 8.2% |
| 15 | gemini-3.5-flash | 7 / 85 | 8.2% |
| 15 | gemini-3.5-flash | 7 / 85 | 8.2% |
| 15 | gemini-3.5-flash | 7 / 85 | 8.2% |
| 15 | claude-fable-5-1 | 7 / 85 | 8.2% |
| 15 | claude-opus-5 | 7 / 85 | 8.2% |
| 15 | gpt-5.5 | 7 / 85 | 8.2% |
| 22 | claude-opus-5 | 6 / 85 | 7.1% |
| 23 | gemini-3.5-flash | 5 / 85 | 5.9% |
| 23 | claude-fable-5 | 5 / 85 | 5.9% |
| 23 | claude-fable-5-1 | 5 / 85 | 5.9% |
| 23 | gemini-3.1-pro-preview | 5 / 85 | 5.9% |
| 23 | gemini-3.1-pro-preview | 5 / 85 | 5.9% |
| 28 | gpt-5.6-sol | 4 / 85 | 4.7% |
| 28 | claude-opus-5 | 4 / 85 | 4.7% |
| 28 | claude-fable-5 | 4 / 85 | 4.7% |
| 28 | claude-fable-5 | 4 / 85 | 4.7% |
| 28 | claude-opus-4-7 | 4 / 85 | 4.7% |
| 28 | claude-opus-5 | 4 / 85 | 4.7% |
| 34 | gemini-3.1-pro-preview | 3 / 85 | 3.5% |
| 34 | claude-fable-5-1 | 3 / 85 | 3.5% |
| 34 | claude-fable-5 | 3 / 85 | 3.5% |
| 34 | claude-opus-5 | 3 / 85 | 3.5% |
| 34 | claude-fable-5 | 3 / 85 | 3.5% |
| 39 | gpt-5.5text-only | 2 / 85 | 2.4% |
| 39 | gpt-5.5 | 2 / 85 | 2.4% |
| 39 | claude-opus-4-8 | 2 / 85 | 2.4% |
| # | Agent | Score | Accuracy |
|---|---|---|---|
| 1 | GPT-6 AstraCodex | 78 / 85 | 91.8% |
| 2 | Fable 5.1Claude Code | 65 / 85 | 76.5% |
| 3 | Opus 5Claude Code | 58 / 85 | 68.2% |
| 4 | Fable 5Claude Code | 43 / 85 | 50.6% |
| 5 | GPT-5.6 SolCodex | 38 / 85 | 44.7% |
| 6 | GPT-5.5Codex | 32 / 85 | 37.6% |
| 7 | Opus 4.8Claude Code | 21 / 85 | 24.7% |
The tasks
Dataset construction
Every puzzle is generated by a program, so its answer is known exactly — then re-rendered as a photorealistic scene that keeps the geometry intact. No human annotation anywhere in the loop.