For AI agents: the public content index is available at https://we0.ai/llms.txt, and the English article bundle is available at https://we0.ai/llms-full.txt.
For AI agents: the complete content index is available at https://we0.ai/llms.txt, the full English article bundle is available at https://we0.ai/llms-full.txt, and this page is available as Markdown at https://we0.ai/articles/vista-visual-harness-explained-lossless-v-bfcc0210.md.
A visual-native harness for multimodal agents has arrived.

A visual-native harness for multimodal agents has arrived.
In the paper “VISTA: A Visual Harness for Reasoning in an Interactive World,” researchers Qiushi Han, Keya Hu, Linlu Qiu, Cathy Wu, and Kaiming He introduce VISTA, a framework designed to help general-purpose multimodal models reason over long interactive tasks without discarding the visual evidence they have already seen.

The basic idea is simple but important.
Instead of converting an interactive environment into text or code and asking a model to rely on a compressed description of the past, VISTA lets the model:
The framework does not require retraining the underlying multimodal model.
That distinction is central to the work. The researchers are not proposing a new foundation model. They are changing the interface around an existing model so that it can preserve, retrieve, and inspect visual evidence over time.
On the 25 public ARC-AGI-3 games, VISTA with Claude Opus 5.0 reaches a perfect 100.00 Relative Human Action Efficiency (RHAE) score and completes all 25 games. The agent uses 57.4% fewer actions than first-time human participants in the paper's reference baseline.
With GPT-5.6 Sol, the same VISTA framework reaches an RHAE score of 99.00.

The team also evaluates the framework on browser games, mazes, line-connection puzzles, and alternative visual representations, suggesting that the underlying idea may extend well beyond one benchmark.
So how does it work?
Visual perception and representation learning have been a recurring theme throughout Kaiming He's research, from ResNet and Mask R-CNN to MAE.
VISTA shifts that line of work toward a different question:
How can an agent accumulate, preserve, and reuse visual experience while interacting continuously with an environment?
The answer is a harness.
Harnesses have become an increasingly important part of long-running AI-agent systems.
Anthropic, for example, previously described how a coding agent can continue a long software task across multiple context windows by saving unfinished work, progress notes, and code state outside the immediate conversation context.

A language-oriented harness can preserve information such as:
This lets an agent divide a complex task into stages and continue even after the original context window has filled up.
That approach works well when the relevant state can be captured in text and code.
But interactive visual environments are different.
A model may need to remember:
A textual summary cannot always preserve those details faithfully.
Worse, the model often does not know when it first sees an image which visual detail will become important later.
One existing approach to ARC-AGI-3 is to convert each visual state into a text grid made of numbers. An agent can then write programs to simulate the environment, infer rules, and test action plans.
That method can be effective, but it imposes an abstraction before the model has fully understood the scene.
As the environment becomes more complex, it becomes harder to encode:
And as the interaction history grows, early visual states may be removed from the model's context or replaced by lossy summaries.
VISTA's research question is therefore more direct:
Can an agent re-examine the original visual evidence later, instead of being forced to trust its first interpretation?
The framework's answer is yes.
Every frame returned by the environment is preserved outside the active context window. When the model needs older evidence, it retrieves the relevant frame and brings only that material back into context.

The researchers describe this mechanism as a form of explicit attention over interaction history.
Instead of automatically feeding the full visual history into every turn, the model decides which earlier evidence is relevant to the question it is currently trying to solve.
This matters because an agent's understanding changes over time.
Imagine that the model sees a small colored block during turn 20. At that moment, it may not understand what the block means.
By turn 80, after discovering more of the game's rules, the model may realize that the old block was an important orientation marker.
If the original frame still exists, the model can revisit it with its new understanding.
If only a summary remains, that evidence may be gone permanently.
VISTA keeps the evidence.

This is one of the paper's most useful conceptual points: memory does not always have to mean a better summary. Sometimes the right memory is the original observation itself.
To preserve and reuse visual experience, VISTA is built around three main components:
Each component handles a different part of the interaction loop.

The first component gives the model direct visual access to the environment.
In the ARC-AGI-3 experiments, the official environment provides a 64×64 visual state. VISTA renders that state as a PNG image and, in the paper's default setup, enlarges it to 512×512 using nearest-neighbor scaling.
This preserves:
The model therefore reasons from an image rather than a serialized text grid.
The important point is not that 512×512 is universally optimal. The paper also studies image scale as an ablation. The point is that the model can directly inspect a visual representation without forcing every object and relationship through a textual encoding step first.
The second component stores what the agent has seen.
After every environment action, VISTA archives all frames returned by the environment, including intermediate animation frames rather than only the final state.
Each frame is indexed by:
This produces an external visual archive that can grow over the course of a long interaction.
The model does not have to decide in advance which visual details are worth remembering.
It can preserve everything first and decide what matters later.
That is why the authors call this a lossless visual memory.
It is “lossless” in the practical sense that the original returned frames remain available for later inspection instead of being represented only by a generated text summary.
The third component lets the model actively retrieve visual evidence.
VISTA's public implementation exposes tools including:
play — execute an environment action.inspect — revisit selected frames or image regions.read_pixels — obtain exact pixel values from a selected image region.history — revisit previous actions and environment results.The inspect tool is especially important.
The model can request:
This makes it possible to answer questions such as:
VISTA combines those tools into a repeated observe-reason-act process.
At the start of each turn, the model sees the current frame and available actions.
It then:

The predict-before-action step is useful because it forces the model to make its current hypothesis explicit.
If the environment behaves differently, that mismatch becomes evidence that the model's current understanding is wrong or incomplete.
Over many turns, the agent gradually builds a better model of the environment without training a new neural network or explicitly synthesizing a complete simulator.
VISTA is visual-native, but it does not eliminate text memory.
The framework uses two text files to maintain continuity:
GUIDE.md — durable knowledge and reusable rules discovered across levels.WORKING.md — temporary state, progress, and plans for the current level.These two files serve different purposes.
GUIDE.md is intended to hold stable understanding: rules that are likely to remain useful.
WORKING.md is closer to a scratchpad: where the agent can keep track of what is currently happening and what it plans to try next.
Describe your idea once, and We0 AI can generate a showcase site, pages, and CMS, then help you attract customers and traffic after launch.
One complete project generation for free registration
Best for trying one complete generation flow and seeing a first project draft quickly.
When the context window approaches its limit, the agent writes a handoff checkpoint before continuing in a fresh context window.
The following state remains available across that transition:
This gives VISTA a hybrid memory system:
Text memory:
GUIDE.md + WORKING.md
↓
compact understanding and current plan
Visual memory:
all archived environment frames
↓
original evidence available for reinspection
The two forms of memory complement each other.
Text is efficient for storing a current hypothesis.
Images are better for preserving evidence that may need to be interpreted differently later.
A key design choice is what VISTA does not do.
Although it stores all historical frames, it does not continuously place all of them inside the model's context window.
That would quickly become expensive and unwieldy.
In the full design, after an action is executed, the model normally receives only the last frame as the next current observation.
Intermediate frames and older observations remain in the external visual archive.
The model retrieves them only when necessary.
This creates a useful division:
Current context
→ only the visual evidence needed now
External visual archive
→ complete historical evidence available on demand
This lets the system preserve far more information than it actively loads into the model at once.
In other words, VISTA separates storage from attention.
The system stores broadly but attends selectively.
Another important part of the work is that VISTA does not introduce a newly trained multimodal model.
The underlying model still performs:
The harness handles:
That means the improvement comes from changing how the model interacts with information rather than changing the model's weights.
This is why the work is relevant to a broader trend in agent research: better scaffolding can reveal capabilities that are already present in a general-purpose model but difficult to use through a minimal interface.
The paper evaluates VISTA on ARC-AGI-3 and three additional visual environments.
ARC-AGI-3 contains interactive visual games in which the rules and objectives are not provided to the agent in advance.
The agent must discover how each environment works through observation and action.
With Claude Opus 5.0, VISTA achieves:
With GPT-5.6 Sol, VISTA reaches:
The paper compares these results with official program-free baselines and several systems that rely on programmatic approaches.
The official minimal implementation reports:
| System | Model | RHAE |
|---|---|---|
| Official implementation | GPT-5.6 Sol | 13.33 |
| Official implementation | Claude Opus 5.0 | 40.68 |
| VISTA | GPT-5.6 Sol | 99.00 |
| VISTA | Claude Opus 5.0 | 100.00 |
The comparison does not mean that the harness alone accounts for every difference between all systems in the table. Different systems can use different interfaces and programmatic strategies.
But the paper's same-model comparisons support the authors' central claim: giving a strong multimodal model better access to its own visual history can dramatically improve long-horizon interactive reasoning.
The team also tests VISTA outside ARC-AGI-3 using the same GPT-5.6 Sol backend.

The reported results are:
| Benchmark | Baseline | VISTA | Notes |
|---|---|---|---|
| GameWorld — success rate | 40.0% | 63.3% | 170 tasks across 34 browser games |
| GameWorld — progress | 62.9% | 79.3% | Fresh-human reference shown as 64.1% in the paper figure |
| AI GameStore | 47.3 | 140.3 | Human median normalized to 100 |
| BabyVision | 41.0% | 63.2% | 39 maze and connect-the-lines questions |
These tasks test different kinds of visual reasoning.
GameWorld and AI GameStore involve interactive browser games.
BabyVision is particularly interesting because its selected tasks are static images rather than long interactive histories.
Even there, the ability to crop a region, enlarge it, and inspect pixel values improves performance.
That suggests active visual inspection is useful not only as memory but also as a way of allocating attention within a single image.
The central contribution of VISTA is not a new benchmark score by itself.
It is a different way to think about multimodal agent context.
Language-agent systems often assume that useful history should eventually be compressed into text.
VISTA argues that this assumption can be harmful in visual environments.
A compressed description answers the question:
What did the model think was important at the time?
A lossless visual archive answers a different question:
What actually happened, and can the model inspect it again now that it knows more?
That difference becomes increasingly important in long-running tasks.
The later a model discovers the rules of an environment, the more useful it can be to reinterpret older evidence.
The paper identifies more embodied and physically grounded settings as a natural direction for future work.
The underlying problem appears in many real-world agent tasks:
VISTA does not prove that its exact design will solve those problems.
The published experiments are still centered on games and visual puzzles.
But the architecture points toward a broader principle: agents acting in dynamic environments may benefit from retaining raw sensory evidence separately from the compact beliefs they derive from it.
The paper has three co-first authors: Qiushi Han, Keya Hu, and Linlu Qiu, with Cathy Wu and Kaiming He as additional authors. All are affiliated with MIT in the paper.
Qiushi “Josh” Han is a PhD student at MIT's Operations Research Center, advised by Cathy Wu.
Keya Hu is a PhD student in MIT's Department of Electrical Engineering and Computer Science. Her research sits at the intersection of language and vision, with an interest in building agents that are more data-efficient and generalize better.
Linlu Qiu is a PhD student affiliated with MIT EECS and CSAIL. Her research interests include natural language processing and machine learning, and she has previously worked with Google Research and Meta FAIR.
Cathy Wu is an associate professor at MIT whose research applies machine learning and reinforcement learning to complex systems, including transportation.
Kaiming He is a professor at MIT EECS and is widely known for major work including ResNet, Mask R-CNN, and MAE. He joined MIT in 2024 after research roles including Microsoft Research Asia and Meta FAIR.
VISTA is a visual harness for multimodal agents developed by researchers at MIT. It preserves raw visual observations outside the active context window and lets the model retrieve, zoom, compare, and inspect earlier frames while reasoning through long tasks.
No. VISTA uses existing multimodal foundation models and changes the surrounding interaction framework. The harness handles visual memory, retrieval, tools, and context continuity while the base model performs reasoning and action selection.
VISTA stores every environment frame returned during interaction, including intermediate animation frames, in its original visual form. The model can later retrieve those frames instead of relying only on a textual summary of what it saw.
The public implementation exposes tools including play, inspect, read_pixels, and history. It also uses GUIDE.md for durable knowledge and WORKING.md as a current-task scratchpad.
With Claude Opus 5.0, VISTA completes all 25 public games and reaches an RHAE score of 100.00. With GPT-5.6 Sol, the published scorecard reports 99.00 RHAE.
Yes. The official GitHub repository is public and licensed under the MIT License. The repository includes the VISTA implementation, Dockerfiles, scripts, environment setup files, and instructions for running supported benchmarks.
Yes. The paper evaluates GPT-5.6 Sol as well, and additional experiments in the technical report test an open-weight GLM-5.3 Flash 320B backend. The framework is designed as a harness around existing multimodal models rather than being tied to one model family.
Not yet as a demonstrated result. The published experiments focus on interactive games and visual puzzles, while embodied physical environments are presented as a future research direction.
VISTA shows that improving an AI agent does not always require training a new model. By preserving original visual observations outside the context window and letting the model retrieve them when needed, the framework gives existing multimodal models a more durable form of visual experience.
That design is especially effective in long-horizon environments where the significance of an earlier frame may only become clear much later. On ARC-AGI-3, VISTA with Claude Opus 5.0 completes all 25 public games and reaches a perfect RHAE score of 100, while the same framework generalizes to browser games and static visual puzzles.
The broader lesson is that agent memory can contain both beliefs and evidence. Text notes preserve what the model currently thinks; lossless visual memory preserves what the environment actually showed.
VISTA's key idea is simple: do not force a multimodal agent to remember only its interpretation of the past—let it look at the past again.
Start from one sentence and have a complete website in minutes.