
Other teams have done this, too, and several other tools have been used to recreate images based on brain scan data. But they’re not good enough, says Irani. Say a person saw a banana. These models can generate an image of a banana, but it would look different, she says. “It wouldn’t have the same structure, the same position.”
A better decoder
The team wanted to more closely recreate the images that had been seen. The first step was to train an AI model on already available data from eight people who each had been shown around 9,000 images while in a high-resolution fMRI scanner.
Crucially, their “brain decoder” has two branches—one to predict the structure of an image (where the colors are, for instance) and a second to predict its content (for example, a bunch of bananas on a plate). The predictions allow a diffusion model, a type of AI best known for creating video and images by gradually cleaning up a noisy mess of pixels, to produce a much more accurate representation of what the person saw.
But to improve the models they needed more data—far more than was actually available.
To get around this problem, she and her colleagues trained another model in the other direction—an encoder that can predict brain activity from an image. The team then used the encoder and decoder together to improve both tools.
It works like this: start with a new image, say, of a leopard. Then use the encoder to predict what the fMRI brain scan of a person would look like when they saw that picture. The decoder is then used to reconstruct the image again. At first, that image probably won’t look much like a leopard, says Irani. But repeatedly training the models this way eventually leads to dramatic improvements.
This approach also allows the team to train their models on as many images as they want, even though they might never have been shown to a person in an fMRI scanner. Irani says that around 70% of the training data is from images that were not originally paired with fMRI scans.
By combining data from multiple studies, they were also able to identify brain regions that seem to share functions across all individuals. One region seemed to respond to images of food, for example, while another responded to images of sports. Irani, a computer scientist, says she is now working with neuroscientists “to see if we can actually use these tools that we’ve developed to really find out new things about the brain.”







