NEWS
Brain-IT Rebuilds Seen Images After One Hour of Scans
Weizmann’s Brain-IT matches 40-hour fMRI reconstructions after one hour on a new person, while 15-minute scans already look like a scene.
Brain-IT, a Weizmann Institute model, matches 40-hour fMRI image reconstructions after one hour of scans on a new person. The system rebuilds a picture someone is viewing, not a private thought, and it does that from blood-flow maps collected in a magnet.
The viral baseball and ski-slope pairs are the demo. The change that labs can actually book is cheaper calibration on a codebook trained across eight brains.
One Hour on a New Brain Matches 40-Hour Models
Roman Beliy, Amit Zalcher, Jonathan Kogman, Navve Wasserman, and Prof. Michal Irani, all in Weizmann’s Department of Computer Science and Applied Mathematics, built Brain-IT around a Brain Interaction Transformer. The ICLR 2026 paper says that with only 1 hour of fMRI from a new subject, the method posts results comparable to 40-hour recordings used by earlier systems.
MindEye2, led by Tanishq Mathew Abraham, already offered a 1-hour mode in 2024. On that same budget, Brain-IT’s CLIP score is 93.0%, the same figure MindEye2 posted after the full 40 hours. Pixel correlation at 1 hour is 0.331 for Brain-IT and 0.195 for MindEye2, averaged on Natural Scenes Dataset subjects 1, 2, 5, and 7.
1 HOUR VERSUS 40 HOURS, SAME TESTS
| Method | Hours | PixCorr | SSIM | CLIP |
|---|---|---|---|---|
| MindEye2 | 40 | 0.322 | 0.431 | 93.0% |
| Brain-IT | 40 | 0.386 | 0.486 | 96.4% |
| MindEye2 | 1 | 0.195 | 0.419 | 79.2% |
| Brain-IT | 1 | 0.331 | 0.473 | 93.0% |
On the full 40-hour set, Brain-IT wins 7 of 8 standard scores against a field that includes MindEye2, MindTuner, NeuroPictor, and Brain-Diffuser. Irani put the gap in plainer words in the institute’s September 14, 2026 note: older models “tend to make mistakes in basic features such as composition and color.”
Eight People, 73,000 Pairs, and a Borrowed Scanner
There is still almost no public fMRI paired with pictures. The team did not scan a new cohort. It trained on the Natural Scenes Dataset, in which eight subjects viewed thousands of scenes inside a 7T scanner at the University of Minnesota’s Center for Magnetic Resonance Research.
WHAT THE NSD SCANS ACTUALLY CONTAIN
- The pairs: Weizmann counts about 73,000 image-fMRI pairs across the eight people.
- The grind: Each person sat through 30 to 40 sessions, with 6 scans of about 10 minutes and about 40 pictures per scan.
- The magnet: Whole-brain 7T scans at 1.8-mm resolution, sampled every 1.6 seconds, using stills drawn from the COCO photo set.
That is a large neuroscience set and a small AI set. Irani’s group treated the shortage like a two-way dictionary. An encoder predicts the scan a picture would produce. A decoder turns a scan back into a picture. Fed images that never went into a magnet, the loop manufactures extra training pairs. Irani said about 70% of the training material came from pictures that were never paired with a real scan.
How Brain-IT Maps 128 Shared Clusters
The same picture lights up different brains in different places. The encoder’s workaround was to split each scan into about 40,000 voxels, then group voxels that behave alike into 128 functional clusters shared across people. Those clusters, not whole-brain fingerprints, are the units the transformer talks to.
During training, the encoder naturally identified 128 functional regions that are shared by all people and perform specific roles in image processing. Some of them are familiar to neuroscientists, but others are entirely new. For example, we discovered a division of roles within the brain region that processes images of places, the PPA, with one part responding to indoor scenes and another to outdoor scenes.
Michal Irani, professor, Weizmann Institute of Science
Irani added that the encoder can match regions that do the same job even when their anatomical seats differ. Color versus black-and-white versions of one picture, run through the encoder, are meant to show where color lives in the predicted scan.
FROM VOXELS TO A PHOTOGRAPH
- Voxel-to-cluster map: Every voxel in every subject is assigned to one of the 128 shared clusters.
- Brain tokens: Each cluster becomes one token. A cross-transformer mixes those tokens and writes local image features.
- Two branches: One branch predicts VGG structure and inverts it with a Deep Image Prior into a coarse layout. The other predicts CLIP-style meaning and steers a diffusion model.
- New person: Most weights stay frozen. Only voxel embeddings are fit, which is why a short scan can be enough.
The institute says that design lets the model learn to read a new person after one hour, where speech brain-computer interfaces for paralysis still train for tens of hours on one patient. Brain-IT is doing a different job. It is matching a viewed still, not spelling a sentence.
Fifteen Minutes Already Looks Like a Picture
The paper’s sharper claim sits under the 1-hour line. On subject 1, transfer learning with 15 minutes of data (450 fMRI samples) already produces reconstructions from 15 minutes the authors call meaningful, with pixel correlation 0.336 and CLIP 91.3%. Thirty minutes on that same subject reaches pixel correlation 0.378 and CLIP 93.3%.
Abraham, whose MindEye2 paper set the prior 1-hour mark, wrote that the Weizmann work was “getting reconstructions that are not complete nonsense from ONLY 15 min of data.” He pointed to the shared voxel clusters and to synthetic scans predicted from COCO photos as the tricks that stretch eight brains this far.
That is still one subject in a 7T research magnet, looking at still photographs on a schedule. It is not a clinic slot, and it is not a thought. It is the first time this literature has put 15-minute outputs next to the old 40-hour pictures and asked a reader to compare them.
A Cake Comes Back as Three Sandwiches
The reconstructions can look photographic and still be wrong in the way a confident generator is wrong. Irani walked through misses on a video call: a cake came back as a pile of three sandwiches, and a dog in a bathtub came back as a goat of the same color in a tub.
Of course we have failures. All in all, really we outperformed the others by a significant margin. Mind reading is a cute, jazzy name for what they’re doing.
Michal Irani, professor, Weizmann Institute of Science
Diffusion models fill gaps with plausible pixels. Brain-IT’s low-level branch is there to pin layout and color before that fill-in starts, which is why a baseball diamond can survive as a diamond instead of a generic sports collage. The sandwich cake is the remainder: semantics close enough, object identity not.
The method also only sees what the magnet sees. fMRI tracks oxygenated blood, not spikes from single cells, and Irani noted that a scan takes about two seconds while video throws dozens of frames in that window.
Scanner Time Still Runs Hundreds of Dollars an Hour
Tommy Sprague, a neuroscientist at the University of California, Santa Barbara, put a price on the old protocol. “None of us can afford 40 hours of imaging for a new subject,” he said. “It’s something like $600 to $1,000 an hour.” At those rates a 40-hour series runs $24,000 to $40,000 before analysis. A 1-hour fit is $600 to $1,000. Fifteen minutes is a fraction of one booked slot.
That arithmetic is why a shared codebook is the product. A lab that could pay for one heavily scanned volunteer can now try the same decoder on many more people, which is the step that makes locked-in communication research thinkable and makes retention, reuse, and consent of neural data a live file.
The hardware has not moved. The person still has to lie still in a 7T research scanner, look at pictures, and give a full hour if the lab wants the 93.0% CLIP band. Judy Illes, a neuroethicist and professor of neurology at the University of British Columbia, called the work “magnificent” and pointed to people with neurologic conditions. The same hour that helps a study also produces a reusable map of how that person’s visual cortex answers a photo.
The authors posted code and pretrained checkpoints for the encoder, the decoders, and the combined diffusion model. Replication no longer requires guessing the 128-cluster map.
Video, Audio, and the EEG Problem
Irani’s lab is already pushing past stills, toward sound and toward video. Dream decoding is the line that travels. “What remains especially challenging is decoding video, for example, during dreaming,” she said. “Dozens of images change every second, while an fMRI scan takes about two seconds. If we overcome all these obstacles, it’s possible that in the future we may even be able to read dreams.”
She has also said the dream version is not in hand. “That’s something we don’t have yet. But we’re striving to achieve it.” Sprague thinks the present encoder-decoder would likely work quite well on pictures a person is imagining rather than viewing, which is a shorter hop than dreams and a longer hop than the baseball demo.
WHERE EXPERTS DISAGREE
- Illes: The therapeutic use for neurologic conditions is the draw, and she called the reconstructions magnificent.
- Sprague: Covert extraction of what someone is thinking is the sci-fi risk, and EEG versions of the same idea deserve more serious ethics than a 7T lab study.
- Ienca: Marcello Ienca, a neuroscientist and philosopher at the Technical University of Munich, called a move to EEG a game changer, because a cap calibrated to one user could later be mined for extra information, and some courts might try to treat reconstructed images as evidence.
Irani said she is not worried about fMRI in its current form. On EEG misuse she was brief: “I’m trying to think only of good things.”
Frequently Asked Questions
How Does Brain-IT Reconstruct a Seen Image?
Two Brain Interaction Transformer models run in parallel, one writing CLIP-style meaning and one writing VGG structure, then a Deep Image Prior turns the VGG features into a coarse photograph that initializes diffusion. Most older decoders crush all voxels into one global embedding first, which is the step Brain-IT skips by sending each of the 128 clusters straight to local image patches.
How Much Scan Time Does a New Person Need?
The published 1-hour scores are the headline comparison, but the supplement on subject 1 also reports 15 minutes (450 samples) and 30 minutes, with only voxel embeddings updated while the shared transformer stays frozen. That freeze is why a new brain does not need its own 40-hour campaign.
What Is the Natural Scenes Dataset Used by Brain-IT?
NSD is a 7T, 1.8-mm, 1.6-second sampling project collected at Minnesota’s CMRR, with pictures from COCO and a continuous recognition task in which each person judged whether a still had appeared earlier in the year-long study. Eight screened adults, two men and six women, ages 19 to 32, completed the core set.
Can Brain-IT Reconstruct Thoughts, Memories, or Dreams?
No in the paper as written: every scored picture is one the person was looking at in the scanner, and the blood-oxygen signal itself peaks several seconds after a still appears. Locked-in speech systems that restore words still train on tens of hours from that one patient and do not share Brain-IT’s 128-cluster map.
What Did the 128 Shared Clusters Reveal?
Besides the indoor/outdoor split inside the parahippocampal place area, Irani said the encoder also isolated clusters that answered food pictures and clusters that answered sports pictures, including some seats neuroscientists did not already treat as separate tools. The 2024 companion encoder paper, “The Wisdom of a Crowd of Brains,” is where those voxel embeddings were learned.
The paper is public, the checkpoints are public, and the magnet is still the appointment you have to keep. Until someone ports the same 128 clusters onto a cap, the pictures come from people who booked the scanner and looked at the stills on purpose.
-
BUSINESS1 month agoMicron’s $22 Billion Deposits Cover Only a Fifth of DRAM
-
NEWS1 month agoOnePlus 16 Bets on 200MP and a Familiar 3x Sony
-
NEWS4 weeks agoC Spire Plans a 200-Mile Fiber Line for One Customer
-
NEWS1 month agoDebian Puts Generative AI Risk on Volunteer Submitters
-
AUTO1 month agoCNG and Hybrids Push Alternative Fuels Past Petrol
-
TECHNOLOGY3 years agoHow to Change Background Color in Epic Hyperspace – Personalizing Your EHR Interface
-
LIFESTYLE3 years agoTypes of Alligators in Florida – Discovering Reptilian Diversity in the Sunshine State
-
LIFESTYLE3 years agoDo Jehovah Witness Go to Church on Saturday – Understanding Sabbath Practices in Jehovah’s Witness Faith
