The brain uses place cells to link sensory experience to locations, and grid cells to track position through movement. Most models explain one by assuming the other already exists.
Our model develops both representations through a single objective: predicting the next sensory observation, without predefined spatial codes. It also reproduces grid fragmentation in hairpin mazes and develops locally ordered grid-like fields alongside localized place-like fields in 3D.
We systematically swept over 1,000 training configurations, varying sensory masking, recurrent noise, activity timescales, and random seeds. We also tested how the learned representations change across open arenas, hairpin mazes, separate and connected rooms, and 3D environments.