Zhaoze Wang

I am a Ph.D. student at the UPenn GRASP Lab. I study predictive world models and compact latent representations for multimodal sequence learning, long-horizon prediction, and planning, drawing inspiration from biological memory and navigation.

I previously interned at Adobe Research, where I worked on video world models, with a focus on efficient generation and long-horizon prediction.

I am seeking research internships from Summer through Fall 2027, with availability for a continuous internship across both seasons.

Email  /  Google Scholar  /  GitHub  /  LinkedIn  /  X

profile photo

Research Interests

The brain's ability to build internal predictive models is a key inspiration behind world models. My research asks how such representations can support not just prediction, but fast updating and replanning as observations and goals change.

I study compact predictive representations, recurrent memory, and learned dynamics, drawing on how the brain integrates sensory experience to guide action. My goal is to translate these principles into fast, adaptive robotic systems.

I study how sensory sequences shape spatial representations [1] and how structured predictive states emerge [2]. I also work on memory-guided planning [3] and efficient video world models. To support model training, I develop compact visual encoders [4] and parallel sensory simulation [5].

Selected Projects

TinyWAM: A Tiny World Action Model
Ongoing project · 2026

A compact-token world-action model for goal-conditioned robot motion, with experiments in simulation and optimized inference on a single GPU.

2023–Present

GPU-parallel simulation combines multi-agent navigation and neural sensor models through Gym-style APIs.

Publications

A Simple Model of Co-Emergence of Grid and Place Fields
Zhaoze Wang, Genela Morris, Dori Derdikman, Pratik Chaudhari†, Vijay Balasubramanian†
NeurIPS 2026

Self-supervised sensory prediction learns spatial representations from experience and motion.

REMI: Reconstructing Episodic Memory During Internally Driven Path Planning
Zhaoze Wang, Genela Morris†, Dori Derdikman†, Pratik Chaudhari†, Vijay Balasubramanian†
NeurIPS 2025

A recurrent model combines cue-triggered goal retrieval, route planning, and sensory reconstruction.

Time Makes Space: Emergence of Place Fields in Networks Encoding Temporally Continuous Sensory Experiences
Zhaoze Wang†, Ronald W. Di Tullio†, Spencer Rooke, Vijay Balasubramanian
NeurIPS 2024

Reconstructing noisy or partially observed sensory sequences learns spatial representations.

Trading Place for Space: Increasing Location Resolution Reduces Contextual Capacity in Hippocampal Codes
Spencer Rooke, Zhaoze Wang, Ronald W. Di Tullio, Vijay Balasubramanian
NeurIPS 2024 Oral

Geometric analysis quantifies the trade-off between spatial precision and memory capacity.

Projects

BtnkMAE image reconstruction results
BtnkMAE: Compact and Decodable Visual Representations
2023–Present

A bottleneck masked autoencoder compresses images into compact embeddings that can be decoded back into images.

NN4N neural network simulation
NN4N: Neural Networks for Neurosimulations
2023–Present

Modular recurrent networks with biologically inspired connectivity support flexible, multi-region neural simulations.

Teaching & Reviewer

Service

NeurIPS 2025 Reviewer

NeurIPS 2026 Reviewer (Top Reviewer)

Teaching Assistant, ESE 5460: Principles of Deep Learning, Fall 2025

Teaching Assistant, PHYS 5585: Comp. and Theoretical Neurosci., Spring 2026


Design and code of this website is adapted from Jon Barron's website.