Captivating Your Audience
The Spatiotemporal Structure of VR Cinematic Narrative
Captivating Your Audience

It is well-understood that VR cinema is an open spatial medium. As the audience steps inside, a more fundamental question emerges: in a world without a “frame,” how do we fulfill the narrative?

In traditional cinema, the director controls everything through cinematic language—focal length, editing, and blocking—precisely guiding the viewer’s gaze. In VR, however, the perspective is unleashed to a full 360 degrees; for the first time, the audience possesses the right to “choose where to look.” And that is exactly where the problem begins. They are left unsure of where to look, what to look at, or how to look at it. Consequently, the narrative becomes fragmented, and traditional film language essentially fails.

This forces creators to undergo a fundamental internal shift: moving from two-dimensional image-based thinking to three-dimensional spatial thinking. Simultaneously, we are no longer “controllers”; we become “experiencers” co-existing with the audience within the same field.

Our early solutions were far from ideal:

  1. Pausing the narrative until the audience’s attention returns to the plot line. — This breaks the rhythm and slows the narrative to a crawl.
  2. Using motion or mechanics to lead the audience into participating in events and choices. — This leans too heavily toward gamification and places high demands on hardware.
  3. Relying on visual cues to direct the viewer’s attention. — This interferes with the aesthetic expression.

These methods, in essence, are antagonizing the audience” rather than understanding the audience.”

In my view, truly effective guidance always arises from three levels: the visual, the auditory, and the somatosensory. In other words, it is not about telling the audience where to look, but rather allowing them to be naturally drawn toward it.

This implies an entirely new narrative logic: one based on perception rather than control; on participation rather than mere observation.

Therefore, the primary core proposition of VR cinema is not how to tell a story well, but rather: How to construct a unified frequency of time and space based on human cognitive, perceptual, and decision-making mechanisms, thereby driving attention itself.

This is no longer just a cinematic issue; it is an interdisciplinary one. We require the collective participation of iconography, psychology, and information science to reconstruct the very act of narrative.

What we must truly create is not a world to be watched, but a world to be co-perceived.

Perhaps one day, we will no longer ask, “What did you see?” Instead, we will ask: “Did we see the same world, at the same moment?”

And perhaps, only on that day will VR technology truly become an indispensable foundation of cinematic language.