Local AI · ComfyUI · Wan2.2 · Claude Code · MCP · Image-to-Video · Camera Movement · Spatial Consistency · Vesper · Local AI · ComfyUI · Wan2.2 · Claude Code · MCP · Image-to-Video · Camera Movement · Spatial Consistency · Vesper · Local AI · ComfyUI · Wan2.2 · Claude Code · MCP · Image-to-Video · Camera Movement · Spatial Consistency · Vesper · Local AI · ComfyUI · Wan2.2 · Claude Code · MCP · Image-to-Video · Camera Movement · Spatial Consistency · Vesper · Local AI · ComfyUI · Wan2.2 · Claude Code · MCP · Image-to-Video · Camera Movement · Spatial Consistency · Vesper · Local AI · ComfyUI · Wan2.2 · Claude Code · MCP · Image-to-Video · Camera Movement · Spatial Consistency · Vesper ·

I’d seen a still image turned into a cinematic camera move. I wanted to know how much of it I could reproduce locally.

The inspiration came from a video by Jason Lee. He used Higgsfield to turn generated architectural images into cinematic camera moves for the web.

I loved the result. My immediate question was: could I get somewhere similar using local AI instead?

That question became Vesper — a fictional cabin-rental site simple enough that I could focus on the production workflow rather than the product.

From prompt to page.

I didn’t simply copy the workflow. I changed one assumption: run the generation locally.

I connected Claude Code to ComfyUI through MCP, with the models running locally on my machine.

Claude could work on the site and orchestrate ComfyUI from the same coding environment.

That was more interesting than MCP itself: code, prompts, generated assets and the website had become one production environment.

  1. Claude Code
  2. MCP
  3. ComfyUI
  4. Local model
  5. Vesper

What was actually generating the movement?

Give it a still image and a camera instruction and it can generate a remarkably convincing short sequence. But there is no camera. And there is no 3D cabin behind the image.

I chained several of those sequences together, using the end of one as the starting point for the next, then assembled them into one continuous journey.

That distinction became the whole experiment: the model wasn’t travelling through a cabin. It was repeatedly inventing what might exist beyond the frame.

And for a while, the illusion worked.

Small camera moves worked. Atmospheric movement worked. Some of the individual transitions looked fantastic.

Put them into the website and the effect was surprisingly convincing.

Then I asked the camera to travel further. That’s where it got interesting.

One of the generated moves used in the final flythrough.

Next time, I’ll design the space before generating the journey.

Next time I’d start with the world, not the image: a floorplan or simple 3D blockout, the relationship between rooms, and the path I want the camera to travel.

Then I’d let AI generate around those constraints. It can still invent the pixels. It just doesn’t have to invent the architecture at the same time.

First attempt
  1. Image
  2. Camera instruction
  3. Video
  4. Next image
  5. Video
  6. Assemble
Next attempt
  1. Floorplan / simple 3D world
  2. Camera path
  3. Key viewpoints
  4. Generated frames
  5. Video transitions
  6. Assemble
  1. 01

    AI can create motion without understanding the world.

    A sequence can look convincing even when the world behind it doesn't exist.

  2. 02

    Camera movement turns image generation into a spatial problem.

    The moment the camera moves, an image problem becomes a spatial problem. The further it travels, the more of the world has to remain true.

  3. 03

    Constraints give generative models something coherent to build around.

    Constraints aren't the enemy of generation. They're what give it something coherent to build around.

  4. 04

    The workflow matters as much as the model.

    Putting code, generation and iteration into one workflow made the experiment dramatically faster.

Local was not the limitation I expected.

I started out testing local versus hosted AI. That turned out to be the less interesting question.

Local wasn’t the constraint. Spatial understanding was.

Vesper

The fictional cabin experience I built while exploring the workflow.

Explore Vesper →
Built locally with ComfyUI, Wan2.2, Claude Code and MCP.
Vesper cabin rental website hero — misty lakeside cabin with booking panel

The model can generate the frames.

Someone still has to direct the camera.