Project Lola / Star Wars, frame by frame
Refik Anadol Studio · Sep 2026 · v0.3
Corpus study6 films13,595 shots36,871 frames

The saga, read by machines

What does a film look like when every frame becomes a number?

Night run · 17 Sep 2026Six Star Wars films were cut into shots, sampled three frames per shot, and passed through a stack of models: color, faces, scene semantics, SigLIP embeddings, motion and depth. On top of the same frames, two generative models were trained: a Flux style LoRA and a from-scratch diffusion transformer that takes a frame's embedding as its only instruction.

FilmsI · II · III · IV (Blu-ray + 35mm) · VI
MachineRTX 4090 · one night
Statusv0.3 · analyses, LoRA, DiT and refined walk
I

Corpus · Act I

6 films13,595 shots · 36,871 frames1080p proxies · 3 frames per shot

One decode pass per film, then everything reads from the proxy. Shots come from PySceneDetect; each shot gives three frames at a quarter, half and three quarters of its length. The 35mm scan of Episode IV sits next to the Blu-ray as a control: same film, different grade.

Cuts per minute, six films stacked
01Cuts / min
Shot rhythm
01 / 06RhythmShot statistics

How fast the saga cuts

Average shot length, film by film

Every film ends faster than it begins.

What the numbers say

Average shot length runs from 3.1 s in A New Hope to 4.4 s in Attack of the Clones. The two prints of Episode IV give almost the same cut profile, which is the first sanity check of the detector: 2,263 shots on the 35mm scan against 2,378 on the Blu-ray, the difference being the Special Edition inserts.

3.15 sShortest average shot: A New Hope, Blu-ray19.1 cuts / min
4.38 sLongest average shot: Attack of the Clones13.7 cuts / min
MethodPySceneDetect · ContentDetector 27min shot 0.5 sshots_*.csv
II

Color · Act II

5 colors per framek-means on 96 px thumbnailsLab brightness · temperature · saturation

Each frame is reduced to five dominant colors and their weights. Laid out shot by shot, the six films become six strips of pigment; read as curves, they become light and temperature over running time.

Palette strips of six films, one row per film, one column per shot
Study 02 · palette strips, top to bottom: IV 35mm, IV Blu-ray, VI, I, II, IIIone column per shot · 5 colors weighted
Color temperature curves per film
02Lab b*
Temperature
02 / 06PigmentColor analysis

Two prints, two colors

The 35mm scan is darker and greener than the Blu-ray

Same frames, different memory of them.

What the strips show

The first two rows are the same film. The 35mm print keeps a low, greenish density; the Blu-ray lifts the blacks and warms the mids. Revenge of the Sith ends in a red block: the lava duel. Attack of the Clones carries the warmest passage of the saga around minute 110, the Geonosis arena.

Blue against red

Counting saturated blue and red pixels gives a crude lightsaber index. The Phantom Menace has the most of both (2.0 % blue, 3.5 % red); Revenge of the Sith is red-heavy (2.3 % red against 0.8 % blue). A New Hope on Blu-ray has almost no saturated blue at all.

Outputsstrip_<film>.pngcurve_L / temp / satcolor.parquet · 36,871 rows
III

Embedding space · Act III

SigLIP so400m1152 dims · DINOv2 as second opinionUMAP 2D and 3D

Every frame becomes a point. Neighbors in this space share what the model considers the same visual idea, regardless of which film it came from. Two things fall out immediately: the eras separate, and the echoes between films become searchable.

UMAP maps, one per film, highlighting where each film sits in the shared space
Study 03 · the same map, one film lit at a timeUMAP · cosine · n_neighbors 30
UMAP of all frames colored by era
03Era
Original · prequel
03 / 06DriftVisual language

Two continents

Original trilogy and prequels barely touch

Twenty years of cinema, visible as a gap.

Reading the map

The prequels (orange) and the originals (blue) occupy separate regions with a thin seam between them. Film grain, lens, digital sets and CGI density all push the same way. The two prints of A New Hope land on top of each other, so the gap is not a grading artefact.

Echoes

For each frame, the nearest frame in every other film is recorded. Between The Phantom Menace and Return of the Jedi the best matches are exactly the recurring bodies: R2-D2, Yoda, Jabba. The mean similarity of the top 200 echoes per film pair puts the two Episode IV prints at 0.98 and Sith against Clones at 0.97.

FilmSpaceShip interiorDesertNatureCityDuel
A New Hope (35mm)11.421.611.02.28.011.3
A New Hope9.219.811.01.810.812.7
Return of the Jedi5.06.96.135.26.69.5
The Phantom Menace3.26.522.912.121.012.6
Attack of the Clones6.66.714.36.722.120.1
Revenge of the Sith10.810.52.55.819.426.0

Share of frames per zero-shot scene class, %. Text prompts scored against SigLIP, centered per class so that people-heavy prompts do not absorb everything. Endor makes Return of the Jedi the nature film; Tatooine and Naboo split The Phantom Menace.

Contact sheets · the most confident frames per classdesert · space · duel
Desert class contact sheet
Desert
Space class contact sheet
Space
Duel class contact sheet
Duel
Echo pairs between The Phantom Menace and Return of the Jedi
03bEchoes
I ↔ VI
03bPairsNearest frame across films

The same shot, twenty-two years apart

Echo pairs between Episode I and Episode VI

How pairs are chosen

Dark frames and the first two minutes and last eight minutes of each film are excluded, so logos and credits do not count as echoes. Each shot appears at most once per side. Left: The Phantom Menace. Right: Return of the Jedi.

Outputsechoes.csvpairs_<A>_<B>.jpg · 15 pairs of filmsmatrix.png
IV

Faces · Act IV

26,670 facesInsightFace detection · ArcFace identity2,947 clusters · ViT expression

Faces are detected in every frame and embedded for identity. Clustered without any labels, the top clusters of each film are its cast, and the same cluster id follows an actor from one film to the next.

Face clusters of A New Hope: eight rows, twelve faces each
Study 04 · A New Hope, eight largest clusters, unlabeledLuke · Obi-Wan · Han · Leia · Tarkin · Chewbacca · pilot Luke · Owen
Face clusters of Revenge of the Sith
04Sith
Clusters
04 / 06IdentityFace analysis

Who is on screen

Clusters, close-ups and expressions

Nobody told the model who Luke is. It found him anyway.

Cluster stability

Cluster #117 is Luke in both prints of A New Hope; #30 and #157 carry over into Return of the Jedi. Agglomerative clustering on cosine distance, threshold 0.55, average linkage. The long tail of 2,947 clusters is background faces and single-shot extras.

FilmAngryHappyNeutralSad
A New Hope (35mm)11.510.836.536.1
A New Hope11.012.229.443.7
Return of the Jedi19.010.021.344.9
The Phantom Menace14.010.124.149.2
Attack of the Clones11.710.925.049.7
Revenge of the Sith14.49.119.654.5

Share of detected faces per expression, %, ViT face-expression model, confidence above 0.5. The model leans sad on dim cinema faces; read the ranking, not the absolute numbers. Revenge of the Sith is the darkest film by this measure.

Open

Clusters still need names. Once the top eight per film are labeled, screen time per character and the ageing drift of one actor across decades come for free from the same embeddings.

V

Motion and depth · Act V

Optical flow per shotsymmetry · thirds · center of detailDepth Anything v2

Composition scores come from the frames themselves; motion energy from two consecutive frames of each shot; depth from a monocular model on a sample. Together they describe how the camera behaves rather than what it sees.

Motion energy heatmap per film over running time
Study 05 · motion energy, dark is still, bright is fastFarneback flow · rolling 15 shots
Depth map contact sheet
05Depth
Depth Anything v2
05 / 06CameraComposition

Where the action is

Set pieces light up as motion

The trench run is a bright bar at the end of a dark film.

Reading the heatmap

A New Hope brightens only in its final twenty minutes. Revenge of the Sith opens at full speed with the space battle over Coruscant. The Phantom Menace peaks around the podrace, minutes 60 to 70. The 35mm print reads slightly hotter than the Blu-ray because its grain adds flow noise.

FilmSymmetryCenter dist.Edge densityMotion
A New Hope (35mm)0.7060.0660.0872.85
A New Hope0.7480.0780.0812.12
Return of the Jedi0.7680.0790.0772.34
The Phantom Menace0.6770.0750.0812.55
Attack of the Clones0.7740.0750.0782.22
Revenge of the Sith0.7600.0810.0782.56

Means per film. Symmetry is horizontal mirror similarity (1 = mirror); center distance is how far the center of detail sits from the frame center; motion is mean flow magnitude in pixels on a 480 px proxy.

Also produced9,389 empty-world backgrounds · SAM2 + inpaintcomposition.parquetdepth/*.npy
VI

Generative · Act VI

Two modelsFlux.1-dev LoRA · 75 minLightningDiT from scratch · SigLIP conditioned

The same frames train two very different models. The LoRA borrows a large model's world and teaches it the saga's look. The DiT knows nothing but these 35,590 frames, and instead of text it takes a frame's embedding as its instruction, which makes the embedding map from Act III a navigable space.

Saga traversal · 30 keyframes, Episode I to VI, condition slerp with fixed noise, refined frame by frame with the Flux LoRA58 s · 696 frames · DiT step 60,000 + img2img 0.5
Flux LoRA · samples at step 2000, trigger-only captionsrank 16 · lr 4e-4 · 1,998 frames
Generated: a lone figure walks across dunes toward two suns
Dunes, two suns
Generated: snow plain at dusk with a dome
Snow plain, dusk
Generated: close-up of a hooded old man
Hooded old man
DiT preview grid at step 20,000: six reference frames, eight samples each
06DiT · step 20k
Rows: reference frames
Reading the gridEach row is one reference frame's SigLIP vector; the eight columns are different noise. The Fox logo, a Tatooine crowd, Vader in a white corridor and the credit crawl appear by step 20,000.
06 / 06ModelsTraining

A model that only knows six films

Embedding in, frame out

Point at any place on the map and ask it to dream there.

Architecture

LightningDiT-B/1 on VA-VAE latents (f16, 32 channels, 256 px). The class embedding is replaced by a small MLP over the 1152-d SigLIP vector; a zero vector is the null condition for classifier-free guidance. Batch 128, 60,000 steps, about four hours on one RTX 4090.

Why conditioning on embeddings

Walking in noise space gives seams; the earlier latent-model runs measured peak-to-mean frame change of 1.7 to 3.5 with no fix from more ODE steps. Walking in condition space with fixed noise moves the meaning instead of the noise: interpolate two frames' vectors and the model renders what lies between Tatooine and Hoth.

Keyframes

Six frames per film, chronological, credits excluded.

Slerp

Spherical interpolation between consecutive SigLIP vectors.

Sample

Fixed noise, CFG 3, 50 Euler steps, VA-VAE decode.

Refine + video

Flux LoRA img2img per frame, then 12 fps motion-interpolated to 24, 1024 px.

Status

trained · 60,000 steps · 3 h 54 min · final loss 0.273
The traversal above walks through six keyframes per film in story order. Each 256 px frame from the DiT is then passed once through the Flux LoRA as image-to-image at 512 px (strength 0.5, fixed seed): composition and motion stay with the DiT, skin and texture come from Flux. Faces and places are recognisable on the way: Tarkin, Luke in the cockpit, Palpatine, young Anakin, Jabba's palace, Geonosis, the lava of Mustafar. What lies between two keyframes is what the model invents.

Artifactsflux_lola_v1.safetensorslola_b1_siglip / checkpointscond_siglip.npy · 35,590 × 1152latent-model recipe
What comes next

The six missing films (V, VII, VIII, IX, Rogue One, Solo) plug into the same pipeline. Naming the face clusters unlocks screen time and ageing curves. Cropping the letterbox out of the 35mm scan cleans the DiT's training set. And the two models can be pointed at the map together: LoRA for fidelity, DiT for the walk.

DataT0c

Pull the remaining six films when they land; re-run T1 to T9 over eleven films.

FacesT8b

Label the top eight clusters per film; compute screen time and the drift of one face from 1977 to 2019.

GenerativeT14b

Longer walks with more keyframes, RIFE instead of motion interpolation, and a 512 px model once the 35mm letterbox is cropped out of the training set.

Idea poolIDEAS.md

Leitmotif detection, Aurebesh OCR, lightsaber detection, palette-conditioned repainting, trilogy LoRA weight interpolation.