A focus pull is an editorial event

Focus is not decoration applied after composition. A deliberate pull changes who or what owns the frame. It can reveal the person watching from behind a foreground action, transfer a reaction, disclose an object, delay information or connect two planes without moving the camera. If the audience would understand the shot equally well without the transfer, shallow depth of field may be a look, but the pull is probably not doing a story job.

Runway's current camera-term library gives Gen-4.5 a clear example: begin on a chef's hand garnishing a plate, then transition to the diner across the table. The example is useful because it names both subjects and their order. It does not merely request a cinematic rack focus. The hand has the frame first; the diner receives it second.

That is the minimum useful instruction for any model. Write the attention change before the optical language: start with A, transfer on a visible cue, finish on B, then hold long enough for the editor and audience to read the new information.

Separate depth of field from focus choreography

Shallow focus, deep focus and rack focus solve different problems. Shallow focus limits the zone of acceptable sharpness and separates one plane. Deep focus keeps near, middle and far information legible together. Rack focus changes the selected plane during the shot. A clip can have shallow depth of field without a pull, and it can pull so gently across a broader depth range that the transition never resembles the familiar snap between two pools of blur.

Runway now publishes generated examples for all three terms. Adobe's current Firefly video prompt guidance also treats depth of field as part of the aesthetic description, with shallow depth of field and bokeh as an example. Those interfaces can understand the vocabulary, but vocabulary is only the top layer. It does not define exactly where the sharp plane begins, how quickly it travels, how blur renders or which visual features must survive the change.

For cinematographers, begin by choosing the depth strategy for the shot. Use deep focus when simultaneous information and blocking matter more than isolation. Use stable shallow focus when one subject owns the beat. Use a pull only when a change of attention is the action.

Runway AI video examples comparing a detailed antique shop in deep focus with a chrome bubble in shallow focus
Runway's official deep-focus antique shop and shallow-focus chrome bubble demonstrate two depth strategies before any focus transfer is added. Frames from Runway's Gen-4.5 camera-term library.

Write a six-field focus-beat card

Give each planned pull a small card with six fields. First subject: the exact feature or object that begins sharp. Second subject: the feature or object that ends sharp. Spatial relationship: foreground, middle distance and background position. Cue: the look, hand movement, line, entrance or sound that motivates the transfer. Timing: when the pull starts, how long it travels and how long the endpoint holds. Protection: the things that must not change while focus moves.

A useful card might read: foreground key in a gloved hand begins sharp; background actor at the doorway finishes sharp; actor turns on the lock click at 1.8 seconds; focus travels over 1.2 seconds; hold the actor for 2 seconds; locked camera; key, glove, doorway, actor identity and exposure remain stable. That is specific enough to prompt, storyboard, review or rebuild conventionally.

Avoid fake precision. A text-to-video model is not a calibrated lens motor, and a prompt containing 1.8 seconds does not prove frame-accurate execution. The timings are acceptance targets. They help the operator reject a transfer that happens before the cue or never gives the endpoint time to register.

Design two readable subject planes

A model cannot create a strong transfer if the composition does not contain two legible places for attention to land. Separate the subjects by depth, silhouette, colour, exposure or movement. Keep the second subject visible enough in the opening that the audience can accept it as part of the space, while reserving clarity or action for the moment it receives focus.

Protect faces, hands, product edges, text, reflections and repeated background shapes at both endpoints. An input image can carry the opening composition and lighting, but it may also lock in weak depth separation. Text-to-video gives more freedom to invent the arrangement, with more risk that subject identity or geography shifts during the transfer. Choose the control mode by the part of the shot that cannot drift.

Google DeepMind's Veo prompt guide separates shot framing and motion, lighting, location and action as useful prompt elements. Apply the same separation here. Describe the static spatial design first, then the action cue, then the focus transfer. If a camera move is also essential, decide whether it supports the same attention change. A dolly, head turn, prop action and focus pull competing inside five seconds is four uncertain systems, not one elegant shot.

Judge the transition, not only the endpoints

Real focus behaviour is continuous. In its March 2026 discussion of Aatma cine-lens design, ZEISS describes focus pulling as part of composition and storytelling, then distinguishes the way different optics move into and out of focus. Focus roll-off can feel abrupt or gradual. Skin, texture and out-of-focus highlights change through that movement. Those qualities are visible over time, not captured by two sharp stills.

Review the whole path at speed and one frame at a time. Does the foreground release before the background is ready, leaving a muddy interval with no visual owner? Does the face sharpen in patches? Do hair, teeth, jewellery or product markings reform as clarity arrives? Do bokeh shapes pop, crawl or change size without a spatial reason? A clean first and last frame can conceal a broken transition.

Do not reject every synthetic irregularity because it differs from a mathematically perfect lens. Optical character can be expressive. Reject it when the behaviour distracts from the story cue, changes the subject, damages continuity or cannot be repeated closely enough across the sequence.

ZEISS optical designer Xiang Lu holding an Aatma cinema lens beside a camera
ZEISS optical designer Xiang Lu discusses Aatma at 85 mm and T1.5. The project treated focus roll-off, skin rendering, bokeh and breathing as moving-image behaviour rather than still-frame attributes. Official image via ZEISS, 30 March 2026.

Check breathing, reframing and invented camera movement

A focus pull should not quietly become a push-in. Physical lenses can change angle of view as focus distance changes, a behaviour known as focus breathing. Cine lenses are often designed to control it: ARRI highlights minimal breathing while racking focus on its Enso Primes, and ZEISS lists breathing control among the modern characteristics carried into Aatma. The amount and character depend on the lens, but the audience should not mistake an unwanted frame-size change for camera intent.

Generated video can produce its own version of the problem. The model may enlarge the destination subject, warp the background, slide the camera or reshape the foreground as it tries to signal attention. Freeze the endpoints and overlay them. Track fixed architecture and frame edges. If the whole scene changes scale, log reframing or synthetic breathing separately from the focus transfer.

Runway notes that even a requested static shot can acquire subtle movement because video models are biased toward motion. Reinforce the lock in plain language: the camera remains entirely motionless; framing and perspective do not change; only the focus plane transfers. If the chosen result still drifts, stabilisation may repair a slight translation, but it cannot restore geometry the model has redesigned.

Review focus across the cut

A successful pull can still be wrong for the sequence. Put the generated clip between its intended neighbours, mute it and watch at speed. The incoming shot should establish where attention begins. The pull should happen on the planned beat. The outgoing shot should inherit the audience's attention or deliberately break it. If the edit arrives before the endpoint settles, the generation spent time on a focus result nobody can read.

Match depth logic across coverage. A wide with deep focus, a close-up with an extremely thin plane and a reverse with soft focus can belong together when the emotional design supports the changes. They fail when the depth treatment wanders only because each shot was generated independently. Record depth strategy, starting plane, ending plane, cue and endpoint hold beside the shot ID.

Review faces and products at delivery size, not only in a small playback window. AI can imitate optical blur while using blur to hide unstable detail. The frame that appears to come into focus must actually resolve the approved identity, geometry and claim.

Buy repeatable attention control

The buyer implication follows the craft. Ask an AI video supplier for a five-shot focus test built from the real format: one stable shallow-focus portrait, one deep-focus composition, one foreground-to-background pull, one background-to-foreground pull and one moving subject that must remain sharp. Supply named cues and protection rules before generation.

Score attention timing, endpoint accuracy, identity stability, geometry stability, breathing, unwanted camera motion, transition artefacts, cut compatibility and human repair time. Record accepted attempts against total attempts. A beautiful rack-focus example is not proof that the workflow can place the transfer on a line reading, preserve a pack shot or match a second angle.

When exact focus timing carries a performance, product claim or editorial reveal, keep the option to build the shot with conventional capture, animation, depth-aware compositing or a controllable 3D camera. AI video is useful when it can preserve the focus decision at the required cost and speed. The cinematographer's value is defining that decision clearly enough that a generated blur transition is never mistaken for the finished job.

Build

Make the focus decision survive the generation

I help creative teams translate cinematography decisions into references, prompts, shot logs and review gates that make AI video directable and finishable.

Launching a business of your own? Founder Launch OS connects the brand, offer, website and visual campaign in one guided Codex or Claude Code workspace.

Build a controlled AI video workflowDiscuss an AI cinematography system

Keep reading

ARRI Signature comparison with a woman shown at zoom focal lengths from 45 mm to 510 mm and a man shown at prime focal lengths from 12 mm to 280 mmAI Cinematography / 9 min readAI Video Lens Prompts Need More Than a Focal LengthCamera operator working beside a cinema camera on a film setAI Cinematography / 8 min readAI Video Prompt Shot Language for CinematographersLuma Dream Machine interface combining an input video, start and end frames, and a character referenceAI Cinematography / 9 min readChoose the Control Before You Prompt AI Video