Comparing AI Video Models: Wan 2.2 vs LTX 2.3 vs Minimax H3

This guide compares three leading open-source AI video models, Wan 2.2, LTX 2.3, and Minimax H3, by testing how each one handles the "Rack Frame" composition. We used the same text-to-video prompt across all three models and evaluated the results based on prompt understanding, camera handling, lighting, and audio quality. Whether you are a filmmaker, content creator, or AI enthusiast, this comparison will help you choose the right model for scenes that demand precise focus shifts and dynamic camera work.

Poster Wan 2.2 vs LTX 2.3 vs Minimax H3

What is the "Rack Frame" composition?

The rack frame, more commonly known in cinematography as a "rack focus" or "pull focus", is a technique where the camera shifts focus from one subject in the foreground to another subject in the background (or vice versa) within a single, unbroken shot. Instead of cutting between two angles, the filmmaker keeps both subjects visible in the frame but changes which one is sharp and which is blurred, guiding the viewer's eye from one character to the other without breaking the spatial continuity of the scene.

This technique is a staple of tense interrogations, standoff scenes, and moments of revelation, where the power dynamic between two characters is shifting second-by-second. By pulling focus from the person in control to the person on the defensive, or the reverse, the rack frame visually communicates who holds the upper hand at any given moment. In AI video generation, achieving a convincing rack focus is particularly challenging because the model must understand depth of field, maintain consistent character placement, and execute a smooth focal transition while keeping the rest of the frame stable.

When to use the "Rack Frame" composition

The rack frame composition is best suited for scenes where two subjects share a frame but the audience's attention needs to shift between them without cutting. It excels in interrogation rooms, standoff confrontations, and dramatic revelations where a character's reaction is as important as the action itself. It is also effective in emotional dialogues where the focus pull mirrors a shift in emotional weight. Avoid this technique when the scene requires fast-paced action or when the two subjects are at similar depths, as the focal transition will be less noticeable and lose its dramatic impact.

  • Tense standoffs where the power dynamic is shifting between two characters
  • Interrogation or confrontation scenes with a detective and a suspect
  • Dramatic revelations where a character's reaction carries the scene
  • Emotional dialogues where focus shifts mirror shifts in emotional weight
  • Moments where cutting would break tension and continuity must be preserved

Pros and Cons of the "Rack Frame" composition

Pros Cons
Directs viewer attention without cutting, preserving spatial continuity Requires precise depth-of-field control, which is difficult for AI models
Creates tension by visually communicating shifting power dynamics Can be confusing if the focal transition is too fast or too subtle
Adds cinematic sophistication and professional polish to a scene Needs clear foreground-background separation to be effective
Allows two characters' reactions to be shown in a single shot Demands consistent character placement throughout the focus shift

Prompt

We used the same text-to-video prompt across all three models to ensure a fair comparison. The prompt describes a classic interrogation room scene with a detective and a suspect, requiring the model to execute a rack focus shift, a camera zoom, and an over-the-shoulder shot, all challenging techniques that test the model's understanding of composition, camera movement, and spatial awareness.

"Two people at the table in a dark interrogation room: detective and suspect. Detective in foreground and suspect in background, focus is shifting between them. In the beginning focus is on the detective, in the end it is on the suspect. Camera slowly zooms in on the face of the suspect. As the focus changes the view is shifting from the detective to the suspect in frame visually. When focus shifts, camera zooms to suspect. Use rack frame composition. Camera behind the detective, over the shoulder shot. Detective says: 'Perhaps it's time to confess. Let's face it, we are at the end of the road.'"

Minimax H3

Minimax H3 produced the most convincing result of the three models. It correctly rendered the interrogation room setting with appropriate dark lighting, maintained two distinct characters (detective and suspect), and executed the rack focus shift smoothly. The camera handling was excellent across all evaluated criteria, and the model also generated coherent audio with the detective's dialogue.

Minimax H3, Rack Frame composition result

LTX 2.3

LTX 2.3 struggled with this prompt. The model rendered two separate scenes and combined them into one, failing to maintain a single coherent shot with the rack focus shift. The camera did not shift focus or view as instructed, and the lighting did not convey the dark interrogation room atmosphere. Additionally, the generated audio contained errors.

LTX 2.3, Rack Frame composition result

Wan 2.2

Wan 2.2 produced a partially successful result. It correctly rendered the dark interrogation room lighting and executed the camera movements (focus shift, view shift, zoom, and over-the-shoulder shot) adequately. However, it failed on character accuracy, rendering two policemen instead of a detective and a suspect, and does not generate audio at all, which limits its usefulness for scenes that require dialogue.

Wan 2.2, Rack Frame composition result

Prompt Understanding

This section evaluates how well each model understood and followed the prompt's instructions regarding characters, location, and composition. A model that misunderstands the characters or the composition cannot produce a usable result, regardless of how well it handles camera movement or lighting.

Minimax H3 LTX 2.3 Wan 2.2
Characters OK OK Errors
Location OK OK OK
Composition OK Errors OK

LTX 2.3 Errors:
- The model rendered two scenes and combined them into one, breaking the single-shot rack frame requirement
- The audio had errors

Wan 2.2 Errors:
- The characters did not follow the prompt: there were two policemen instead of a detective and a suspect
- The model does not generate audio

Winner: Minimax H3, the only model that correctly understood all three aspects of the prompt.

Camera Handling

Camera handling is critical for the rack frame composition, which depends on precise focus shifts, view changes, zoom control, and correct shot types. This section evaluates how well each model executed the camera instructions in the prompt, including the over-the-shoulder shot and the gradual zoom on the suspect's face.

Minimax H3 LTX 2.3 Wan 2.2
Shifting focus Excellent No OK
Shifting view Excellent No OK
Zoom handling Excellent No OK
Camera movement Excellent No OK
Shot type (over-the-shoulder shot) OK Errors OK

Minimax H3 delivered excellent camera handling across all criteria, executing the focus shift, view shift, zoom, and camera movement exactly as described in the prompt. Wan 2.2 produced acceptable results, the camera movements were present but not as smooth or precise. LTX 2.3 failed to execute any of the camera instructions, likely because it split the scene into two separate renders.

Winner: Minimax H3

Lighting

The prompt calls for a dark interrogation room, which requires the model to understand and render low-key lighting with dramatic shadows. This section evaluates whether each model captured the intended moody, high-contrast atmosphere.

Minimax H3 LTX 2.3 Wan 2.2
Lighting Excellent Errors Excellent

Both Minimax H3 and Wan 2.2 understood the lighting concept of a dark interrogation room and produced appropriately moody, high-contrast results. LTX 2.3 failed to capture the dark room atmosphere, likely due to the scene-splitting issue that disrupted its overall rendering.

Audio

The prompt includes dialogue for the detective, making audio generation an important evaluation criterion. This section assesses whether each model produced coherent, error-free audio that matches the prompt's dialogue.

Minimax H3 LTX 2.3 Wan 2.2
Audio OK Errors No audio

Minimax H3 generated coherent audio with the detective's dialogue, making it the clear winner in this category. LTX 2.3 produced audio but it contained errors. Wan 2.2 does not support audio generation at all, which is a significant limitation for scenes that rely on dialogue.

Winner: Minimax H3

Overall Comparison Summary

The table below provides a side-by-side summary of all evaluation criteria. Minimax H3 is the only model that performed well across every category, making it the clear choice for rack frame compositions and similar camera-intensive scenes.

Criterion Minimax H3 LTX 2.3 Wan 2.2
Characters OK OK Errors
Location OK OK OK
Composition OK Errors OK
Shifting focus Excellent No OK
Shifting view Excellent No OK
Zoom handling Excellent No OK
Camera movement Excellent No OK
Shot type (over-the-shoulder) OK Errors OK
Lighting Excellent Errors Excellent
Audio OK Errors No audio

Overall winner: Minimax H3, the only model that delivered strong results across prompt understanding, camera handling, lighting, and audio simultaneously.

The tool I have used

I used the VideoCool software to generate 5 seconds of video in 1366×768 (HD) resolution for each model. You can see the exact steps in the video below.

How to generate the videos, step-by-step walkthrough

Recommended book

Learn more about AI video editing

Learn more about AI video editing by reading the book How to Make AI Videos by Gyula Rabai. In this book you can find very important information about: Camera Movement, Lighting, Colors, Composition, Consistency, Character emotion, Character motion, Video style, Action Control and Sound design This information will make your videos more professional, and more engaging.

Get your copy here:
https://videcool.com/p_3707-how-to-make-ai-videos-by-gyula-rabai-book.html

Conclusion

The rack frame composition is one of the most demanding techniques for AI video models, requiring precise control over focus shifts, camera movement, lighting, and character consistency, all within a single shot. Our comparison shows that Minimax H3 is currently the best-performing open-source model for this type of scene, delivering excellent camera handling, correct prompt understanding, appropriate dark lighting, and coherent audio. Wan 2.2 showed promise with strong lighting and acceptable camera work, but its failure to render the correct characters and lack of audio generation limit its usefulness for dialogue-driven scenes. LTX 2.3 struggled the most, splitting the scene into two separate renders and failing to execute the camera instructions or capture the intended lighting.

For filmmakers and prompt engineers working with rack frame compositions or similar camera-intensive techniques, Minimax H3 is the recommended choice. As open-source models continue to evolve, we expect improvements in focus control and scene coherence, but for now, Minimax H3 stands out as the most capable option for this challenging composition.