Comparing AI Video Models: Wan 2.2 vs LTX 2.3 vs Minimax H3
This guide compares three leading open-source AI video models, Wan 2.2, LTX 2.3, and Minimax H3, by testing how each one handles the "Rack Frame" composition. We used the same text-to-video prompt across all three models and evaluated the results based on prompt understanding, camera handling, lighting, and audio quality. Whether you are a filmmaker, content creator, or AI enthusiast, this comparison will help you choose the right model for scenes that demand precise focus shifts and dynamic camera work.
What is the "Rack Frame" composition?
The rack frame, more commonly known in cinematography as a "rack focus" or "pull focus", is a technique where the camera shifts focus from one subject in the foreground to another subject in the background (or vice versa) within a single, unbroken shot. Instead of cutting between two angles, the filmmaker keeps both subjects visible in the frame but changes which one is sharp and which is blurred, guiding the viewer's eye from one character to the other without breaking the spatial continuity of the scene.
This technique is a staple of tense interrogations, standoff scenes, and moments of revelation, where the power dynamic between two characters is shifting second-by-second. By pulling focus from the person in control to the person on the defensive, or the reverse, the rack frame visually communicates who holds the upper hand at any given moment. In AI video generation, achieving a convincing rack focus is particularly challenging because the model must understand depth of field, maintain consistent character placement, and execute a smooth focal transition while keeping the rest of the frame stable.
When to use the "Rack Frame" composition
The rack frame composition is best suited for scenes where two subjects share a frame but the audience's attention needs to shift between them without cutting. It excels in interrogation rooms, standoff confrontations, and dramatic revelations where a character's reaction is as important as the action itself. It is also effective in emotional dialogues where the focus pull mirrors a shift in emotional weight. Avoid this technique when the scene requires fast-paced action or when the two subjects are at similar depths, as the focal transition will be less noticeable and lose its dramatic impact.
- Tense standoffs where the power dynamic is shifting between two characters
- Interrogation or confrontation scenes with a detective and a suspect
- Dramatic revelations where a character's reaction carries the scene
- Emotional dialogues where focus shifts mirror shifts in emotional weight
- Moments where cutting would break tension and continuity must be preserved
Pros and Cons of the "Rack Frame" composition
| Pros | Cons |
|---|---|
| Directs viewer attention without cutting, preserving spatial continuity | Requires precise depth-of-field control, which is difficult for AI models |
| Creates tension by visually communicating shifting power dynamics | Can be confusing if the focal transition is too fast or too subtle |
| Adds cinematic sophistication and professional polish to a scene | Needs clear foreground-background separation to be effective |
| Allows two characters' reactions to be shown in a single shot | Demands consistent character placement throughout the focus shift |
Prompt
We used the same text-to-video prompt across all three models to ensure a fair comparison. The prompt describes a classic interrogation room scene with a detective and a suspect, requiring the model to execute a rack focus shift, a camera zoom, and an over-the-shoulder shot, all challenging techniques that test the model's understanding of composition, camera movement, and spatial awareness.
"Two people at the table in a dark interrogation room: detective and suspect. Detective in foreground and suspect in background, focus is shifting between them. In the beginning focus is on the detective, in the end it is on the suspect. Camera slowly zooms in on the face of the suspect. As the focus changes the view is shifting from the detective to the suspect in frame visually. When focus shifts, camera zooms to suspect. Use rack frame composition. Camera behind the detective, over the shoulder shot. Detective says: 'Perhaps it's time to confess. Let's face it, we are at the end of the road.'"
Minimax H3
Minimax H3 produced the most convincing result of the three models. It correctly rendered the interrogation room setting with appropriate dark lighting, maintained two distinct characters (detective and suspect), and executed the rack focus shift smoothly. The camera handling was excellent across all evaluated criteria, and the model also generated coherent audio with the detective's dialogue.
LTX 2.3
LTX 2.3 struggled with this prompt. The model rendered two separate scenes and combined them into one, failing to maintain a single coherent shot with the rack focus shift. The camera did not shift focus or view as instructed, and the lighting did not convey the dark interrogation room atmosphere. Additionally, the generated audio contained errors.
Wan 2.2
Wan 2.2 produced a partially successful result. It correctly rendered the dark interrogation room lighting and executed the camera movements (focus shift, view shift, zoom, and over-the-shoulder shot) adequately. However, it failed on character accuracy, rendering two policemen instead of a detective and a suspect, and does not generate audio at all, which limits its usefulness for scenes that require dialogue.
Prompt Understanding
This section evaluates how well each model understood and followed the prompt's instructions regarding characters, location, and composition. A model that misunderstands the characters or the composition cannot produce a usable result, regardless of how well it handles camera movement or lighting.
| Minimax H3 | LTX 2.3 | Wan 2.2 | |
|---|---|---|---|
| Characters | OK | OK | Errors |
| Location | OK | OK | OK |
| Composition | OK | Errors | OK |
LTX 2.3 Errors:
- The model rendered two scenes and combined them into one, breaking the single-shot rack frame requirement
- The audio had errors
Wan 2.2 Errors:
- The characters did not follow the prompt: there were two policemen instead of a detective and a suspect
- The model does not generate audio
Winner: Minimax H3, the only model that correctly understood all three aspects of the prompt.
Camera Handling
Camera handling is critical for the rack frame composition, which depends on precise focus shifts, view changes, zoom control, and correct shot types. This section evaluates how well each model executed the camera instructions in the prompt, including the over-the-shoulder shot and the gradual zoom on the suspect's face.
| Minimax H3 | LTX 2.3 | Wan 2.2 | |
|---|---|---|---|
| Shifting focus | Excellent | No | OK |
| Shifting view | Excellent | No | OK |
| Zoom handling | Excellent | No | OK |
| Camera movement | Excellent | No | OK |
| Shot type (over-the-shoulder shot) | OK | Errors | OK |
Minimax H3 delivered excellent camera handling across all criteria, executing the focus shift, view shift, zoom, and camera movement exactly as described in the prompt. Wan 2.2 produced acceptable results, the camera movements were present but not as smooth or precise. LTX 2.3 failed to execute any of the camera instructions, likely because it split the scene into two separate renders.
Winner: Minimax H3
Lighting
The prompt calls for a dark interrogation room, which requires the model to understand and render low-key lighting with dramatic shadows. This section evaluates whether each model captured the intended moody, high-contrast atmosphere.
| Minimax H3 | LTX 2.3 | Wan 2.2 | |
|---|---|---|---|
| Lighting | Excellent | Errors | Excellent |
Both Minimax H3 and Wan 2.2 understood the lighting concept of a dark interrogation room and produced appropriately moody, high-contrast results. LTX 2.3 failed to capture the dark room atmosphere, likely due to the scene-splitting issue that disrupted its overall rendering.
Audio
The prompt includes dialogue for the detective, making audio generation an important evaluation criterion. This section assesses whether each model produced coherent, error-free audio that matches the prompt's dialogue.
| Minimax H3 | LTX 2.3 | Wan 2.2 | |
|---|---|---|---|
| Audio | OK | Errors | No audio |
Minimax H3 generated coherent audio with the detective's dialogue, making it the clear winner in this category. LTX 2.3 produced audio but it contained errors. Wan 2.2 does not support audio generation at all, which is a significant limitation for scenes that rely on dialogue.
Winner: Minimax H3
Overall Comparison Summary
The table below provides a side-by-side summary of all evaluation criteria. Minimax H3 is the only model that performed well across every category, making it the clear choice for rack frame compositions and similar camera-intensive scenes.
| Criterion | Minimax H3 | LTX 2.3 | Wan 2.2 |
|---|---|---|---|
| Characters | OK | OK | Errors |
| Location | OK | OK | OK |
| Composition | OK | Errors | OK |
| Shifting focus | Excellent | No | OK |
| Shifting view | Excellent | No | OK |
| Zoom handling | Excellent | No | OK |
| Camera movement | Excellent | No | OK |
| Shot type (over-the-shoulder) | OK | Errors | OK |
| Lighting | Excellent | Errors | Excellent |
| Audio | OK | Errors | No audio |
Overall winner: Minimax H3, the only model that delivered strong results across prompt understanding, camera handling, lighting, and audio simultaneously.
The tool I have used
I used the VideoCool software to generate 5 seconds of video in 1366×768 (HD) resolution for each model. You can see the exact steps in the video below.
Recommended book
Learn more about AI video editing
Learn more about AI video editing by reading the book How to Make AI Videos by Gyula Rabai. In this book you can find very important information about: Camera Movement, Lighting, Colors, Composition, Consistency, Character emotion, Character motion, Video style, Action Control and Sound design This information will make your videos more professional, and more engaging.
Get your copy here:
https://videcool.com/p_3707-how-to-make-ai-videos-by-gyula-rabai-book.html
Conclusion
The rack frame composition is one of the most demanding techniques for AI video models, requiring precise control over focus shifts, camera movement, lighting, and character consistency, all within a single shot. Our comparison shows that Minimax H3 is currently the best-performing open-source model for this type of scene, delivering excellent camera handling, correct prompt understanding, appropriate dark lighting, and coherent audio. Wan 2.2 showed promise with strong lighting and acceptable camera work, but its failure to render the correct characters and lack of audio generation limit its usefulness for dialogue-driven scenes. LTX 2.3 struggled the most, splitting the scene into two separate renders and failing to execute the camera instructions or capture the intended lighting.
For filmmakers and prompt engineers working with rack frame compositions or similar camera-intensive techniques, Minimax H3 is the recommended choice. As open-source models continue to evolve, we expect improvements in focus control and scene coherence, but for now, Minimax H3 stands out as the most capable option for this challenging composition.