Qwen, Z-Image Turbo, and Anima on a 24 GB RTX 5090 Laptop
The earlier comparison was wrong in an important way: it forced three models through one sampler recipe. That made the result easy to tabulate, but it did not test what Qwen Image or Z-Image Turbo are designed to do.
This rerun uses each model's documented native recipe, the official or official-mirror model files, and two fixed seeds. It also records the failure that initially made Qwen look unusable: on this RTX 5090 Laptop, ComfyUI's --use-sage-attention startup flag produced black Qwen images. Removing that flag fixed the images without changing the Qwen weight, prompt, VAE, seed, or sampler settings. ComfyUI's own Blackwell guidance warns that the Triton SageAttention launch flag can produce black Qwen and Wan outputs. ComfyUI discussion
The scene prompt was supplied by Doubao. The setup, validation, and visual review were performed in the current Codex session. Scores are reproducible assistant review notes, not a human preference study or a general leaderboard.
The practical result
For this hand-painted mountain-cabin scene, Qwen Image 2512 FP8 was the best overall image once the attention backend was corrected. Z-Image Turbo was the best speed-quality tradeoff. Anima Aesthetic v1.1 gave the clearest anime-specific look and the lowest VRAM use, but it leaned toward contemporary high-contrast illustration rather than the muted watercolor animation brief.
| Model | Native recipe | Seed-42 time | Peak VRAM | Result on this scene |
|---|---|---|---|---|
| Qwen Image 2512 FP8 | Euler, 50 steps, CFG 4.0 | 220.772 s | 23,781 MiB | Best brief adherence and atmosphere |
| Z-Image Turbo BF16 | Euler, 9 steps, CFG 1.0 | 21.239 s | 23,144 MiB | Strong scene quality at roughly one tenth of Qwen's time |
| Anima Aesthetic v1.1 | ER-SDE, 30 steps, CFG 4.0 | 30.409 s | 6,888 MiB | Clean modern anime illustration, less faithful to the vintage brief |
All final T2I runs used the same native 2:3 canvas, 1056 × 1584. That is Qwen's documented 2:3 bucket; using it avoids treating an unsupported 1280 × 1792 canvas as a model failure. Qwen's native ComfyUI workflow
What changed from the first run
Three changes mattered.
- Qwen moved from a Q6 GGUF to the official
qwen_image_2512_fp8_e4m3fn.safetensorsrelease. The BF16 alternative is about 40.9 GB before runtime memory, so it is not a viable 24 GB option. The 20,430,679,144-byte FP8 file passed SHA-256 verification before use. - Z-Image Turbo remained on its verified BF16 release, but switched from the inappropriate UniPC / CFG 3.5 / 7-step setup to the official Turbo recipe: Euler, 9 steps, CFG 1.0,
simplescheduler, and AuraFlow shift 3.0. Z-Image's official repository - The ComfyUI service no longer launches with
--use-sage-attention. This was the black-frame root cause for Qwen on this Blackwell laptop. The previous launch configuration was backed up before the managed service was restarted.
The Qwen FP8 model is therefore the highest-quality official Qwen weight that this 24 GB card can load in the tested stack. It is usable after the attention fix, but it runs with very little VRAM headroom. Z's BF16 release also fits, although it uses similarly high VRAM.
Models and provenance
| Model | File used | Why it was selected |
|---|---|---|
| Qwen Image 2512 | qwen_image_2512_fp8_e4m3fn.safetensors | Official FP8 file; BF16 is not realistic on 24 GB |
| Z-Image Turbo | z_image_turbo_bf16.safetensors | Official quality-preserving Turbo release already installed and hash-verified |
| Anima Aesthetic v1.1 | anima-aesthetic-v1.1.safetensors | Current native-ComfyUI anime and non-photorealistic art model; selected as a domain-specific challenger |
Anima is a useful anime challenger, not an asserted public-benchmark champion. Its model card describes training focused on anime and non-photorealistic art, but I found no independent, reproducible public anime leaderboard that would justify declaring it a universal winner before testing. Its weights are non-commercial; treat this as an evaluation model unless its license fits the intended use. Anima model card
Prompts and native settings
The three models received prompt wording adapted to their intended text encoders. No LoRA, ControlNet, or IP-Adapter was loaded.
Qwen Image 2512
Studio Ghibli-inspired hand-painted animation scene at dusk: a young woman stands on the wooden terrace of a mountain cabin, with long hair moving in the wind. Layered forested mountains recede through thin evening mist. Soft diffused amber sunset light, restrained pastel color, vintage cel painting, subtle watercolor paper texture, faint film grain, quiet nostalgic mood, wide cinematic composition.
Negative prompt:
low resolution, low quality, deformed limbs, malformed fingers, oversaturated image, waxy skin, featureless face, overly smooth image, AI artifacts, confused composition, blurred or distorted text
AuraFlow shift 3.1 · Euler · simple · 50 steps · CFG 4.0 · 1056 x 1584
Z-Image Turbo
Studio Ghibli-inspired hand-painted animation scene at dusk: a young woman stands on the wooden terrace of a mountain cabin, with long hair moving in the wind. Layered forested mountains recede through thin evening mist. Soft diffused amber sunset light, restrained pastel color, vintage cel painting, subtle watercolor paper texture, faint film grain, quiet nostalgic mood, wide cinematic composition.
AuraFlow shift 3.0 · Euler · simple · 9 steps · CFG 1.0 · 1056 x 1584
Anima Aesthetic v1.1
masterpiece, best quality, hand-painted Japanese fantasy animation, young woman on the wooden terrace of a mountain cabin at dusk, long hair in the wind, layered forested mountains, evening mist, soft amber sunset light, restrained pastels, vintage cel painting, watercolor texture, faint film grain, cinematic wide shot
Negative prompt:
worst quality, low quality, score_1, score_2, score_3, blurry, jpeg artifacts, text, signature
AuraFlow shift 3.0 · ER-SDE · simple · 30 steps · CFG 4.0 · 1056 x 1584
T2I outputs
Qwen Image 2512 FP8


Both Qwen images contain the full requested scene: the cabin and terrace, a wind-blown figure, layered mountains, atmospheric mist, and a warm low sun. The watercolor paper texture is present without turning the characters into glossy 3D renderings. The main tradeoff is runtime: a full-quality image takes about three and a half minutes.
Z-Image Turbo BF16


Z now follows the brief reliably. It preserves the mountain cabin, terrace, figure, dusk, and mist across both seeds, with a softer, more film-like treatment than the first failed run. The result is especially compelling for a nine-step model, although the figure and cabin are slightly less controlled than Qwen's.
Anima Aesthetic v1.1


Anima clearly understands anime illustration: the anatomy, line work, and architectural shapes are clean. It consistently pushes the image toward saturated cyan shadows, dramatic ink lines, and a modern key-visual composition. That is a good option when that look is desired, but it is not the closest match to a restrained, soft, older-animation frame.
Review scores
Scores are 1–5. They assess the actual images above, not the models' claims or benchmark cards.
| Model / seed | D1 style fit | D2 brief adherence | D3 structure | D4 light and depth | D5 overall | Verdict |
|---|---|---|---|---|---|---|
| Qwen / 42 | 4 | 5 | 5 | 5 | 5 | Best overall |
| Qwen / 1337 | 4 | 5 | 5 | 5 | 5 | Best overall |
| Z / 42 | 4 | 5 | 5 | 5 | 4 | Best speed-quality tradeoff |
| Z / 1337 | 4 | 5 | 5 | 5 | 4 | Best speed-quality tradeoff |
| Anima / 42 | 3 | 5 | 5 | 3 | 4 | Strong modern anime look |
| Anima / 1337 | 3 | 5 | 5 | 4 | 4 | Strong modern anime look |
The post-generation upscale stage
Each seed-42 image went through the same post stage:
ImageUpscaleWithModel(4x-UltraSharp.pth) → ImageScale(Lanczos, 2112 × 3168)
The model first performs its native four-times upscale; the image is then downscaled to a practical 2x delivery size. The stage is separate from diffusion and does not change the seed, prompt, or model ranking.
| Source model | Upscale time | Peak VRAM | Visual effect |
|---|---|---|---|
| Qwen Image 2512 | 9.397 s | 21,317 MiB | Cleaner railings, hair, and cabin texture; stronger ink edges |
| Z-Image Turbo | 9.554 s | 19,912 MiB | Better foliage separation and wood detail; can slightly sharpen painted softness |
| Anima Aesthetic v1.1 | 8.203 s | 6,184 MiB | Crisp line art; least likely to need it because the base image is already contrasty |



The stage improves usable delivery resolution and local crispness, but it does not invent trustworthy scene detail. For Qwen and Z, use it for a larger export; for Anima, prefer the native output when the softer original line treatment is important.
Image editing: Qwen Image Edit 2511
The I2I portion uses the supplied two-character illustration and the current official Qwen Image Edit 2511 FP8mixed release. It is intentionally separate from the T2I ranking: Z-Image Turbo does not publish an equivalent native editing workflow, so a forced VAE-encode I2I path would not be a fair model comparison.
The 20,533,762,817-byte Edit FP8mixed checkpoint passed SHA-256 verification before use. The source was resized without distortion to Qwen's 1056 x 1584 bucket and padded around the artwork. The workflow used Qwen Image Edit's reference-conditioning nodes, AuraFlow shift 3.1, CFGNorm 1.0, Euler/simple, 40 steps, CFG 3.0, and denoise 0.65. Both tasks ran serially at seeds 42 and 1337.

I2I-A: style-preserving restyle
Prompt:
Restyle image 1 as a Studio Ghibli-inspired hand-painted animation. Preserve the full composition, the two character poses, the clothing, the camera framing, and every object placement. Add restrained cel-painted color, subtle watercolor texture, and faint film grain only.
| Seed | Time | Peak VRAM | R1 composition and pose | R2 style transfer | R3 artifact control |
|---|---|---|---|---|---|
| 42 | 320.447 s | 23,670 MiB | 5 | 4 | 5 |
| 1337 | 304.687 s | 22,838 MiB | 5 | 4 | 5 |


This is a very strong preservation result. The model leaves the two poses, clothing, spatial relationship, and framing essentially untouched, and subtly cleans up the cel-and-watercolor treatment. It does not over-style the image, which is the right behavior for this task.
I2I-B: weather edit to rainy night
Prompt:
Keep image 1's full composition, the two character poses, the clothing, the camera framing, and every object placement unchanged. Change only the weather and lighting to a rainy night, with visible heavy rain, wet reflective surfaces, and soft night ambient light in a Studio Ghibli-inspired hand-painted animation style.
| Seed | Time | Peak VRAM | E1 instruction execution | E2 style consistency | E3 structure fidelity |
|---|---|---|---|---|---|
| 42 | 321.011 s | 23,670 MiB | 3 | 4 | 5 |
| 1337 | 308.631 s | 23,098 MiB | 3 | 4 | 5 |


The edit preserves the image remarkably well and clearly adds rain, droplets, wet clothing, and puddles. It does not deliver the requested night lighting: both samples remain bright and warm. That is a useful failure to record rather than hide. At denoise 0.65, reference preservation wins over the stronger lighting change; a real production pass would either raise denoise slightly or use a masked/background-focused edit.
Recommendation
Use Qwen Image 2512 FP8 for the best final stills when a roughly 3.5-minute render time and near-full 24 GB VRAM use are acceptable. Use Z-Image Turbo BF16 when speed matters: it delivered a strong result in about 20 seconds with the correct nine-step workflow. Use Anima Aesthetic v1.1 for a bold modern anime key visual and only where its non-commercial-weight license is acceptable.
The most important operational change is not a different prompt. It is keeping --use-sage-attention out of this Blackwell ComfyUI launch profile. With it, Qwen generated black frames; without it, the same verified FP8 model and native workflow produced the leading images in this test.