Advertisement

Qwen, Z-Image Turbo, and Anima on a 24 GB RTX 5090 Laptop

2026-08-12

The earlier comparison was wrong in an important way: it forced three models through one sampler recipe. That made the result easy to tabulate, but it did not test what Qwen Image or Z-Image Turbo are designed to do.

This rerun uses each model's documented native recipe, the official or official-mirror model files, and two fixed seeds. It also records the failure that initially made Qwen look unusable: on this RTX 5090 Laptop, ComfyUI's --use-sage-attention startup flag produced black Qwen images. Removing that flag fixed the images without changing the Qwen weight, prompt, VAE, seed, or sampler settings. ComfyUI's own Blackwell guidance warns that the Triton SageAttention launch flag can produce black Qwen and Wan outputs. ComfyUI discussion

The scene prompt was supplied by Doubao. The setup, validation, and visual review were performed in the current Codex session. Scores are reproducible assistant review notes, not a human preference study or a general leaderboard.

The practical result

For this hand-painted mountain-cabin scene, Qwen Image 2512 FP8 was the best overall image once the attention backend was corrected. Z-Image Turbo was the best speed-quality tradeoff. Anima Aesthetic v1.1 gave the clearest anime-specific look and the lowest VRAM use, but it leaned toward contemporary high-contrast illustration rather than the muted watercolor animation brief.

ModelNative recipeSeed-42 timePeak VRAMResult on this scene
Qwen Image 2512 FP8Euler, 50 steps, CFG 4.0220.772 s23,781 MiBBest brief adherence and atmosphere
Z-Image Turbo BF16Euler, 9 steps, CFG 1.021.239 s23,144 MiBStrong scene quality at roughly one tenth of Qwen's time
Anima Aesthetic v1.1ER-SDE, 30 steps, CFG 4.030.409 s6,888 MiBClean modern anime illustration, less faithful to the vintage brief

All final T2I runs used the same native 2:3 canvas, 1056 × 1584. That is Qwen's documented 2:3 bucket; using it avoids treating an unsupported 1280 × 1792 canvas as a model failure. Qwen's native ComfyUI workflow

What changed from the first run

Three changes mattered.

  1. Qwen moved from a Q6 GGUF to the official qwen_image_2512_fp8_e4m3fn.safetensors release. The BF16 alternative is about 40.9 GB before runtime memory, so it is not a viable 24 GB option. The 20,430,679,144-byte FP8 file passed SHA-256 verification before use.
  2. Z-Image Turbo remained on its verified BF16 release, but switched from the inappropriate UniPC / CFG 3.5 / 7-step setup to the official Turbo recipe: Euler, 9 steps, CFG 1.0, simple scheduler, and AuraFlow shift 3.0. Z-Image's official repository
  3. The ComfyUI service no longer launches with --use-sage-attention. This was the black-frame root cause for Qwen on this Blackwell laptop. The previous launch configuration was backed up before the managed service was restarted.

The Qwen FP8 model is therefore the highest-quality official Qwen weight that this 24 GB card can load in the tested stack. It is usable after the attention fix, but it runs with very little VRAM headroom. Z's BF16 release also fits, although it uses similarly high VRAM.

Models and provenance

ModelFile usedWhy it was selected
Qwen Image 2512qwen_image_2512_fp8_e4m3fn.safetensorsOfficial FP8 file; BF16 is not realistic on 24 GB
Z-Image Turboz_image_turbo_bf16.safetensorsOfficial quality-preserving Turbo release already installed and hash-verified
Anima Aesthetic v1.1anima-aesthetic-v1.1.safetensorsCurrent native-ComfyUI anime and non-photorealistic art model; selected as a domain-specific challenger

Anima is a useful anime challenger, not an asserted public-benchmark champion. Its model card describes training focused on anime and non-photorealistic art, but I found no independent, reproducible public anime leaderboard that would justify declaring it a universal winner before testing. Its weights are non-commercial; treat this as an evaluation model unless its license fits the intended use. Anima model card

Prompts and native settings

The three models received prompt wording adapted to their intended text encoders. No LoRA, ControlNet, or IP-Adapter was loaded.

Qwen Image 2512

Studio Ghibli-inspired hand-painted animation scene at dusk: a young woman stands on the wooden terrace of a mountain cabin, with long hair moving in the wind. Layered forested mountains recede through thin evening mist. Soft diffused amber sunset light, restrained pastel color, vintage cel painting, subtle watercolor paper texture, faint film grain, quiet nostalgic mood, wide cinematic composition.

Negative prompt:

low resolution, low quality, deformed limbs, malformed fingers, oversaturated image, waxy skin, featureless face, overly smooth image, AI artifacts, confused composition, blurred or distorted text

AuraFlow shift 3.1 · Euler · simple · 50 steps · CFG 4.0 · 1056 x 1584

Z-Image Turbo

Studio Ghibli-inspired hand-painted animation scene at dusk: a young woman stands on the wooden terrace of a mountain cabin, with long hair moving in the wind. Layered forested mountains recede through thin evening mist. Soft diffused amber sunset light, restrained pastel color, vintage cel painting, subtle watercolor paper texture, faint film grain, quiet nostalgic mood, wide cinematic composition.

AuraFlow shift 3.0 · Euler · simple · 9 steps · CFG 1.0 · 1056 x 1584

Anima Aesthetic v1.1

masterpiece, best quality, hand-painted Japanese fantasy animation, young woman on the wooden terrace of a mountain cabin at dusk, long hair in the wind, layered forested mountains, evening mist, soft amber sunset light, restrained pastels, vintage cel painting, watercolor texture, faint film grain, cinematic wide shot

Negative prompt:

worst quality, low quality, score_1, score_2, score_3, blurry, jpeg artifacts, text, signature

AuraFlow shift 3.0 · ER-SDE · simple · 30 steps · CFG 4.0 · 1056 x 1584

T2I outputs

Qwen Image 2512 FP8

Qwen Image 2512 FP8, seed 42

Qwen Image 2512 FP8, seed 1337

Both Qwen images contain the full requested scene: the cabin and terrace, a wind-blown figure, layered mountains, atmospheric mist, and a warm low sun. The watercolor paper texture is present without turning the characters into glossy 3D renderings. The main tradeoff is runtime: a full-quality image takes about three and a half minutes.

Z-Image Turbo BF16

Z-Image Turbo BF16, seed 42

Z-Image Turbo BF16, seed 1337

Z now follows the brief reliably. It preserves the mountain cabin, terrace, figure, dusk, and mist across both seeds, with a softer, more film-like treatment than the first failed run. The result is especially compelling for a nine-step model, although the figure and cabin are slightly less controlled than Qwen's.

Anima Aesthetic v1.1

Anima Aesthetic v1.1, seed 42

Anima Aesthetic v1.1, seed 1337

Anima clearly understands anime illustration: the anatomy, line work, and architectural shapes are clean. It consistently pushes the image toward saturated cyan shadows, dramatic ink lines, and a modern key-visual composition. That is a good option when that look is desired, but it is not the closest match to a restrained, soft, older-animation frame.

Review scores

Scores are 1–5. They assess the actual images above, not the models' claims or benchmark cards.

Model / seedD1 style fitD2 brief adherenceD3 structureD4 light and depthD5 overallVerdict
Qwen / 4245555Best overall
Qwen / 133745555Best overall
Z / 4245554Best speed-quality tradeoff
Z / 133745554Best speed-quality tradeoff
Anima / 4235534Strong modern anime look
Anima / 133735544Strong modern anime look

The post-generation upscale stage

Each seed-42 image went through the same post stage:

ImageUpscaleWithModel(4x-UltraSharp.pth) → ImageScale(Lanczos, 2112 × 3168)

The model first performs its native four-times upscale; the image is then downscaled to a practical 2x delivery size. The stage is separate from diffusion and does not change the seed, prompt, or model ranking.

Source modelUpscale timePeak VRAMVisual effect
Qwen Image 25129.397 s21,317 MiBCleaner railings, hair, and cabin texture; stronger ink edges
Z-Image Turbo9.554 s19,912 MiBBetter foliage separation and wood detail; can slightly sharpen painted softness
Anima Aesthetic v1.18.203 s6,184 MiBCrisp line art; least likely to need it because the base image is already contrasty

Qwen seed 42 after the 2x stage

Z-Image seed 42 after the 2x stage

Anima seed 42 after the 2x stage

The stage improves usable delivery resolution and local crispness, but it does not invent trustworthy scene detail. For Qwen and Z, use it for a larger export; for Anima, prefer the native output when the softer original line treatment is important.

Image editing: Qwen Image Edit 2511

The I2I portion uses the supplied two-character illustration and the current official Qwen Image Edit 2511 FP8mixed release. It is intentionally separate from the T2I ranking: Z-Image Turbo does not publish an equivalent native editing workflow, so a forced VAE-encode I2I path would not be a fair model comparison.

The 20,533,762,817-byte Edit FP8mixed checkpoint passed SHA-256 verification before use. The source was resized without distortion to Qwen's 1056 x 1584 bucket and padded around the artwork. The workflow used Qwen Image Edit's reference-conditioning nodes, AuraFlow shift 3.1, CFGNorm 1.0, Euler/simple, 40 steps, CFG 3.0, and denoise 0.65. Both tasks ran serially at seeds 42 and 1337.

I2I source: the supplied two-character illustration

I2I-A: style-preserving restyle

Prompt:

Restyle image 1 as a Studio Ghibli-inspired hand-painted animation. Preserve the full composition, the two character poses, the clothing, the camera framing, and every object placement. Add restrained cel-painted color, subtle watercolor texture, and faint film grain only.
SeedTimePeak VRAMR1 composition and poseR2 style transferR3 artifact control
42320.447 s23,670 MiB545
1337304.687 s22,838 MiB545

Qwen Image Edit 2511 I2I-A, seed 42

Qwen Image Edit 2511 I2I-A, seed 1337

This is a very strong preservation result. The model leaves the two poses, clothing, spatial relationship, and framing essentially untouched, and subtly cleans up the cel-and-watercolor treatment. It does not over-style the image, which is the right behavior for this task.

I2I-B: weather edit to rainy night

Prompt:

Keep image 1's full composition, the two character poses, the clothing, the camera framing, and every object placement unchanged. Change only the weather and lighting to a rainy night, with visible heavy rain, wet reflective surfaces, and soft night ambient light in a Studio Ghibli-inspired hand-painted animation style.
SeedTimePeak VRAME1 instruction executionE2 style consistencyE3 structure fidelity
42321.011 s23,670 MiB345
1337308.631 s23,098 MiB345

Qwen Image Edit 2511 I2I-B, seed 42

Qwen Image Edit 2511 I2I-B, seed 1337

The edit preserves the image remarkably well and clearly adds rain, droplets, wet clothing, and puddles. It does not deliver the requested night lighting: both samples remain bright and warm. That is a useful failure to record rather than hide. At denoise 0.65, reference preservation wins over the stronger lighting change; a real production pass would either raise denoise slightly or use a masked/background-focused edit.

Recommendation

Use Qwen Image 2512 FP8 for the best final stills when a roughly 3.5-minute render time and near-full 24 GB VRAM use are acceptable. Use Z-Image Turbo BF16 when speed matters: it delivered a strong result in about 20 seconds with the correct nine-step workflow. Use Anima Aesthetic v1.1 for a bold modern anime key visual and only where its non-commercial-weight license is acceptable.

The most important operational change is not a different prompt. It is keeping --use-sage-attention out of this Blackwell ComfyUI launch profile. With it, Qwen generated black frames; without it, the same verified FP8 model and native workflow produced the leading images in this test.

Advertisement