Krea 2 SVDQuant in ComfyUI – How to Quantize Raw Finetunes to W4A4

If you found a Krea 2 Raw finetune you like, but it only comes in BF16 or FP8 and you need it to take up less VRAM or run faster on GPUs that do not have native FP8 Tensor Core acceleration but do have fast INT4 paths, you can convert the diffusion model to SVDQuant W4A4 from inside ComfyUI. This gives you a much smaller checkpoint and can make Krea 2 far faster on the right hardware.

This is especially useful on RTX 30-series Ampere GPUs such as the RTX 3090, but the node pack runs on newer Ada (4xxx) and Blackwell (5xxx) cards too. Ampere has fast INT8 and INT4 Tensor Core paths, but no native FP8 Tensor Core path. The Krea 2 SVDQuant node pack uses ComfyUI’s W4A4 backend and adds a BF16 low-rank correction branch to recover part of the quality lost by plain 4-bit quantization.

So, after converting your Krea 2 checkpoint to SVDQuant W4A4, you can benefit from:

  • Less VRAM required to run the model overall.
  • Much faster inference (image generation) on NVIDIA Ampere GPUs like the 3090.

All this with practically little to no perceivable loss in the quality of your outputs.

In the example below, just to show you how it works, I converted the Cat Tower v2.0 Raw FP8 finetune to a rank-256 SVDQuant checkpoint. The source file was 13.16GB, the finished file was 9.10GB, and the full conversion took about 5 minutes on an RTX 3090 with 100 refinements. You could also aim for even smaller quantizations, if you decide to select a lower rank setting.

You might also like: Best GPUs for Local AI Image Generation – My Top List

TL;DR – If You Already Know What You’re Doing

  1. Put the source model you want to quantize in models/diffusion_models
  2. Install the Krea 2 SVDQuant node pack
  3. Import the Krea2 SVDQuant Quantize node into an empty ComfyUI workspace.
  4. Select the svdq format, keep refinement at 100, pick rank 64 or 256, and queue the standalone quantizer node.
  5. You end up with a fully quantized model.

Which GPUs Benefit From Krea 2 SVDQuant?

NVIDIA Ampere cards like the RTX 3090 don’t have native FP8 Tensor Core acceleration. This means that if you’re running inference using an FP8 checkpoint, the values have to be upcast to FP16/BF16 before they can be processed by the card’s Tensor Cores. The W4A4 format can use lower-precision integer hardware that Ampere already has, but the complex scaling overhead means it requires specialized, custom kernels to actually make inference faster. And that’s the whole story.

The node pack measured a warm 8-step Turbo run at 18.80 seconds in BF16 and 7.77 seconds with rank-256 SVDQuant on its RTX 3090. That is the source of the roughly 2.4x figure. Against its INT8 build, rank-256 SVDQuant was only 1.18x faster per sampling step. These are of course Turbo benchmarks, not guaranteed Raw-finetune results.

And before you ask, yes, according to my quick tests the difference is pretty much the same when using a converted SVDQuant Raw model with a Turbo LoRA.

And a quick summary if you’re wondering whether it’s worth it to use SVDQuant on non-Ampere GPUs:

GPU familyWhat to expect
RTX 30-series, AmpereW4A4 and INT8 use hardware paths that the Ampere architecture natively supports, and the SVDQuant checkpoint is far smaller than BF16.
RTX 20-series, TuringThe checkpoint is smaller so you can benefit from less VRAM usage, but the current INT4 kernel path brings much less speed gain than on Ampere.
RTX 40-series, AdaAda has native FP8 support. SVDQuant still cuts the checkpoint size, but FP8 is still a stronger option here than it is on a 3090.
RTX 50-series, BlackwellBlackwell introduces new native FP4 paths that these specific INT4/W4A4 kernels don’t fully exploit yet, so a native FP8/FP4 implementation may still yield better performance here.
AMD, Intel, Apple SiliconThis particular Krea 2 node pack targets NVIDIA CUDA. AMD, Intel, and Apple Silicon are not supported by this implementation.

SVDQuant is not the only useful format here. The same project also supports INT8 and plain W4A4. INT8 keeps more fidelity and uses more VRAM. Plain W4A4 is smaller and faster but drops the low-rank correction branch, which is one of the best quality-preserving features of SVDQuant. Still, these are the other options to pick if you’re up for some experimentation.

What You Need Before You Start

  • A current ComfyUI install with a PyTorch build based on CUDA 13 or newer. The fast comfy_kitchen CUDA backend is disabled on older CUDA builds. Use the Krea2 SVDQuant Env Check node after installing the pack – it can tell you whether your environment is fully ready for the quantization process.
  • Krea-2-SVDQuant-ComfyUI installed in ComfyUI/custom_nodes/. You can install the node pack via a Git clone or a traditional ZIP install. Restart ComfyUI after copying it in.
  • Your chosen Krea 2 finetune in ComfyUI/models/diffusion_models/. BF16 is the best source for maximum fidelity but you can use FP8 versions too. Keep in mind that the converter will expand FP8 weights to BF16 for the conversion math, but of course it cannot restore information that was already lost in the FP8 source.
  • Enough free disk space for another 8 to 10GB checkpoint. Rank 64 is about 7.9GB and rank 256 is about 9.1GB in the project tests, so make sure you have enough space on your disk for the converted checkpoint to be saved.
Krea2 SVDQuant Quantize custom node shown in the ComfyUI node search panel.
After installing the node pack, search for Krea2 SVDQuant Quantize in ComfyUI and add it to an empty workflow.

How To Quantize a Krea 2 Raw Finetune

The conversion node works standalone, so you do not need to connect the quantizer to any generation graph for the basic conversion. Add the Krea2 SVDQuant Quantize node to an empty workflow, select your finetune, set the fields, then queue the workflow just as if you were generating a new image.

Krea2 SVDQuant Quantize node configured for a Krea 2 Raw FP8 finetune with rank 256 and 100 refinement iterations.
These are the settings I used for the Cat Tower v2.0 Raw finetune. Rank 256 is the LoRA-heavy choice, not a hard requirement.
Node settingValue to start withWhat it does
source_modelYour Raw finetuneSelect the BF16 or FP8 diffusion model file you want to convert.
formatsvdqCreates W4A4 weights and activations with the BF16 low-rank correction branch.
rank256 for LoRAs, 64 for a smaller fileRank 256 holds up better in the node pack’s LoRA tests. Rank 64 is a good smaller choice if you never load LoRAs on top.
rank_allocuniformThe default allocation. The experimental GQA allocation generally did not show a clear quality gain in the published tests.
refine_iters100Refines the low-rank branch against quantization error. Setting this to 0 makes conversion much faster, but higher ranks lose most of their point without refinement.
groupsize256The default ConvRot group size used by the pack.
variantbaseUse this for a Raw mode finetune, and you can change it for other types. This setting changes naming and metadata, not the quantization math itself.
output_nameA custom name for the outputThis is optional. A blank field makes the node derive a filename automatically.
overwritefalseLeave this off. This way the node will refuse to replace an existing file with the same name.
act_statsBlank for the simple passLeave it empty for the basic conversion shown here. A finetune-specific activation capture can improve fidelity, covered below.
seed0Keeps the randomized SVD reproducible on the same device.

For my example I used rank 256, 100 refinement iterations, group size 256, uniform rank allocation, and the base variant, leaving act_stats empty.

The pack’s own tests that you can find in the repository found little difference between rank 64 and rank 256 without a LoRA. Rank 256 becomes more useful once LoRAs are added on top of the quantized checkpoint. Thus, I would advise you to leave it at 256 if you have enough memory/disk space on hand, as you will probably end up using additional LoRAs with your checkpoint anyway.

Start the Conversion

Click the Queue Prompt button just as if you were generating a new image. The node unloads models that are already sitting in VRAM, takes over the GPU for the conversion, and blocks the queue until your new checkpoint is written.

ComfyUI console showing the Krea 2 SVDQuant conversion starting and detecting the Krea 2 architecture.
At the start, the console will show the source and output paths, selected format, rank, refinement count, and the detected krea2 architecture.

The console then will slowly report the transformer layers as they are processed. The current converter works by quantizing the 224 transformer-block linear layers and leaving smaller or more sensitive parts of the model at higher precision.

Krea 2 SVDQuant conversion in progress in ComfyUI with quantized layer counts shown in the console.
A rank-256 run with refinement takes a few minutes. The console updates the number of quantized layers as the process moves through the model.

On my RTX 3090, the example run took just under five minutes. The node pack itself quotes anything from under a minute for a quick split to roughly six minutes for a refined build on the same class of GPU.

Finished Krea 2 SVDQuant W4A4 conversion showing a 9.10GB rank-256 output file and a roughly five-minute run.
The finished Cat Tower v2.0 conversion produced a 9.10GB rank-256 SVDQuant file from the 13.16GB FP8 source.

Load the New SVDQuant Checkpoint

After the conversion is done, refresh ComfyUI’s model list. You can do this by pressing the “R” key.

Load your newly made svdq checkpoint with the Krea2 SVDQuant W4A4 Loader node. You can keep using the regular Krea 2 text encoder and VAE from your standard workflow.

If you add LoRAs on top, use the Krea2 SVDQuant LoRA Loader from the same node pack. The stock ComfyUI LoRA loader can patch current quantized Krea 2 weights, but it rewrites and requantizes the weight delta, which is rather inefficient. The dedicated loader keeps the 4-bit weight untouched and applies compatible LoRAs through a parallel branch.

Raw and Turbo Need Different Sampling Settings

Quantization quite understandably won’t turn a Raw finetune into a Turbo model. It only changes how the weights are stored and processed. The variant dropdown does not distill the model either.

Krea’s official starting point for Krea 2 Raw is 52 steps and CFG 3.5 at around 1024px. Krea 2 Turbo is the distilled model and runs at 8 steps with guidance disabled in the official code. The SVDQuant pack’s own Turbo workflow uses CFG 1.0 in ComfyUI.

Follow the recommended settings for your chosen workflow just as if you’d be using the regular, non-converted models.

SVDQuant does not require Euler, UniPC, Simple, Normal, or any other sampler and scheduler pair. If you want to experiment a bit, you can try different pairs, but be advised that not every standard pair will work here, and you can end up with solid black outputs at times.

For Better Fidelity, Capture Activation Stats From the Finetune

The simple conversion above works with an empty act_stats field. The node pack, however, has a better quality path for users who want to spend a little more time on the conversion.

Use the Krea2 SVDQuant Capture Start and Capture Save nodes with the same finetune you are converting. This calibration path loads the BF16 source, runs representative prompts, saves the activation statistics, then feeds that file into the quantizer’s act_stats input.

The activation-aware branch has the same size and runtime cost as the plain SVDQuant branch, so the extra work happens only during the build.

Quick Troubleshooting

Two basic problems are very common with this setup. One is a quantized Krea 2 build running slower than expected. The other is fuzzy or broken output on some SageAttention setups. The checks below cover both cases.

ProblemWhat to check
SVDQuant is slower than FP8 or BF16Run Krea2 SVDQuant Env Check. A PyTorch build below CUDA 13 drops the W4A4 path to a slow fallback. The Check node will tell you whether your current ComfyUI setup can make full use of the quantized model.
No real speed gain on RTX 20-seriesThis is expected with the current Turing kernels. Keep the format for memory savings, but don’t expect the Ampere-class speed-up shown in the project’s RTX 3090 benchmarks.
Black, fuzzy, or broken Krea 2 outputUpdate ComfyUI and the SVDQuant pack, then test once with SageAttention disabled. Users have reproduced fuzzy or broken Krea 2 output on some SageAttention paths, but it is not the only possible cause. Another common cause for this is using the SVDQuant loaded models with uncensor LoRAs with their weights set too high.
Out of memory on an 8GB cardTry rank 64, lower the image resolution, close other GPU-heavy programs and apps, and let ComfyUI offload parts of the model.
Pin error. appears on WindowsThis comes from ComfyUI’s pinned-memory handling. The node pack documents it as harmless, with a small load or offload speed cost. In most cases you can safely ignore it.

You might also like: Anima In ComfyUI – Quick Starter Guide

Tom Smigla
Tom Smiglahttps://tomsmigla.com/
Tom is the founder of TechTactician.com with years of experience as a professional tech journalist and hardware & software reviewer. Armed with a master's degree in Cultural Studies / Cyberculture & Media, he created the "Consumer Usability Benchmark Methodology" to ensure all the content he produces is practical and real-world focused.

Check out also:

Latest Articles