Four-bit weights used to be a hobbyist compromise — the thing you did to squeeze a 70B model onto a couple of consumer GPUs and accepted some quality loss as the price. That framing is now obsolete. OpenAI’s gpt-oss models shipped with MXFP4 native quantization of the MoE weights, evaluated in that precision, which is how a 117-billion-parameter model fits on a single 80 GB GPU. Quantization has moved from “what you do after training” to “what the model is,” and the practical question has shifted accordingly: not “should I quantize?” but “which format, which kernel, and what am I actually giving up?”
This post is a decision guide for 4-bit and 8-bit post-training quantization in 2026: how the major formats differ mechanically, why the kernel your inference engine picks matters more than the format you choose, what the quality cost really looks like, and when you should pay for 8-bit instead.
Why 4-Bit Won
The memory arithmetic is unforgiving. Weights at BF16 cost two bytes per parameter, so a 117B-parameter model needs roughly 234 GB just for weights — two to three accelerators before you have stored a single KV cache entry. At 4 bits, the same weights occupy about 60 GB, fitting comfortably on one H100 with room for activations and cache. That difference is not an optimization; it is the boundary between “needs a multi-GPU cluster” and “one box.”
Two developments made the quality question survivable. First, the formats got better: modern block-scaled schemes (GGUF K-quants, MXFP4, NVFP4) allocate scale factors per small block of weights rather than per tensor, so the effective precision is closer to 5 bits even when the payload is 4. Second, MoE architectures helped by accident — in a mixture-of-experts model only a fraction of parameters is active per token, and quantization error in rarely-used experts matters less than error in the dense layers everyone touches. OpenAI quantized the MoE weights of gpt-oss to MXFP4 while keeping sensitive components at higher precision, and published the models with evals run at that precision. Kimi followed suit with K2-Thinking, released with natively INT4-quantized weights, and DeepSeek trains and ships at native FP8 — the same philosophy (low-precision by design rather than as an afterthought), one tier higher. When frontier labs ship 4-bit as the reference precision, the “quantization is lossy” objection loses most of its force for inference workloads.
The Three Format Families
Almost everything you will encounter falls into three camps, and they were built by different communities for different hardware.
GPTQ: Error-Minimizing Rounds
GPTQ quantizes layer by layer, using second-order (Hessian) information to adjust the remaining weights each time one is rounded to 4 bits. The idea is that quantization error in one weight can be partially compensated by nudging its neighbors, so the total layer output deviates less than naive rounding would cause. It produces W4A16 models — 4-bit weights, 16-bit activations — and remains the most widely supported GPU format. The downside is that calibration is slow and the format is rigid: a GPTQ checkpoint is committed to its group size and bit width.
AWQ: Protect What Activations Touch
The AWQ paper observed that a small fraction of weight channels carries most of the quantization-sensitive signal, identifiable from activation magnitudes alone. AWQ scales those channels up before quantization and compensates afterward — no backpropagation, no Hessian, just a calibration pass over a few hundred samples. The result is usually better quality than GPTQ at the same bit width, with faster conversion. Like GPTQ it targets W4A16 GPU inference, and it is supported across most of the serving stack.
GGUF K-Quants: Block Scales for the CPU World
GGUF is not a quantization algorithm — it is the container format used by llama.cpp and its derivatives, and the weights inside are packed with GGML’s block-wise schemes. The K-quant family (Q4_K_M being the default 4-bit choice) splits each 256-weight block into sub-blocks with their own scales, mixing 4-bit and 6-bit representations so that important coefficients keep more precision. This is why a GGUF Q4_K_M file often beats same-width GPTQ on perplexity despite no error-compensation math: the effective bit rate is higher (~4.8 bits per weight versus ~4.5 for the legacy Q4_0). GGUF also supports importance-matrix variants (IMatrix), which bring AWQ-like calibration ideas into the llama.cpp world. The trade-off is ecosystem: GGUF runs best on CPU and unified-memory hardware, and GPU serving engines have historically given it second-class kernel support.
There is a fourth, newer camp: the hardware-native microscaling formats. MXFP4 (used by gpt-oss) and NVFP4 (NVIDIA’s variant with an extra precision tier in each block) store shared scale factors alongside 4-bit FP payloads, so dequantization is free inside the tensor-core pipeline. These are the formats the current hardware generation accelerates natively, and they are where new deployments of large MoE models should start.
The Kernel Matters More Than the Format
Here is the finding that reorders most people’s mental model. In serving engines like vLLM, each format can run through different matrix-multiplication kernels, and the gap between a default kernel and an optimized one dwarfs the differences between formats. The optimized Marlin kernel family achieves near-ideal speedups for FP16×INT4 workloads — on the order of 4× over FP16 compute at low batch sizes — but only if your format-and-hardware combination actually routes to it.
The practical consequences:
- The same AWQ model can run several times faster or slower depending on whether the engine selects the Marlin path or a fallback kernel. Check the engine logs at startup — the selected kernel per layer type is printed, and “using default kernel” on a quantized layer is a red flag.
- Kernel support is hardware-gated. Marlin requires SM 7.5 (Turing) or newer for standard INT4, and MXFP4 variants need Ampere or newer. Older cards silently fall back to slow paths.
- FP8 (W8A8) kernels need Ada (SM 8.9) or Hopper-class hardware. On older GPUs, INT8 W8A8 via other backends is the available 8-bit option.
So the decision order is: pick the serving engine first, check which kernels it can use on your GPU, then pick the format that routes to the fastest available kernel. Choosing “the best format” in the abstract and hoping the kernel works out is how teams end up with 4-bit models that serve slower than the BF16 original.
What Quality Actually Costs
Head-to-head comparisons of the same model in different 4-bit formats show a consistent pattern. Perplexity deltas against the BF16 baseline are small for the well-blocked formats — roughly +0.2 on WikiText-2-scale numbers — while code generation is the first capability to degrade measurably, dropping several points on HumanEval-style pass@1 at 4 bits. GPTQ tends to fare worst in these comparisons, AWQ and GGUF K-quants land close to each other, and the gap between any good 4-bit format and FP16 is real but modest for conversational use.
Three rules of thumb fall out of the published numbers:
- Conversational and summarization workloads barely notice 4-bit. Perplexity shifts of a few tenths are within run-to-run variance for many tasks.
- Code, math, and structured output are the sensitive cases. If your product generates JSON, SQL, or programs, benchmark those tasks specifically before committing to 4-bit — and consider 8-bit if the benchmarks regress.
- Exotic quants (2-bit, 3-bit) are experiments, not deployments. The quality cliff below 4 bits is steep, and the “large model at low bits beats small model at high bits” folklore only holds down to about 4.
One more trap: quantizing the KV cache is a separate decision with separate trade-offs. FP8 or INT8 KV cache roughly doubles cache capacity (more concurrent tokens per GPU) at a smaller quality cost than weight quantization of the same magnitude, and it stacks with whatever you did to the weights.
A Practical Decision Path
For a serving deployment on modern NVIDIA hardware (Ampere or newer), the path that almost always works:
# Serve a pre-quantized MXFP4 model (like gpt-oss) with vLLM
vllm serve openai/gpt-oss-120b \
--quantization mxfp4 \
--gpu-memory-utilization 0.92 \
--max-model-len 32768
# Serve a model quantized to INT4 via llm-compressor (W4A16)
vllm serve org/model-int4 \
--quantization compressed-tensors \
--gpu-memory-utilization 0.92
If the model has no official 4-bit release, quantize it yourself with the engine’s recommended toolkit rather than downloading a random community checkpoint — conversion quality varies enormously, and a badly calibrated AWQ or GPTQ checkpoint will underperform its format’s reputation. Validate with two things before rollout: perplexity on a domain-representative sample, and pass@1 on a code benchmark if the model writes code. If perplexity moved less than ~0.3 and your code benchmark moved less than a couple of points, ship it.
Reach for 8-bit when: the model is small enough that 8-bit fits anyway and you want zero-thought quality preservation; the workload is code-heavy and 4-bit benchmarks regressed; or you are already on FP8-capable hardware where W8A8 is nearly free. Stay at 4-bit when memory is the binding constraint — which, for any model above ~30B parameters on single-GPU or dual-GPU setups, it almost always is.
Wrapping Up
Quantization in 2026 is a deployment decision with well-understood edges. 4-bit is the default shipping precision for large open-weight models, the format families differ mainly in calibration philosophy and ecosystem, and the kernel your engine selects matters more than the format you picked. Choose the engine and hardware first, verify the kernel routing, benchmark code generation specifically, and treat 8-bit as the escape hatch for quality-sensitive workloads. The teams that get this wrong are the ones treating quantization as a checkbox — the ones that get it right treat it as part of the serving architecture.