Validating Low-Bit LLM Quantization: Separate Weight Compression, KV Cache and Business Regressions

A smaller four-bit model file does not guarantee proportional runtime-memory savings, faster inference or unchanged quality. This article distinguishes weights, activations and KV cache, then develops a reproducible memory example, calibration-data requirements and controlled comparisons. Paired business cases reveal regressions that average scores can conceal.

Quantization acceptance involves three separate questions: what was compressed, whether the target runtime actually uses that compression, and which tasks changed afterward. Focusing on post-training quantization, this article offers constructed calculations and comparison procedures for adopting, restricting or rolling back a configuration. No models were benchmarked; all capacity and quality numbers illustrate the method only.

Specify the complete precision configuration, not just 'four-bit'

Quantization maps continuous or high-precision values into a finite representation. In a simplified uniform four-bit scheme, sixteen codes cover a selected interval. Widening that interval increases spacing between adjacent values; narrowing it can clip values outside the range. Real implementations may use groups, scales, zero points and nonuniform representations, so bit width alone cannot determine error.

Weight quantization concerns persistent model parameters. Activation quantization concerns intermediate values during inference. KV-cache quantization concerns the historical keys and values used by attention. W4A16 usually indicates four-bit weights and sixteen-bit activations; it does not also specify the cache precision. Record all three separately, including layers retained at their original precision.

Distinguish FP16, BF16, INT8 and different FP8 formats. Equal bit widths can have different numerical ranges, rounding behavior and computational support. There is no universally recommended format here: evaluate each candidate together with its architecture, target hardware, inference engine and exact version.

GPTQ and AWQ explain mechanisms, not guaranteed deployment performance

The original GPTQ paper proposes post-training weight quantization using approximate second-order information. Its limitations explicitly distinguish reduced memory movement from reduced computation and note that the study did not cover activation quantization.[1] Smaller weights therefore do not imply a proportional reduction in task latency.

The AWQ project describes observing activations to select channel scaling that protects weights important to quantization error.[2] Using activation information to select weight representations does not mean that AWQ automatically quantizes the activations themselves. The calibration procedures and runtime kernels of different methods are not interchangeable switches.

Software constraints also determine practical availability. vLLM's quantization documentation lists support by implementation and hardware and warns that the matrix changes.[3] Loading a file does not necessarily establish use of the intended low-bit kernel. Record the actual backend, hardware, drivers and framework versions, and check for format conversion, fallback computation or unsupported layers. The online documentation was consulted on October 10, 2026; it cannot replace a frozen deployment environment.

Memory example: smaller weights do not remove long-request state

Consider a constructed model with eight billion parameters. Excluding metadata, sixteen-bit weights contain approximately 16 billion bytes, or 16 GB; four-bit weights contain about 4 GB. This is raw payload arithmetic using decimal GB. Scales, zero points, unquantized layers, alignment and copies created during loading change actual occupancy.

Assume further a conventional full-attention KV cache with 32 layers, eight KV heads per layer, head dimension 128, and two bytes per value for both keys and values. Cache bytes per token equal two times the number of layers, times KV-head count, times head dimension, times bytes per value. The result is 131,072 bytes, or 0.125 MiB per token.

A request retaining 8,192 tokens then has about 1 GiB of raw KV cache. Eight equally long requests without shared cache require about 8 GiB. These are binary GiB, which must not be confused with the decimal GB used above. Retained length includes processed input and generated output still needed by attention; input length alone does not describe the eventual peak.

This simplified model excludes block-allocation waste, prefix sharing, sliding windows, eviction, tensor-parallel sharding and other architectural differences. It illustrates independent growth of weights and request state, not a purchasing specification for an actual model. Runtime memory must additionally accommodate workspaces, activations, engine reservations and other processes, and be measured under real workloads.

Calibration should represent the workload without becoming the final exam

For methods requiring calibration data, cover the actual serving distribution: principal languages, short questions, long documents, number-heavy text, structured outputs and tool arguments. Calibrating only on brief chat and deploying long-document extraction does not establish quality evidence for the target task. More calibration examples cannot compensate for missing types.

Retain the tokenizer, chat template, truncation direction and maximum length. Incorrect truncation removes business information, which is a different fault from numerical approximation. Fix input processing first. Manage calibration samples, development cases used to choose the configuration and final acceptance cases separately; repeated tuning on the same cases is not independent validation.

Treat KV cache separately. vLLM documents cache quantization and scale-calibration options, with some configurations constrained by attention-backend support.[4] Fix the weight scheme before comparing native-precision and candidate caches. Changing weights, activations and cache together makes it difficult to locate the source of long-context regressions.

Retain calibration-data versions, quantization tools and settings, original-weight identifiers, output checksums and runtime loading configuration. The same 'four-bit' label can describe different group sizes, retained layers and kernels. Without those records, reproduction and rollback become guesswork.

Paired comparisons should expose new errors, not just average-score differences

Define a native-precision reference and hold weight provenance, tokenizer, prompt template, retrieved inputs, tool definitions, generation limits and sampling conditions constant. Change one quantization factor in the first round and run identical tasks case by case. Fix random seeds where feasible, but do not automatically classify differences between kernels as quantization errors: apply the business acceptance standard.

In a constructed set of 1,000 tasks, both configurations pass 900; the reference alone passes 25; the candidate alone passes 15; and both fail 60. Reference acceptance is 92.5%, candidate acceptance 91.5%, a one-percentage-point net difference. Nevertheless, the candidate introduces 25 failures; reporting only ten fewer passes hides that distinction.

Inspect those 25 cases for categories and consequences: monetary units, negation, required arguments, evidence near the end of a long document, or answering when abstention was required. Unacceptable critical errors cannot be offset by gains elsewhere in an average. These figures describe no actual model, and the sample count alone is insufficient to establish a reliable confidence interval.

The reference is not absolute ground truth. The 60 shared failures belong in a separate system-defect list; the candidate's 15 repairs also need independent evidence. Adjudicate disputed labels before scoring, and use blind review or explicit rubrics for open-ended outputs. Repeat stochastic generations where relevant and report task counts separately from generation counts, avoiding treatment of correlated repeats as independent business cases.

Hold workload constant before attributing a speed improvement

After quality passes, compare reference and candidate on the same hardware, concurrency and length distribution, separating short and long inputs and outputs. Report client-observed time to first token, complete-task duration, output rate, peak memory and failures. Specify percentiles, warm-up conditions and test duration.

First hold concurrency constant to observe the effect of substitution. Only in a second round use released memory to increase concurrency and establish effective capacity under the defined quality and latency requirements. These rounds answer different questions. Comparing increased-batch candidate throughput with reference single-request latency does not show that every user receives a faster response.

If the low-bit configuration is slower, inspect kernel support, dequantization overhead, input-processing versus generation stages, batching conditions and waiting in other components. Quantization may be more useful in memory-bandwidth-limited phases; if retrieval or tools dominate, weight compression may not address the main bottleneck. These are hypotheses to measure, not universal acceleration ratios.

Service-capacity acceptance for private deployment can incorporate the findings: attach the accepted precision configuration to permitted context lengths, concurrency, task mix and timing commitments. State both the stable operating range and the rules for throttling, queueing or changing routes outside it.

Release a bounded configuration with reproducible rollback

Make one of three decisions. If all critical segments and workload conditions pass, adopt within that scope. If only a subset passes, restrict tasks or context length explicitly. If core quality or stability requirements fail, retain the reference. A smaller file alone is not a release criterion.

The release package should contain original and quantized model identifiers, complete precision and calibration settings, the software environment, paired case results, workload traces, a rollback target and triggers. A later change to cache precision, engine version or hardware creates a new configuration requiring at least targeted retesting.

Rollback involves more than keeping old weights. Confirm that the current engine can load the old format, enough memory remains available, and queued requests have defined handling. A resource-constrained single machine may need an agreed maintenance window or human takeover. Discovering only after failure that the native-precision version can no longer start is not a recovery plan. Reproducible calculations, diagnosable differences and executable rollback turn low-bit optimization into an engineering deliverable.

References

  1. [1] Frantar et al. — GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (ICLR 2023, v2)
  2. [2] MIT Han Lab — AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
  3. [3] vLLM — Quantization (documentation consulted 10 October 2026)
  4. [4] vLLM — Quantized KV Cache (documentation consulted 10 October 2026)
Back to insights
鲁ICP备2024109755号-2
Drag to move. Right-click, touch and hold, or press Shift+F10 to choose a corner.