Membership Inference Attacks on Quantized LLMs

The problem

Post-training quantization (AWQ, GPTQ, bitsandbytes NF4/LLM.int8()) is usually framed purely as a compression and efficiency technique. Whether it also changes a model’s exposure to membership inference attacks, where an attacker tries to determine if a specific example was in the training set, was largely untested. If quantization happened to weaken that signal as a side effect, that would matter for anyone deploying compressed models trained on sensitive data. If it didn’t, that assumption needed to stop being taken for granted.

Key finding

Signal preserved
4-bit and above: p < 0.05
All three quantization frameworks (AWQ, GPTQ, bitsandbytes) keep a statistically significant membership signal at 4-bit precision and above. Compression alone does not protect training-data privacy.
Weakened, not eliminated
2-bit / 3-bit GPTQ
Only the most aggressive GPTQ settings measurably weaken the membership signal. Even there the signal is reduced, not removed, so the exposure does not go away.

What was built

Membership-signal vector, per sequence36 features
Min-K% · 11
Max-K% · 11
Perturbation ppl · 12
Token-level loss 1
Min-K% scores 11
Max-K% scores 11
ZLIB compression ratio 1
Perturbation-based perplexity 12
Model: Pythia-12B
Framework: extends Maini et al. Dataset Inference
Tests: two-sample t-test + Cohen's d
Domains
6 PILE domains
Compute
SLURM HPC cluster
Hardware
NVIDIA A100 80GB
Score
Linear regressor
Wikipedia
ArXiv
GitHub
Common Crawl
HackerNews
Math
An IID-controlled dataset-inference pipeline: a linear regressor combines the 36 features above into a single membership score per sequence, and the t-test with Cohen's d checks whether that score separates members from non-members across quantization configurations.
FPpreserved
4-bitpreserved
3-bitweakening
2-bitweakening
From full precision down to 4-bit, the membership signal stays statistically significant (p < 0.05) across all three frameworks. The weakening only appears at the most aggressive GPTQ 3-bit and 2-bit settings, and even there the signal is reduced rather than eliminated.

A methodological finding along the way

Artifact ruled outCorpus-level leakage was distribution shift, not memorization
Apparent reading: large cross-corpus leakage differences reflect true memorization
→
Actual cause: residual distribution shift between member and non-member splits
Diagnosed with a secondary spaCy analysis of five corpus-level attributes:
7-gram overlap Named-entity density Mean dependency length Lexical specificity Vocabulary diversity
A pitfall worth flagging for anyone evaluating membership inference attacks across heterogeneous corpora.

Status

This is an NYU Abu Dhabi capstone project, currently under peer review at a security-focused academic venue (venue withheld per double-blind submission policy). Code and full results will be published once the review concludes.

Technologies: Python PyTorch Hugging Face Transformers spaCy NumPy Pandas

Concepts / Algorithms: LLM ML Security ML Privacy Membership Inference Attacks Quantization (AWQ, GPTQ, bitsandbytes) Statistical Testing