LLMLingua is Microsoft Research’s MIT-licensed prompt-compression toolkit, a Python library that deletes the tokens a small model judges unimportant before a large model ever reads the prompt.
LLMLingua is the peer-reviewed ancestor of this category’s compression products, and the newest independent evaluation hands back an uncomfortable verdict: its token pruning is exactly the kind that fails on code, the input this category exists to serve.
What it is #
A family of four methods behind one pip install llmlingua entry point: LLMLingua (perplexity-based token pruning with a small causal LM such as GPT-2 or a sub-8GB quantized Llama-2-7B, claiming up to 20x compression with little loss), LongLLMLingua (question-aware compression plus document reordering that attacks lost-in-the-middle, claiming up to 21.4 percent better RAG performance on a quarter of the tokens), LLMLingua-2 (compression distilled from GPT-4 into a BERT-class token classifier, task-agnostic, 3x to 6x faster and better out-of-domain), and SecurityLingua (security-aware compression that exposes jailbreak intent, CoLM 2025).
It is a library for pipeline builders, not an agent tool: delivery is the Python API plus integrations written into LangChain (document compressor), LlamaIndex (node postprocessor), and Microsoft Prompt flow, with no MCP server, no proxy, and no hooks.
MIT-licensed, built by the Microsoft Research group behind the papers (Huiqiang Jiang, Chin-Yew Lin, Lili Qiu, and colleagues), running anywhere a Hugging Face model runs.
Status #
Research-mature and in maintenance: 6,742 stars and 435 forks since 2023-07-07, pushed 2026-09-10, 124 open issues, as of 2026-10-10 (GitHub API).
The PyPI package has been frozen at 0.2.2 since 2024-04-09 (14 releases total), and the GitHub release train stopped at the same version, while main keeps receiving pushes that never reach pip users. The team’s news feed has moved to KV-cache infrastructure (SCBench, RetrievalAttention, MInference), the layer underneath a prompt compressor once token pruning is done. The launch HN thread (December 2023) drew 149 points and 47 comments, a discussion footprint most of this category’s young tools would envy.
Strengths #
- The evidence base is the strongest in the category: two peer-reviewed compression papers (EMNLP 2023, ACL 2024 main and Findings) plus SecurityLingua at CoLM 2025, where every other compressor here is self-benchmarked.
- LLMLingua-2’s distillation insight, formulating compression as token classification with a bidirectional encoder, is the design this category’s newer products are reinventing with worse evaluation.
- The compressors are small (mBERT or XLM-RoBERTa-large, or a sub-8GB quantized 7B for the original method), so the pruning pass fits on commodity hardware.
- Distribution already exists where long prompts live, through the LangChain, LlamaIndex, and Prompt flow integrations.
Cautions #
- The independent multi-dimensional evaluation (April 2026) found code completion, few-shot, and structure-dependent tasks degrade or fail under compression, precisely the inputs coding agents would feed it; summarization and QA held up.
- The same study found end-to-end speedups above 1.3x only in narrow conditions on optimized serving stacks, and LLMLingua-1 missed requested compression rates badly enough to make API bills and quality unpredictable; only LLMLingua-2 was called practical.
- Compression here is destructive with no retrieve path: what the small model deletes is gone, so an aggressive rate silently removes the fact the answer needed.
- The package has been frozen at 0.2.2 since April 2024 while development continues without releases, so fixes do not reach pip users, and contributions require a Microsoft CLA.
Pricing #
Free and open source under MIT; there is no paid tier, so pricing does not apply. The costs are your own hardware for the compressor model and the tokens you fail to save.
Compared to #
- Headroom: the productized successor surface (proxy, wrap, MCP) whose reversible compress-then-retrieve design answers LLMLingua’s worst property, backed by seeded self-benchmarks instead of peer review.
- rtk: compresses the other channel, command output after the model call, where LLMLingua compresses the prompt before it.
- TOON: the lossless extreme, re-encoding structured data exactly rather than deleting tokens approximately.
Bottom line #
Recommended for pipeline builders compressing very long natural-language prompts (RAG chunks, transcripts, few-shot demonstrations) who will measure downstream quality at their target ratio, ideally on LLMLingua-2. Not for code context, which the independent evidence says does not survive token pruning, and not as an agent-loop component unless the original prompt is kept. My disagreeable claim: the newer compression products in this category are selling LLMLingua’s ideas back to practitioners with weaker evaluation, and a buyer who starts from the papers will demand the reversibility the papers never offered.
Changes #
- 2026-10-10 - Created from the 2026-10-10 Meirtz/Awesome-Context-Engineering scan (the list’s LongLLMLingua entry), with the April 2026 independent evaluation recorded as the critical source.
See also #
- Headroom - the agent-native compression layer built on the same token-classification idea
- rtk - the output-side filter in the same token-bill fight
- TOON - the lossless packing counterpoint
- Context Engines Feature Matrix - the category comparison this note joins as the fifteenth column
References #
https://api.github.com/repos/microsoft/LLMLingua - repository stats, MIT license, and activity as of 2026-10-10
https://raw.githubusercontent.com/microsoft/LLMLingua/main/README.md - the method family, quick-start surfaces, framework integrations, and the KV-cache news line
https://pypi.org/pypi/llmlingua/json - the package, frozen at 0.2.2 since 2024-04-09, 14 releases
https://arxiv.org/abs/2310.05736 - the LLMLingua paper (EMNLP 2023): coarse-to-fine compression, up to 20x with little loss
https://arxiv.org/abs/2310.06839 - the LongLLMLingua paper (ACL 2024): long-context compression and document reordering
https://arxiv.org/abs/2403.12968 - the LLMLingua-2 paper (ACL 2024 Findings): distillation into BERT-class token classification
https://arxiv.org/abs/2604.02985 - the independent critical evaluation (April 2026): latency overhead, rate-adherence failures, and the code-completion degradation
https://hn.algolia.com/api/v1/items/38689653 - the launch thread, 149 points and 47 comments, December 2023
https://www.microsoft.com/en-us/research/project/llmlingua/ - the Microsoft Research project page and its stated insights