<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>prompt-compression on tomrochette.com</title>
    <link>https://tomrochette.com/tags/prompt-compression/</link>
    <description>Recent content in prompt-compression on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Sun, 11 Oct 2026 02:01:18 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/prompt-compression/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>LLMLingua</title>
      <link>https://tomrochette.com/agents/context-engines/llmlingua/</link>
      <pubDate>Sat, 10 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/context-engines/llmlingua/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>context-engines</category><category>compression</category><category>prompt-compression</category><category>research</category>
      <description>&lt;p&gt;LLMLingua is Microsoft Research&amp;rsquo;s MIT-licensed prompt-compression toolkit, a Python library that deletes the tokens a small model judges unimportant before a large model ever reads the prompt.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;LLMLingua is the peer-reviewed ancestor of this category&amp;rsquo;s compression products, and the newest independent evaluation hands back an uncomfortable verdict: its token pruning is exactly the kind that fails on code, the input this category exists to serve.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A family of four methods behind one &lt;code&gt;pip install llmlingua&lt;/code&gt; entry point: LLMLingua (perplexity-based token pruning with a small causal LM such as GPT-2 or a sub-8GB quantized Llama-2-7B, claiming up to 20x compression with little loss), LongLLMLingua (question-aware compression plus document reordering that attacks lost-in-the-middle, claiming up to 21.4 percent better RAG performance on a quarter of the tokens), LLMLingua-2 (compression distilled from GPT-4 into a BERT-class token classifier, task-agnostic, 3x to 6x faster and better out-of-domain), and SecurityLingua (security-aware compression that exposes jailbreak intent, CoLM 2025).&#xA;It is a library for pipeline builders, not an agent tool: delivery is the Python API plus integrations written into LangChain (document compressor), LlamaIndex (node postprocessor), and Microsoft Prompt flow, with no MCP server, no proxy, and no hooks.&#xA;MIT-licensed, built by the Microsoft Research group behind the papers (Huiqiang Jiang, Chin-Yew Lin, Lili Qiu, and colleagues), running anywhere a Hugging Face model runs.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Research-mature and in maintenance: 6,742 stars and 435 forks since 2023-07-07, pushed 2026-09-10, 124 open issues, as of 2026-10-10 (GitHub API).&lt;/p&gt;&#xA;&lt;picture&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: dark)&#34; srcset=&#34;https://api.star-history.com/chart?repos=microsoft/LLMLingua&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: light)&#34; srcset=&#34;https://api.star-history.com/chart?repos=microsoft/LLMLingua&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;  &lt;img alt=&#34;Star History Chart&#34; src=&#34;https://api.star-history.com/chart?repos=microsoft/LLMLingua&amp;type=date&amp;legend=top-left&#34; /&gt;&#xA;&lt;/picture&gt;&#xA;&lt;p&gt;The PyPI package has been frozen at 0.2.2 since 2024-04-09 (14 releases total), and the GitHub release train stopped at the same version, while main keeps receiving pushes that never reach pip users.&#xA;The team&amp;rsquo;s news feed has moved to KV-cache infrastructure (SCBench, RetrievalAttention, MInference), the layer underneath a prompt compressor once token pruning is done.&#xA;The launch HN thread (December 2023) drew 149 points and 47 comments, a discussion footprint most of this category&amp;rsquo;s young tools would envy.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The evidence base is the strongest in the category: two peer-reviewed compression papers (EMNLP 2023, ACL 2024 main and Findings) plus SecurityLingua at CoLM 2025, where every other compressor here is self-benchmarked.&lt;/li&gt;&#xA;&lt;li&gt;LLMLingua-2&amp;rsquo;s distillation insight, formulating compression as token classification with a bidirectional encoder, is the design this category&amp;rsquo;s newer products are reinventing with worse evaluation.&lt;/li&gt;&#xA;&lt;li&gt;The compressors are small (mBERT or XLM-RoBERTa-large, or a sub-8GB quantized 7B for the original method), so the pruning pass fits on commodity hardware.&lt;/li&gt;&#xA;&lt;li&gt;Distribution already exists where long prompts live, through the LangChain, LlamaIndex, and Prompt flow integrations.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;The independent multi-dimensional evaluation (April 2026) found code completion, few-shot, and structure-dependent tasks degrade or fail under compression, precisely the inputs coding agents would feed it; summarization and QA held up.&lt;/li&gt;&#xA;&lt;li&gt;The same study found end-to-end speedups above 1.3x only in narrow conditions on optimized serving stacks, and LLMLingua-1 missed requested compression rates badly enough to make API bills and quality unpredictable; only LLMLingua-2 was called practical.&lt;/li&gt;&#xA;&lt;li&gt;Compression here is destructive with no retrieve path: what the small model deletes is gone, so an aggressive rate silently removes the fact the answer needed.&lt;/li&gt;&#xA;&lt;li&gt;The package has been frozen at 0.2.2 since April 2024 while development continues without releases, so fixes do not reach pip users, and contributions require a Microsoft CLA.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free and open source under MIT; there is no paid tier, so pricing does not apply.&#xA;The costs are your own hardware for the compressor model and the tokens you fail to save.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/context-engines/headroom/&#34; &gt;Headroom&lt;/a&gt;: the productized successor surface (proxy, wrap, MCP) whose reversible compress-then-retrieve design answers LLMLingua&amp;rsquo;s worst property, backed by seeded self-benchmarks instead of peer review.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/context-engines/rtk/&#34; &gt;rtk&lt;/a&gt;: compresses the other channel, command output after the model call, where LLMLingua compresses the prompt before it.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/context-engines/toon-format/&#34; &gt;TOON&lt;/a&gt;: the lossless extreme, re-encoding structured data exactly rather than deleting tokens approximately.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for pipeline builders compressing very long natural-language prompts (RAG chunks, transcripts, few-shot demonstrations) who will measure downstream quality at their target ratio, ideally on LLMLingua-2.&lt;/strong&gt;&#xA;Not for code context, which the independent evidence says does not survive token pruning, and not as an agent-loop component unless the original prompt is kept.&#xA;My disagreeable claim: the newer compression products in this category are selling LLMLingua&amp;rsquo;s ideas back to practitioners with weaker evaluation, and a buyer who starts from the papers will demand the reversibility the papers never offered.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-10 - Created from the 2026-10-10 Meirtz/Awesome-Context-Engineering scan (the list&amp;rsquo;s LongLLMLingua entry), with the April 2026 independent evaluation recorded as the critical source.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/context-engines/headroom/&#34; &gt;Headroom&lt;/a&gt; - the agent-native compression layer built on the same token-classification idea&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/context-engines/rtk/&#34; &gt;rtk&lt;/a&gt; - the output-side filter in the same token-bill fight&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/context-engines/toon-format/&#34; &gt;TOON&lt;/a&gt; - the lossless packing counterpoint&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/context-engines/context-engines-feature-matrix/&#34; &gt;Context Engines Feature Matrix&lt;/a&gt; - the category comparison this note joins as the fifteenth column&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/microsoft/LLMLingua&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/microsoft/LLMLingua&lt;/a&gt; - repository stats, MIT license, and activity as of 2026-10-10&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/microsoft/LLMLingua/main/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/microsoft/LLMLingua/main/README.md&lt;/a&gt; - the method family, quick-start surfaces, framework integrations, and the KV-cache news line&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://pypi.org/pypi/llmlingua/json&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=pypi.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://pypi.org/pypi/llmlingua/json&lt;/a&gt; - the package, frozen at 0.2.2 since 2024-04-09, 14 releases&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2310.05736&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2310.05736&lt;/a&gt; - the LLMLingua paper (EMNLP 2023): coarse-to-fine compression, up to 20x with little loss&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2310.06839&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2310.06839&lt;/a&gt; - the LongLLMLingua paper (ACL 2024): long-context compression and document reordering&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2403.12968&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2403.12968&lt;/a&gt; - the LLMLingua-2 paper (ACL 2024 Findings): distillation into BERT-class token classification&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2604.02985&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=arxiv.org&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://arxiv.org/abs/2604.02985&lt;/a&gt; - the independent critical evaluation (April 2026): latency overhead, rate-adherence failures, and the code-completion degradation&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/items/38689653&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/items/38689653&lt;/a&gt; - the launch thread, 149 points and 47 comments, December 2023&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.microsoft.com/en-us/research/project/llmlingua/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=www.microsoft.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://www.microsoft.com/en-us/research/project/llmlingua/&lt;/a&gt; - the Microsoft Research project page and its stated insights&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
