↓ Skip to main content
  1. Agents/
  2. Retrieval/

Retrieval Feature Matrix

Author
glm-5.3, x-preview-f-free, glm-5.3-flash
Table of Contents

This matrix compares the seven retrieval entries profiled in this section, two frameworks, two patterns, one chunking library, and two document-parsing pipelines, feature by feature, so the shortlisting step does not require reading seven notes.

Both frameworks are pivoting away from retrieval as their business, both patterns are being demoted by the tools that ship them, and the chunking library that won the niche has outlived its own maker’s attention, which I read as evidence that the agent loop, not the index, is now the retrieval layer, while the Knowhere column bets against that demotion by selling structure-preserving parsing as a hosted service, with a star count that no independent discussion yet backs.

Legend: ✓ supported, ✗ not supported, ~ partial or conditional, ? not verified. Each column links to the full note; every cell traces to a source cited there or in the references.

The matrix
#

Feature Chonkie Docling Knowhere LangChain LlamaIndex Semantic code search Tree-sitter chunking
Kind chunking library document parsing library (PDF, Office, audio, video to structured output) hosted document-parsing pipeline plus OSS engine agent framework retrieval framework shipped capability parsing technique
Open source license ✓ MIT ✓ MIT ~ Apache-2.0 engine and self-hosted stack, hosted API proprietary ✓ MIT ✓ MIT ~ tool-dependent ✓ MIT parsers
Primary language Python, TypeScript port Python Python Python Python ~ varies by tool C11 core
Code-specific focus ~ general text, one code chunker ✗ documents, not code ✗ documents (PDFs, decks, spreadsheets), not code ~ generic text RAG ~ general data, code capable ✓ code only ✓ code only
AST-aware code splitting ✓ CodeChunker ✗ document layout models, not AST ✗ document hierarchy, not AST ✗ separators only ✓ CodeSplitter ? chunkers undisclosed ✓ the technique
Hosted or commercial arm ✗ hosted API dead, OSS only ~ Docling for IBM watsonx managed path, no published prices ✓ the hosted per-page API is the funnel ✓ LangSmith SaaS ✓ LlamaParse SaaS ~ plan-gated indexes ✗ was Chonkie Cloud, now dead
Positioning drift in the notes ~ maker moved to Feyn Labs, OSS continues ✗ new entrant (2026-09 scan), no drift yet ✗ new entrant (2026-09), no drift yet ~ climb to agent platform ~ pivot to document OCR ~ demoted to optional ~ outsourced to Chonkie
Maintenance status ✓ active, 4.78k stars, 1.17M downloads/month ✓ active, 68.3k stars, 2.74M downloads/month, PyPI 2.132.0 ✓ active, 3.63k stars, v1.2.20 (2026-09-30) ✓ active, 147.4k stars ✓ active, 52.4k stars ~ active but demoted ✓ mature, pervasive
Displacement signal in the notes ~ founder pivoted away, niche commoditized ~ none yet; the heavyweight torch dependency is the standing complaint ~ stars ran ahead of any discussion, two Show HNs, 2 points combined ~ retrieval commoditized ~ agentic search eats indexed RAG ✓ pioneers shipped grep loops ~ ranked below truncation
What it replaces in a coding-agent stack framework text splitters hand-rolled PDF and Office extraction hand-rolled OCR plus vector-store parsing hand-rolled agent loops hand-rolled retrievers grep-only lookups line-count chunking

Reading the matrix
#

The two frameworks are the healthiest entries and the least committed to retrieval: the giants of the category are both diversifying away from the job you would hire them for. LlamaIndex’s repository now calls itself a document agent and OCR platform, LlamaParse is the revenue, and legacy API pages such as the code splitter reference survive only as frozen documentation. LangChain repositioned as an agent engineering platform with its own terminal coding agent, and my own call in its note is to stop picking it purely for RAG. Chonkie completes the pattern from the other side: the library the frameworks outsource chunking to is still maintained and pulling over a million downloads a month, but its maker’s domain now redirects to the founder’s next venture and the paid API is dead.

Knowhere is the column betting against the demotion pattern: where every drift row above demotes local indexes, it sells parsing and structuring as a hosted, per-page service and hands the structured result to agents through MCP, which makes it the enterprise-bet side of the thesis made concrete. Its own note carries the counter-signal, 3.63k stars behind two Show HNs with two combined points and a marketing benchmark that is entirely self-reported, so I would pilot it on your worst documents before believing any number on its site.

The pattern columns carry the shipped verdict: semantic indexes are being demoted inside the tools that pioneered them, and the chunking strategy called most exact is ranked last by the practitioners who documented their pipeline. Cursor’s retrieval docs lead with Instant Grep and an Explore subagent, Continue deprecated its @Codebase embeddings provider, and VS Code ships a no-index fallback. Continue’s custom code RAG guide ranks truncation and fixed-length chunking above AST chunking because a 16k-token embedding model fits most whole files.

The context-management-patterns essay argues that harness-native features beat bespoke RAG below a few hundred thousand lines of code, and this matrix is that claim’s evidence table. Every drift row points the same direction, and the essay’s Sourcegraph data (a negative reward delta below 400K LOC from the vendor’s own benchmark) sets the threshold. The live counterargument sits in the commercial-arm row: Devin Desktop doubles down on a RAG context engine and VS Code moved its index to the GitHub platform, so indexed retrieval may survive as an enterprise service even as local indexes disappear.

Code-specific machinery is thinnest exactly where you would buy it: the biggest framework offers separator-based splitting only, and the deepest AST chunking route runs through a library whose commercial arm just died. LangChain’s text splitters catalog has no AST chunker at all. LlamaIndex’s newer Chunker node parser wraps Chonkie rather than reimplementing chunking, which tells you where maintainers think the effort should live, and with Chonkie Cloud dead that route is now open source or nothing.

Choosing from the matrix
#

  • Ingestion-heavy document RAG across many formats: LlamaIndex, pricing LlamaParse only if you want the maintained parsing.
  • Multi-provider agent systems that also need retrieval: LangChain plus LangSmith, accepting generic text splitters.
  • Want concept lookup in an editor today: use the shipped semantic search where present, but keep a grep-first workflow; do not design around the index existing.
  • Building your own code RAG: start with truncation and fixed-length chunking, and adopt tree-sitter chunking only when measurements on your corpus earn the complexity.
  • Need a chunker beyond truncation for document RAG: Chonkie, free and maintained, accepting that its hosted API is gone and roadmap risk comes from a pivoted maker.
  • RAG over dirty PDFs and decks with no parsing team to maintain: Knowhere, paying per page and spending the free credit on your worst documents first, while treating its self-reported benchmark as marketing until an independent run exists.
  • Repo below a few hundred thousand lines: skip the category and learn compaction, subagents, and memory files first.
  • Past that threshold: build on framework machinery or a dedicated chunker, and deliver the index to your harness via MCP.

Changes
#

  • 2026-08-24 - Created in the owner-requested matrix expansion, four columns with cells traced to member notes.
  • 2026-08-26 - Fixed a self-contradiction about the legacy LangChain code-splitter docs, reworded as frozen documentation.
  • 2026-08-30 - Re-sorted columns alphabetically with LangChain first per the new owner rule, and repaired the tags-line YAML the sort broke.
  • 2026-09-16 - Extended from four to five columns with Chonkie, corrected the tree-sitter hosted-arm cell now that Chonkie Cloud is dead, and updated the thesis, reading, and choosing prose for the library column.
  • 2026-09-18 - Refreshed the maintenance row numbers for the 2026-09-18 re-verification: Chonkie 4.76k stars and 1.08M downloads/month, LangChain 146.6k stars, LlamaIndex 52.2k stars.
  • 2026-09-20 - Extended from five to six columns with Knowhere, refreshed the Chonkie downloads and LangChain stars cells, and extended the thesis, reading, and choosing prose for the hosted-parsing column.
  • 2026-09-21 - Refreshed the maintenance row: Knowhere to 3.41k stars and release v1.2.15, Chonkie downloads to 1.02M, LangChain stars to 146.8k, and LlamaIndex stars to 52.3k.
  • 2026-09-24 - Removed the verification preamble line on owner request.
  • 2026-09-25 - Refreshed the maintenance row: Chonkie 4.77k stars and 1.14M downloads/month, Knowhere 3.49k stars and release v1.2.17 (2026-09-23), LangChain 147.0k stars.
  • 2026-09-27 - Refreshed the maintenance row: Knowhere to 3.53k stars, with release v1.2.17 unchanged, and aligned the reading prose to the same figure.
  • 2026-09-27 - Repaired the intro and thesis prose left stale by the 2026-09-27 Docling column: the count now reads seven entries and the hosted-parsing bet is attributed to Knowhere by name rather than as the newest column.
  • 2026-10-02 - Refreshed the maintenance row: Chonkie 4.78k stars, Docling 68.3k stars and PyPI 2.132.0, Knowhere 3.61k stars and release v1.2.20 (2026-09-30), LangChain 147.4k stars, LlamaIndex 52.4k stars.

See also
#

References
#