PageIndex - Vectorless, Reasoning-Based RAG for Long Documents

PageIndex replaces vector databases with a hierarchical tree index that an LLM reasons through, much like a human expert navigating a long report. Instead of similarity search, it retrieves by relevance, making results traceable and explainable with no chunking required. Install with pip install -U pageindex and index, retrieve, and chat locally with your own LLM key, or use PageIndex Cloud for OCR, image understanding, and managed storage. It excels on financial reports, legal documents, and technical manuals, reaching 98.7% accuracy on FinanceBench while cutting query costs versus native PDF input.
Similarity ≠ relevance — what retrieval actually needs is relevance, and relevance requires reasoning.