PageIndex
Type: agentic-systems/types/generated-review.md
Evidence basis: code-grounded analysis of https://github.com/VectifyAI/PageIndex at d2693d80791a86345ef78b3234834f5fe53a70a0, with an analysis cutoff of 2026-09-29.
Evidence basis and boundary. The analysis covers the open-source PageIndex Python SDK at commit d2693d80791a86345ef78b3234834f5fe53a70a0 in local mode, whole-system, code-grounded. Every operation conclusion is uninspected: no indexing run, test run or model call was made, so everything below about what the code does is implementation reading (SRC-1) and design prose (SRC-2). Accuracy, cost and timing figures are README reports only (SRC-3; CLM-2, CLM-3). The cloud service, the hosted App, host agent runtimes and model providers are outside the boundary.
Characterization and claimed work. PageIndex claims to replace vector similarity with an agent that reasons over a document tree (CLM-1) and to run entirely on the caller's machine with the caller's model key (CLM-4). The code supports the local-storage claim at the implementation level and contains no embedding or vector component (ABS-4). Its retrieval claim is a claim about model behavior, which no record here observed.
Indexing. A caller submits a PDF. The default flash route (RTE-1) builds a section tree from layout statistics and, by default, embedded bookmarks, then a model writes node summaries and one document description. Bookmarks are validated for range and order and merged by tier (EPI-RTE-1); the flash structure comes from program code, and the model supplies summaries and expand proposals (CMP-1). A refinement pass (RTE-8) lets the model propose child headings for large nodes, keeps a proposal only if its title string occurs in the index-time page text, and then a cost rule expands the node or keeps it collapsed. The standard route (RTE-3) has the model propose the structure and then check its own entries against page text, with accept, repair and fall-back thresholds (RTE-9). Repair can return a tree that still holds entries judged wrong, with no flag in the stored tree (EPI-OBJ-5). The model-written summaries and description are never compared with their pages (EPI-ABS-1, EPI-RTE-2). Scanned or layout-less long PDFs are refused rather than stored in degraded form (RTE-1).
Storage. Indexing writes three JSON files per document and a manifest cache: the tree (OBJ-1), the page text from a second extractor (OBJ-2) and metadata with the model-written description and any caller-supplied metadata (OBJ-3). The tree is written before use and never changed by use. The store grows by submission and shrinks by deletion, and a changed document is deleted and resubmitted (ABS-1, MEM-ABS-1). The only other writer is a CLI trace log that nothing reads (MEM-OBJ-1). Index-time text and query-time text come from different parsers and are reconciled only by page-range bounds (EPI-OBJ-9, OBJ-2).
Query-time retrieval. A chat call runs a single model agent over four read-only tools that list documents, return an outline without node text, and return page text up to a character budget (RTE-2). Local discovery has no ranking: the model matches document names and descriptions itself. When the caller supplies a document id, code adds that document's metadata as the first user message (MEM-RTE-1) and restricts tool lookups to it. That restriction holds only on the chat lanes of a local client; the host adapters have no document scope and can expose irreversible deletion when the host opts in (RTE-4, RTE-5). The instruction and tool text is prompt-level guidance (BAP-1, OBJ-4). Stored summaries, page text and the description reach the model unframed, so the model reads them as data with no injection defense at query time (ABS-3, BAP-2). The memory profile classifies this store as files, natural-language and symbolic, imported, compiled and authored, read back by pull plus one identifier-selected push, with no curation and no trace learning; the profile treats the store as memory because each submission adds a retained document, and that scope choice is stated in the profile (see Limitations).
Answers and citations. With citations requested, the model is instructed to emit page tags. The code resolves each cited document name to an id and does not check that the page exists or supports the claim (EPI-RTE-3, EPI-ABS-2). Answers reach the caller with no check or acceptance step (EPI-OBJ-6). The README's claim of traceable and explainable retrieval (CLM-5) holds for the addressability of a tag, not for verification of the cited statement.
Theory-builder conditions, learning, reflection, autonomy. The only stated theory the system applies is the search-cost model in RTE-8. Condition 1 (localized content) is wired: the rule and its constants are separately named in one file. Condition 2 (the theory guides what the system does) is wired: expansion depends on the formulas. Condition 3 (criticism aimed at the theory's content) is absent, and condition 4 (iteration on the result of criticism) is absent (ABS-2). The standard-route verify and repair step (RTE-9) is not a theory route. Whether criticism of a consumed theory improved capacity for future action is therefore absent, as is learning, and the system is not reflective (ABS-1, ABS-2). Autonomy in RTE-8 is stated role by role: a model proposes, computation decides, the user configures inputs; that description is wired and is not a claim of autonomy for the whole system. Self-improvement at the declared boundary is absent. Repository commit history, where maintainers may revise these rules, was not examined.
Scenario-relative assessment. For a caller who indexes a text-layer PDF and needs a stable, inspectable, per-document tree that a model can navigate, the code delivers one built once, stored as plain files, read through a small read-only tool set, without embedding or ranking machinery. For a scenario in which retrieved content must be correct, PageIndex supplies no verification: summaries, model-verified structure and answers carry no independent check, and the warrant for an answer stays with the caller. For a scenario that expects the system to improve with use, nothing here adapts.
What would change this assessment. A run of the indexing routes on PDFs with known outlines, compared with stored summaries and trees, would give observed evidence on structure accuracy and summary fidelity (EPI-ABS-1). A run of the chat lanes with logged tool calls would show whether the model follows the reading workflow and whether the outline changes which pages it reads (RTE-2, MEM-RTE-1). A run comparing answers with and without retrieved pages would test dependence on retrieved content (MEM-ABS-2). The external benchmark repositories, if frozen and tied to this commit, would let the README figures be judged (CLM-2, CLM-3). The hosted service's code or served instructions would be needed for any cloud conclusion (RTE-6, BAP-4).
Limitations
| limitation | affected source, record, or route IDs | inspected boundary | conclusion prevented | evidence that would resolve it |
|---|---|---|---|---|
| No indexing, test or model run was made, so every operation conclusion is uninspected | SRC-1, RTE-1, RTE-2, RTE-3, RTE-4, RTE-5, RTE-7, RTE-8, RTE-9, EPI-RTE-1, EPI-RTE-2, EPI-RTE-3, MEM-RTE-1 | pageindex/ and run_pageindex.py read statically at the frozen commit |
Any statement that a route behaves as coded, that models follow the instructions, or that the summaries or outline change answers | Run the test suite, indexing runs and logged chat runs at the frozen commit |
| README accuracy, cost and timing figures come from external benchmark repositories that are not frozen | SRC-3, CLM-2, CLM-3 | README.md and chart assets at the frozen commit |
Any observed or causal claim about accuracy, cost or comparison with vector retrieval or native PDF input | Freeze the benchmark runners and data at a commit tied to this SDK and run a one-component contrast |
| Example outputs carry no run configuration, model or code revision | SRC-4, OBJ-1 | examples/documents/results/ at the frozen commit |
Attributing the sample trees to the frozen pipeline, a mode or a model | Regenerate sample outputs with a recorded configuration |
| PageIndex Cloud, the hosted App and the cloud MCP server are excluded; only client code was read | RTE-6, BAP-4, CLM-1, CLM-5 | pageindex/cloud_api.py, pageindex/mcp_bridge.py as the client half |
Any conclusion about cloud indexing, storage, ranking, served instructions or the managed chat | Inspect the hosted service or observed cloud runs |
| Host adapters and protocol lanes were read only as far as the tool wrappers; host loops, approvals and persistence belong to host runtimes | RTE-4, RTE-5 | pageindex/integrations/ and the lane entry points |
Host-side memory, approval or retry behavior, and provider prompt-cache behavior | Trace a host runtime that uses the adapters |
| The memory profile treats submission as use, so the document store is inside the memory boundary; under a reading that counts only query-time writes the boundary would be empty | MEM-RTE-1, OBJ-1, OBJ-2, OBJ-3, RTE-1, MEM-ABS-1 | Store and its write and read routes | Any conclusion that PageIndex learns or accumulates memory through querying; and the positive profile values if the stricter reading were used | A type-level ruling on whether submission counts as use, or a comparison set that fixes the boundary rule |
The manual write-agency value rests on one caller-supplied metadata field, and the identifier read-back signal on a caller-supplied document id |
OBJ-3, MEM-RTE-1 | pageindex/local_api.py, pageindex/agent_tools.py |
Whether these values would survive a stricter reading of retained memory | A stricter definition of retained memory, or a second write route beyond submission |
The trace-learning negative is encoded with the basis wired, because the evidence vocabulary has no negative |
ABS-1, MEM-ABS-1, MEM-OBJ-1 | Bounded code searches of pageindex/ and the CLI |
Reading wired as a positive route; the negative rests on searches only |
A vocabulary value for bounded negatives, or a run showing no trace-fed writes |
| Absences rest on regex searches of listed terms, so an unlisted synonym could be missed | ABS-1, ABS-2, ABS-3, ABS-4, MEM-ABS-1, MEM-ABS-2, EPI-ABS-1, EPI-ABS-2 | The file sets and term lists in each record | Certainty that no write-back, curation, sanitization, vector component or check exists | A different search method or a code review of the same files |
| Repository commit history and human edits to rules were not examined | ABS-2, RTE-8, RTE-9 | Frozen commit only | Whether maintainers revise the cost rule or prompts from evidence over time, which bears on theory-builder conditions 3 and 4 at repository level | Read the commit history and issue record of the rule files |
| Index-time page text (pdfium) and query-time page text (PyPDF2) come from different parsers, and their agreement is unchecked | EPI-OBJ-9, OBJ-2, EPI-OBJ-4 | pageindex/local_api.py, pageindex/flash/ as consumed |
Whether a heading confirmed at index time appears in the text the agent reads, and extraction fidelity | Compare the two extractions on sample PDFs |
| Summary and description fidelity, layout heading precision and bookmark correctness were not tested | EPI-OBJ-1, EPI-OBJ-3, EPI-OBJ-8, EPI-OBJ-2 | Consumers of these outputs only | Whether the model-written and detected structure is accurate | Compare stored summaries and trees with their pages and printed outlines |
| The Markdown pipeline was traced only as far as structure, and its output has no shipped reader | RTE-7 | pageindex/page_index_md.py and run_pageindex.py |
Any conclusion about its content routes | Trace the pipeline and any consumer |
| The four amendments correct records that stay unchanged in their members | RTE-2, CLM-5, RTE-8, ABS-1 | This Reconciliation | A reader of a member alone sees the uncorrected wording | Read the amendments with the members |