Skip to main content
The storage layer is the bridge between the SDK’s Python objects and FalkorDB’s graph database. Two classes handle everything: GraphStore manages node and relationship writes, and VectorStore manages embeddings, indexes, and search. A third class, EntityDeduplicator, handles post-ingestion entity merging. This document explains how each one works, what Cypher queries they generate, and how to tune them.

The Big Picture


GraphStore — Writing Nodes and Relationships

GraphStore is the single write path for all graph data. It uses parameterized Cypher to prevent injection and batched UNWIND for performance.

Batched UNWIND MERGE

Nodes and relationships are written in batches of 500 items per query. The pattern:
Why MERGE, not CREATE? MERGE is idempotent — if a node or relationship already exists, it updates the properties rather than creating a duplicate. This makes re-ingestion safe.

Label Hints

Relationship MATCH queries use label hints to speed up endpoint lookups. Instead of searching all nodes, FalkorDB only looks at nodes with the specified label: Unknown edge types default to (__Entity__, __Entity__).

Per-Item Fallback

If a batch upsert fails (e.g., a single malformed property causes the whole batch to error), GraphStore falls back to per-item upserts. The behavior differs slightly: for nodes, the first per-item failure raises a DatabaseError (remaining items in that batch are not attempted); for relationships, failures are logged as warnings and processing continues through the batch.

None-ID Guard

Before writing, nodes with None or empty IDs are filtered out. These come from bad LLM extraction (the LLM sometimes returns entities without proper names). The guard prevents phantom nodes from polluting the graph.

Property Cleaning

All properties go through _clean_properties() before writing:

Cypher Safety

Node labels and relationship types are sanitized via sanitize_cypher_label() to prevent Cypher injection. Properties are applied via parameter maps (e.g., SET n += item.properties), so property keys and values are never interpolated into the Cypher string.
VectorStore handles everything related to vector and fulltext operations.

Index Creation

Vector indexes use FalkorDB’s native vector index syntax:
Fulltext indexes use the RediSearch-based fulltext API:
Relationship vector indexes use the edge index syntax:
All index creation is idempotent — if the index already exists, the error is silently caught and logged at debug level.

Chunk Indexing (index_chunks)

When chunks are ingested, their text is embedded and stored:
  1. Batch embed: All chunk texts are passed to the embedder via a single aembed_documents call. The underlying provider controls how these are internally batched (e.g., via a configurable batch_size) and may split them across multiple API requests.
  2. Batch write: Vectors are written to Chunk nodes using UNWIND (500 per batch):
  3. Fallback: If batch embedding fails, chunks are embedded one at a time. If batch writing fails, items are written individually.

Entity Embedding Backfill (backfill_entity_embeddings)

After all documents are ingested, entity nodes need embeddings for vector search. This is done during finalize():
  1. Query: Find entities missing embeddings: WHERE e.embedding IS NULL
  2. Embed: Batch-embed entity names via aembed_documents
  3. Write: Store vectors using UNWIND:
  4. Loop: Repeat until no more entities with NULL embeddings remain. Each batch naturally returns the next un-embedded set.

Relationship Embedding (embed_relationships)

RELATES edges with a fact property but no embedding are batch-embedded:
  1. Query: Find edges with r.embedding IS NULL AND r.fact IS NOT NULL
  2. Embed: Batch-embed the fact text
  3. Write: Store vectors on each edge individually (using internal FalkorDB edge IDs):

Search Methods

Vector search on Chunk nodes:
Vector search on Entity nodes:
Vector search on RELATES edges (with fallback):
Fulltext search:
Special characters in fulltext queries are escaped for RediSearch compatibility (commas, brackets, operators, etc.).

Stored Embedding Optimization

During retrieval, the rerank_chunks() function uses stored embeddings when possible. Instead of re-embedding all candidate chunks (which would require an expensive API call), it fetches the vectors already stored on Chunk nodes and computes cosine similarity locally. This makes reranking instant when stored embedding coverage is >= 90%.

ensure_indices()

Creates all 5 standard indexes in one call. Tracks state internally (_indices_ensured) to avoid redundant creation. Called automatically after each ingest() call, and re-run during finalize() (which resets the flag).

EntityDeduplicator — Merging Duplicate Entities

After ingesting multiple documents, the same real-world entity might exist as multiple nodes (e.g., “Alice” from doc 1 and “Alice” from doc 5). The deduplicator merges them.

Phase 1: Exact Name Match (Always Runs)

  1. Fetch all __Entity__ nodes with their primary label (the non-__Entity__ label)
  2. Group by (canonical_name, label) — the name is case-folded with accents, punctuation, inner dots (A.I. = AI) and a leading English article removed, Surname, Given un-inverted, common abbreviations expanded and one trailing legal form (Ltd, Inc) dropped; word order and the letters of non-Latin scripts are kept. Grouping by label prevents cross-type merging (Person “Paris” and Location “Paris” stay separate)
  3. Fold acronyms: a 3–6 letter name joins the group of the single same-label entity whose initials it spells (AIHS / Ashford Island Historical Society). An acronym with several candidate expansions is left alone
  4. For each group with duplicates:
    • Survivor: A node a declared table wrote (it carries the row’s key, so the table’s next re-sync still finds it), then a real node over a placeholder, then the long form over an acronym, then the best-connected node (most RELATES and MENTIONED_IN edges), then the longest description, then the longest name. Two nodes written from declared keys are never merged into each other, whatever they are called
    • Remap edges: All RELATES and MENTIONED_IN edges from duplicates are redirected to the survivor. The survivor is MATCHed before every MERGE — a path-MERGE with an unbound (s {id}) creates an empty ghost __Entity__ node whenever the edge does not exist yet — the RELATES MERGE is keyed on rel_type so two facts between one pair stay two edges, and RELATES source_chunk_ids are unioned rather than overwritten:
    • Absorb and delete in one statement: the duplicate’s source_chunk_ids are unioned into the survivor’s, its description appended to the survivor’s with " | " (each segment kept once), its name added to the survivor’s aliases when it is not a spelling variant, and every other property the survivor lacks — a typed column a table signed, is_stub — copied over, with a value already on the survivor always winning; then DETACH DELETE dup. Because the survivor is matched in the same statement, a survivor removed mid-pass makes this a no-op — the duplicate and its edges stay in place and the merge is not counted.

Phase 2: Fuzzy Embedding Match (Optional)

Enabled via fuzzy=True. Catches near-duplicates that have slightly different names (e.g., “J. Doe” and “Jane Doe”):
  1. Fetch all surviving entities with their labels
  2. Batch-embed entity names
  3. Normalize vectors and compute pairwise cosine similarity in blocks (1000 entities per block to avoid OOM)
  4. Merge pairs above the similarity threshold (default: 0.95 — at 0.90 the pass wrongly merged 2 of 33 hard-negative pairs for no added recall), but only within the same label (no cross-type merging)
  5. Remap and delete as in Phase 1

Phase 3: Resolver Across Sources (Default in finalize())

The resolution strategy finalize(resolver=...) names — by default an LLMVerifiedResolution over the instance’s llm and embedder, asking the model about pairs from cosine 0.6 — is shown every surviving entity of the graph as one GraphData, with the RELATES edges between them, and asked which are one thing. It decides identity; the merge follows Phase 1’s rules (survivor rank, absorb-and-delete in one statement), so a table’s row always survives a mention of it and two keyed rows never fold into each other. A NO is written to the graph as a DISTINCT_FROM edge and the pair is never asked again; finalize(resolve=False) skips this phase.

Phase 4: LLM-Judged Cross-Document Dedup (Default in finalize())

Ingest resolves per document, so Airbus in one file and Airbus SE in another never meet until here. LLMJudgeDeduplicator (storage/judge_dedup.py) runs once over the whole graph, after Phases 1–3:
  1. Embed every entity’s name (written to e.embedding and kept for retrieval; finalize() runs backfill_entity_embeddings() afterwards for whatever is still missing one) and its description (written to e.description_embedding with e.description_embedding_hash, the digest of the text it was computed from, the embedder’s model_name and the vector’s dimension — a later run reuses the vector while the hash still matches, and re-embeds what a merge or a re-ingest changed or another model embedded; a cached vector of another dimension than the ones embedded now is re-embedded with a warning, not dropped); an entity with no description is nominated by name only
  2. Nominate candidate pairs through three LLM-free doors: top-10 name neighbours at cosine ≥ 0.65, top-10 description neighbours at ≥ 0.55, and “A’s name appears inside B’s description”
  3. Group candidates into dense sets of at most 8: more than half of a set’s pairs must be nominated edges, so a chain A–B–C–D where only neighbours are similar (exactly half) cannot form; an edge left between two sets is its own 2-member set. Whole sets are packed into prompts of ≤ 3,000 tokens; the entity text is rendered one line per entity, flattened and quoted, and the prompt tells the model it is untrusted document data
  4. Judge, pass 1 — the LLM partitions each set into same-referent groups; names, labels and descriptions only, no similarity scores shown
  5. Judge, pass 2 (judge_vote=True, the default) — the same sets, members shuffled and reversed; an answer that changes with the order is a guess. With judge_vote=False there is no second pass: the first pass’s groups are the agreed pairs and nothing is linked for disagreement (a cross-set pair, step 7, is still linked, with agreement=1)
  6. Both passes agree → merge (with judge_vote=False: the single pass groups them), through the same absorb path as Phase 1: survivor by the Phase 1 rank (a table’s row over a mention, then degree, then longest description); every member’s description kept in descriptions and joined with " | " in description; other names kept in aliases; source_chunk_ids and missing properties carried; edges remapped to the survivor first, then the loser deleted. On top of that the survivor gains every member’s label, as Cypher labels and on record in merged_labels, so a later merge carries them on. Agreements are unioned within one set only: two sets that share a member (a straddling edge) never chain into one merge, the cross-set pair is linked instead. Two keyed rows, a mention two rows could own, and any pair a resolver remembered as DISTINCT_FROM are never put to the judge, never grouped with each other through a third member, and never merged or linked. If the DISTINCT_FROM pairs cannot be read at all, the phase does not run — an unavailable protection set is not an empty one — and last_judge_stats reports skipped_reason; the resolver and fuzzy phases follow the same rule
  7. Only one pass agrees → link: a SAME_AS edge (source='llm_judge', agreement=1; for a pair agreed on across two sets, agreement is the number of passes that saw it — 2 with the vote, 1 without, so agreement=2 always means both passes agreed) between the two; both nodes and both descriptions stay. A set whose prompt failed in either pass is left unjudged — its pairs are neither merged nor linked
  8. Neither → nothing
Measured on the benchmark corpus: the two-pass vote cut wrong merges by two thirds at 2× judge cost; gpt-4.1 made 9 wrong merges where gpt-4o-mini made 51, so pass a gpt-4.1-class judge_llm. Statistics (candidates, sets, calls, agreed/disagreed pairs, merged, linked) are in EntityDeduplicator.last_judge_stats and FinalizeResult.judge_stats. finalize(judge=False) skips this phase (the resolver pass still runs; finalize(resolve=False, judge=False) calls no model). The standalone deduplicate_entities() runs the judge only when asked (judge=True).

FalkorDB-Specific Notes

Vector Storage

FalkorDB stores vectors as vecf32 — a native 32-bit float vector type. All vectors in the SDK are stored via:

Vector Search API

Node vector search uses 4 arguments:
Relationship vector search (FalkorDB >= 4.2):

Graph Deletion

For fast graph deletion, use GRAPH.DELETE (the Redis-level command):
This is much faster than MATCH (n) DETACH DELETE n on large graphs.

Retry Behavior

The connection layer retries transient query failures up to 3 times using exponential backoff with jitter (retry_delay * 2^attempt * random(0.5, 1.5)) and employs a circuit breaker to short-circuit repeated failures. Non-transient errors (containing “already indexed”, “already exists”, or “unknown index”) are raised immediately.

File Reference