The Big Picture
GraphStore — Writing Nodes and Relationships
GraphStore is the single write path for all graph data. It uses parameterized Cypher to prevent injection and batched UNWIND for performance.
Batched UNWIND MERGE
Nodes and relationships are written in batches of 500 items per query. The pattern:Label Hints
Relationship MATCH queries use label hints to speed up endpoint lookups. Instead of searching all nodes, FalkorDB only looks at nodes with the specified label:
Unknown edge types default to
(__Entity__, __Entity__).
Per-Item Fallback
If a batch upsert fails (e.g., a single malformed property causes the whole batch to error), GraphStore falls back to per-item upserts. The behavior differs slightly: for nodes, the first per-item failure raises aDatabaseError (remaining items in that batch are not attempted); for relationships, failures are logged as warnings and processing continues through the batch.
None-ID Guard
Before writing, nodes withNone or empty IDs are filtered out. These come from bad LLM extraction (the LLM sometimes returns entities without proper names). The guard prevents phantom nodes from polluting the graph.
Property Cleaning
All properties go through_clean_properties() before writing:
Cypher Safety
Node labels and relationship types are sanitized viasanitize_cypher_label() to prevent Cypher injection. Properties are applied via parameter maps (e.g., SET n += item.properties), so property keys and values are never interpolated into the Cypher string.
VectorStore — Embeddings, Indexes, and Search
VectorStore handles everything related to vector and fulltext operations.
Index Creation
Vector indexes use FalkorDB’s native vector index syntax:Chunk Indexing (index_chunks)
When chunks are ingested, their text is embedded and stored:- Batch embed: All chunk texts are passed to the embedder via a single
aembed_documentscall. The underlying provider controls how these are internally batched (e.g., via a configurablebatch_size) and may split them across multiple API requests. - Batch write: Vectors are written to Chunk nodes using UNWIND (500 per batch):
- Fallback: If batch embedding fails, chunks are embedded one at a time. If batch writing fails, items are written individually.
Entity Embedding Backfill (backfill_entity_embeddings)
After all documents are ingested, entity nodes need embeddings for vector search. This is done duringfinalize():
- Query: Find entities missing embeddings:
WHERE e.embedding IS NULL - Embed: Batch-embed entity names via
aembed_documents - Write: Store vectors using UNWIND:
- Loop: Repeat until no more entities with NULL embeddings remain. Each batch naturally returns the next un-embedded set.
Relationship Embedding (embed_relationships)
RELATES edges with afact property but no embedding are batch-embedded:
- Query: Find edges with
r.embedding IS NULL AND r.fact IS NOT NULL - Embed: Batch-embed the
facttext - Write: Store vectors on each edge individually (using internal FalkorDB edge IDs):
Search Methods
Vector search on Chunk nodes:Stored Embedding Optimization
During retrieval, thererank_chunks() function uses stored embeddings when possible. Instead of re-embedding all candidate chunks (which would require an expensive API call), it fetches the vectors already stored on Chunk nodes and computes cosine similarity locally. This makes reranking instant when stored embedding coverage is >= 90%.
ensure_indices()
Creates all 5 standard indexes in one call. Tracks state internally (_indices_ensured) to avoid redundant creation. Called automatically after each ingest() call, and re-run during finalize() (which resets the flag).
EntityDeduplicator — Merging Duplicate Entities
After ingesting multiple documents, the same real-world entity might exist as multiple nodes (e.g., “Alice” from doc 1 and “Alice” from doc 5). The deduplicator merges them.Phase 1: Exact Name Match (Always Runs)
- Fetch all
__Entity__nodes with their primary label (the non-__Entity__label) - Group by
(canonical_name, label)— the name is case-folded with accents, punctuation, inner dots (A.I.=AI) and a leading English article removed,Surname, Givenun-inverted, common abbreviations expanded and one trailing legal form (Ltd,Inc) dropped; word order and the letters of non-Latin scripts are kept. Grouping by label prevents cross-type merging (Person “Paris” and Location “Paris” stay separate) - Fold acronyms: a 3–6 letter name joins the group of the single same-label entity whose initials it spells (
AIHS/Ashford Island Historical Society). An acronym with several candidate expansions is left alone - For each group with duplicates:
- Survivor: A node a declared table wrote (it carries the row’s key, so the table’s next re-sync still finds it), then a real node over a placeholder, then the long form over an acronym, then the best-connected node (most RELATES and MENTIONED_IN edges), then the longest description, then the longest name. Two nodes written from declared keys are never merged into each other, whatever they are called
- Remap edges: All RELATES and MENTIONED_IN edges from duplicates are redirected to the survivor. The survivor is
MATCHed before everyMERGE— a path-MERGEwith an unbound(s {id})creates an empty ghost__Entity__node whenever the edge does not exist yet — the RELATESMERGEis keyed onrel_typeso two facts between one pair stay two edges, and RELATESsource_chunk_idsare unioned rather than overwritten: - Absorb and delete in one statement: the duplicate’s
source_chunk_idsare unioned into the survivor’s, its description appended to the survivor’s with" | "(each segment kept once), its name added to the survivor’saliaseswhen it is not a spelling variant, and every other property the survivor lacks — a typed column a table signed,is_stub— copied over, with a value already on the survivor always winning; thenDETACH DELETE dup. Because the survivor is matched in the same statement, a survivor removed mid-pass makes this a no-op — the duplicate and its edges stay in place and the merge is not counted.
Phase 2: Fuzzy Embedding Match (Optional)
Enabled viafuzzy=True. Catches near-duplicates that have slightly different names (e.g., “J. Doe” and “Jane Doe”):
- Fetch all surviving entities with their labels
- Batch-embed entity names
- Normalize vectors and compute pairwise cosine similarity in blocks (1000 entities per block to avoid OOM)
- Merge pairs above the similarity threshold (default: 0.95 — at 0.90 the pass wrongly merged 2 of 33 hard-negative pairs for no added recall), but only within the same label (no cross-type merging)
- Remap and delete as in Phase 1
Phase 3: Resolver Across Sources (Default in finalize())
The resolution strategy finalize(resolver=...) names — by default an LLMVerifiedResolution over the instance’s llm and embedder, asking the model about pairs from cosine 0.6 — is shown every surviving entity of the graph as one GraphData, with the RELATES edges between them, and asked which are one thing. It decides identity; the merge follows Phase 1’s rules (survivor rank, absorb-and-delete in one statement), so a table’s row always survives a mention of it and two keyed rows never fold into each other. A NO is written to the graph as a DISTINCT_FROM edge and the pair is never asked again; finalize(resolve=False) skips this phase.
Phase 4: LLM-Judged Cross-Document Dedup (Default in finalize())
Ingest resolves per document, so Airbus in one file and Airbus SE in another never meet until here. LLMJudgeDeduplicator (storage/judge_dedup.py) runs once over the whole graph, after Phases 1–3:
- Embed every entity’s name (written to
e.embeddingand kept for retrieval;finalize()runsbackfill_entity_embeddings()afterwards for whatever is still missing one) and its description (written toe.description_embeddingwithe.description_embedding_hash, the digest of the text it was computed from, the embedder’smodel_nameand the vector’s dimension — a later run reuses the vector while the hash still matches, and re-embeds what a merge or a re-ingest changed or another model embedded; a cached vector of another dimension than the ones embedded now is re-embedded with a warning, not dropped); an entity with no description is nominated by name only - Nominate candidate pairs through three LLM-free doors: top-10 name neighbours at cosine ≥ 0.65, top-10 description neighbours at ≥ 0.55, and “A’s name appears inside B’s description”
- Group candidates into dense sets of at most 8: more than half of a set’s pairs must be nominated edges, so a chain A–B–C–D where only neighbours are similar (exactly half) cannot form; an edge left between two sets is its own 2-member set. Whole sets are packed into prompts of ≤ 3,000 tokens; the entity text is rendered one line per entity, flattened and quoted, and the prompt tells the model it is untrusted document data
- Judge, pass 1 — the LLM partitions each set into same-referent groups; names, labels and descriptions only, no similarity scores shown
- Judge, pass 2 (
judge_vote=True, the default) — the same sets, members shuffled and reversed; an answer that changes with the order is a guess. Withjudge_vote=Falsethere is no second pass: the first pass’s groups are the agreed pairs and nothing is linked for disagreement (a cross-set pair, step 7, is still linked, withagreement=1) - Both passes agree → merge (with
judge_vote=False: the single pass groups them), through the same absorb path as Phase 1: survivor by the Phase 1 rank (a table’s row over a mention, then degree, then longest description); every member’s description kept indescriptionsand joined with" | "indescription; other names kept inaliases;source_chunk_idsand missing properties carried; edges remapped to the survivor first, then the loser deleted. On top of that the survivor gains every member’s label, as Cypher labels and on record inmerged_labels, so a later merge carries them on. Agreements are unioned within one set only: two sets that share a member (a straddling edge) never chain into one merge, the cross-set pair is linked instead. Two keyed rows, a mention two rows could own, and any pair a resolver remembered asDISTINCT_FROMare never put to the judge, never grouped with each other through a third member, and never merged or linked. If theDISTINCT_FROMpairs cannot be read at all, the phase does not run — an unavailable protection set is not an empty one — andlast_judge_statsreportsskipped_reason; the resolver and fuzzy phases follow the same rule - Only one pass agrees → link: a
SAME_ASedge (source='llm_judge',agreement=1; for a pair agreed on across two sets,agreementis the number of passes that saw it —2with the vote,1without, soagreement=2always means both passes agreed) between the two; both nodes and both descriptions stay. A set whose prompt failed in either pass is left unjudged — its pairs are neither merged nor linked - Neither → nothing
judge_llm. Statistics (candidates, sets, calls, agreed/disagreed pairs, merged, linked) are in EntityDeduplicator.last_judge_stats and FinalizeResult.judge_stats. finalize(judge=False) skips this phase (the resolver pass still runs; finalize(resolve=False, judge=False) calls no model). The standalone deduplicate_entities() runs the judge only when asked (judge=True).
FalkorDB-Specific Notes
Vector Storage
FalkorDB stores vectors asvecf32 — a native 32-bit float vector type. All vectors in the SDK are stored via:
Vector Search API
Node vector search uses 4 arguments:Graph Deletion
For fast graph deletion, useGRAPH.DELETE (the Redis-level command):
MATCH (n) DETACH DELETE n on large graphs.
Retry Behavior
The connection layer retries transient query failures up to 3 times using exponential backoff with jitter (retry_delay * 2^attempt * random(0.5, 1.5)) and employs a circuit breaker to short-circuit repeated failures. Non-transient errors (containing “already indexed”, “already exists”, or “unknown index”) are raised immediately.