Graph-Based Spreading Activation
Overview
Graph-Based Spreading Activation extends the hybrid memory system with typed relationships between facts, enabling zero-LLM recall through graph traversal. This feature addresses a key limitation of pure vector search: finding conceptually or causally related items that may not be semantically similar.
The Problem
Traditional vector search finds semantically similar text but often misses:
- Causal relationships: “Why did X fail?” → “Decision Y caused bug Z”
- Hierarchical relationships: “What are the components of system X?”
- Temporal relationships: “What superseded this decision?”
- Dependency relationships: “What does this feature depend on?”
Pure keyword (FTS5) search also fails to capture these structural relationships.
The Solution
The system now maintains a graph of typed relationships between facts in the memory_links table. During recall, the system:
- Vector/FTS search finds initial relevant facts (starting nodes)
- Graph traversal (BFS) follows typed links to discover connected facts
- Results merging combines both approaches
This is zero-LLM — no embedding calls are needed for graph traversal, making it extremely fast.
See it live. The Memory Graph app renders this graph as an interactive constellation at
http://127.0.0.1:7700/graph— strongest memories in the center, real-time pulses as memories are recalled, community clustering, and in-place curation (add/edit/remove/link, plus gap & contradiction views).
Architecture
Database Schema
memory_links Table
CREATE TABLE memory_links (
id TEXT PRIMARY KEY,
source_fact_id TEXT NOT NULL,
target_fact_id TEXT NOT NULL,
link_type TEXT NOT NULL,
strength REAL NOT NULL DEFAULT 1.0,
created_at INTEGER NOT NULL,
FOREIGN KEY (source_fact_id) REFERENCES facts(id) ON DELETE CASCADE,
FOREIGN KEY (target_fact_id) REFERENCES facts(id) ON DELETE CASCADE
);
-- Indexes for efficient graph traversal
CREATE INDEX idx_links_source ON memory_links(source_fact_id);
CREATE INDEX idx_links_target ON memory_links(target_fact_id);
CREATE INDEX idx_links_type ON memory_links(link_type);
CREATE INDEX idx_links_source_type ON memory_links(source_fact_id, link_type);
Link Types
The system supports five typed relationships (inspired by Zep/Graphiti):
| Type | Description | Example |
|---|---|---|
SUPERSEDES | One fact replaces/updates another | “New API key supersedes old API key” |
CAUSED_BY | Causal relationship | “Bug X was caused by decision Y” |
PART_OF | Hierarchical/component relationship | “Feature X is part of project Y” |
RELATED_TO | General semantic relationship | “Dark mode preference relates to VS Code usage” |
DEPENDS_ON | Dependency relationship | “Feature X depends on library Y” |
Link Strength
Each link has a strength value (0.0-1.0):
1.0= Strong relationship (manually created or high confidence)0.7-0.9= Moderate relationship (auto-linked with high similarity)0.5-0.6= Weak relationship (auto-linked with moderate similarity)
Configuration
Add the graph section to your plugin config:
{
"plugins": {
"entries": {
"openclaw-hybrid-memory": {
"config": {
"graph": {
"enabled": true,
"autoLink": true,
"autoLinkStrength": 0.7,
"autoLinkSimilarityThreshold": 0.7,
"autoLinkLimit": 2,
"maxTraversalDepth": 2,
"useInRecall": false,
"coOccurrenceWeight": 0.2
},
"graphRetrieval": {
"enabled": true,
"defaultExpand": true,
"maxExpandDepth": 2,
"maxExpandedResults": 12
},
"retrieval": {
"strategies": ["semantic", "fts5", "graph"],
"graphWalkDepth": 2
}
}
}
}
}
}
Configuration Options
| Option | Type | Default | Description |
|---|---|---|---|
enabled | boolean | true | Enable graph features |
autoLink | boolean | false | Auto-create RELATED_TO links when storing facts |
autoLinkStrength | number | 0.7 | Strength (0.0–1.0) stored on auto-created RELATED_TO links |
autoLinkSimilarityThreshold | number | 0.7 | Minimum embedding cosine similarity for semantic auto-linking |
autoLinkMinScore | number | — | Deprecated. Alias for both strength and threshold when the new keys are unset |
autoLinkLimit | number | 3 | Max similar facts to auto-link per storage |
maxTraversalDepth | number | 2 | Max hops for legacy graph traversal in recall |
useInRecall | boolean | true | Legacy flat graph add-on in memory_recall when GraphRAG expand is off |
coOccurrenceWeight | number | 0.3 | Link strength for temporal co-occurrence RELATED_TO edges (facts recalled together) |
autoSupersede | boolean | true | Auto-create a SUPERSEDES edge and supersede the old fact when an entity+key conflict is detected on store |
Person and organization enrichment (entity layer)
When graph.enabled is true, each successful memory_store also runs (asynchronously) multilingual NER: franc detects the primary text language (ISO 639-3); a small LLM extracts PERSON and ORG spans with offsets. Before results are stored in SQLite (fact_entity_mentions, organizations, contacts, org_fact_links), mention rows are normalized, deduplicated per fact/label, and filtered for short or generic noise. This complements known-entity autoLinkEntities behavior—it does not remove or replace structured entity fields on facts.
Agents use memory_directory (list_contacts, org_view) for stable contact lists and org-scoped fact ids—not for ad-hoc ranked search (use memory_recall for that). Facts stored before this feature, or imports without mentions, can be backfilled with openclaw hybrid-mem enrich-entities. See MULTILINGUAL-SUPPORT.md and issue #985–#987.
Structured contact profiles (#2014)
NER populates contacts.display_name and weak primary_org_id heuristics, but operators also keep email, role, and board membership in facts and curated CONTACTS.md rosters. When contacts.profileEnrichment is enabled (default), each successful memory_store / auto-capture also:
- Parses email, phone, and role patterns from the fact text (reuses
extractStructuredFieldsplus person-aware heuristics). - Merges into the matched
contactsrow (fill empty fields; overwrite only when the incoming source outranks the existingprofile_updated_by, or on import/manual supersede). - Resolves partial NER names (
Daniel→ existingDaniel Thunberg) when there is a single unambiguous prefix match.
CLI
openclaw hybrid-mem contacts import --from workspace/CONTACTS.md [--dry-run]
openclaw hybrid-mem contacts sync # re-import when CONTACTS.md mtime changes
openclaw hybrid-mem contacts merge <fromId> <intoId> # each arg accepts a contact id or a name
Config (contacts block): profileEnrichment (default true), requireSurname (skip single-token PERSON rows from NER unless co-mentioned org sets context), importPath (default for contacts sync). See CONFIGURATION.md.
memory_directory: list_contacts includes email, role, board status, and updated_at; org_view groups people under Board / Management when board_status is set.
Recommendation: Start with autoLink: false and manually create links using memory_link to establish high-quality relationships. Enable autoLink: true later once you have a critical mass of facts.
Contact profile enrichment vs. NER (issue #2014)
NER (above) creates thin contacts/organizations rows from PERSON/ORG spans — just display_name and, when the same fact also mentions an org, primary_org_id. Contact profile enrichment is a second, focused step that fills in the rest of the contact record from the same fact text:
- After NER runs, if the fact mentions exactly one PERSON (ambiguous when there are two or more — attributing an email to the wrong person is worse than not enriching), the fact text is scanned for an email address, a phone number, and a job-title/role keyword (English + Swedish, e.g.
CEO,Sverigechef,styrelseledamot). - Matched fields are merged onto that contact’s
email,phone,role, andboard_statuscolumns — filling empty fields, and only overwriting already-populated fields when the new write has equal or higher priority (manual>agent/import>ner).notesare appended (deduplicated), never overwritten. - Example: storing
"Daniel Thunberg email daniel.thunberg@avoki.com Sverigechef Avoki"setscontacts.email = "daniel.thunberg@avoki.com",contacts.role = "Sverigechef Avoki", andcontacts.board_status = "management"on the “Daniel Thunberg” contact. - This is a lightweight regex heuristic, not an LLM/NLP title parser — it will miss unusual phrasing. Non-goal for v1: full CRM sync.
Entity resolution (duplicate contacts). NER often creates near-duplicate contacts from partial names (e.g. “Daniel” and “Daniel Thunberg” as two rows). Two mechanisms address this without manual SQL:
- Auto-merge: when NER is about to create a new contact and there is exactly one existing contact whose name is a token prefix/suffix of the new name (e.g. “Daniel” ⊂ “Daniel Thunberg”), the mention is merged into that existing contact instead of creating a duplicate — the shorter/older name is kept as an alias. Ambiguous cases (zero or 2+ candidates) are left alone.
- Manual merge:
openclaw hybrid-mem contacts suggest-mergeslists remaining candidates;openclaw hybrid-mem contacts merge <fromId> <intoId>(each arg accepts a contact id or a name) repoints allfact_entity_mentionsandfacts.entity_contact_idtointoId, folds infromId’s profile fields (manual merges always win field conflicts) and aliases, and deletesfromId. - Set
contacts.requireSurname: trueto stop NER from creating a new PERSON contact from a single-token surface (e.g. bare “Daniel”) unless the same fact also mentions an organization — reduces one-word noise contacts at the cost of occasionally not tracking a first-name-only mention. Defaultfalse(matches pre-#2014 behavior).
Roster import: openclaw hybrid-mem contacts import --from <file> [--dry-run] [--no-part-of] upserts organizations/contacts/roster facts from a CONTACTS.md-style file. It also stores a per-org summary roster fact and links each person’s roster fact to it with a PART_OF edge (use --no-part-of to skip). See CONFIGURATION.md for the file format and the full contacts CLI reference.
Usage
1. Auto-Linking (Optional)
When graph.autoLink is enabled, memory_store automatically creates RELATED_TO links to the most similar existing facts:
// Store a new fact
memory_store({
text: "Project X uses TypeScript for backend",
category: "fact"
});
// If autoLink is enabled, the system finds the top 3 most similar facts
// (e.g., "TypeScript requires strict mode") and creates RELATED_TO links
Output:
Stored: "Project X uses TypeScript for backend" [decay: stable] (linked to 2 related facts)
2. Manual Link Creation
Use memory_link to create typed relationships:
memory_link({
sourceFact: "fact-id-1", // "Decision to migrate to TypeScript"
targetFact: "fact-id-2", // "Project X uses TypeScript"
linkType: "CAUSED_BY",
strength: 0.9
});
Output:
Created CAUSED_BY link from "Decision to migrate to TypeScript..." to "Project X uses TypeScript..." (strength: 0.9)
3. Enhanced Recall with Graph Traversal
When graph.useInRecall is enabled, memory_recall automatically traverses the graph:
memory_recall({
query: "TypeScript configuration"
});
How it works:
- Vector/FTS search finds initial matches (e.g., “TypeScript configuration requires strict mode”)
- Graph traversal finds connected facts up to
maxTraversalDepthhops (e.g., “Decision to migrate to TypeScript”, “Project X uses TypeScript”) - Results are merged and returned
Output:
Found 5 memories (includes 2 graph-connected facts):
1. [sqlite/technical] TypeScript configuration requires strict mode (95%)
2. [sqlite/fact] Project X uses TypeScript for backend (88%)
3. [sqlite/decision] Decision to migrate to TypeScript was made in Q1 2024 (50%)
4. [sqlite/fact] TypeScript compiler options include strict: true (50%)
5. [sqlite/preference] Team prefers TypeScript over JavaScript (45%)
Memories 3-5 were discovered via graph traversal, not by text similarity.
4. Graph Exploration
Use memory_graph to explore the relationship graph:
memory_graph({
factId: "fact-id-xyz",
depth: 2 // max 3
});
Output:
Fact: "Project X uses TypeScript for backend"
Direct links (3):
→ [PART_OF] TypeScript configuration requires strict mode (strength: 1.00)
← [CAUSED_BY] Decision to migrate to TypeScript was made in Q1 2024 (strength: 0.80)
← [DEPENDS_ON] TypeScript configuration requires strict mode (strength: 0.70)
Total connected facts (depth 2): 5
5. Backfilling orphan links (#2127)
autoLink only creates links at store time. A store that enabled autoLink late, imported facts in bulk, or simply has a lot of pre-existing facts can end up with a large fraction of active facts having no explicit graph link at all — memory_health reports this as orphanRate and warns above 70%.
openclaw hybrid-mem maintenance graph-link-enrichment backfills one safe, deterministic signal for those orphans: facts that share a recorded source event (facts.provenance_json.sourceEventIds) are linked via RELATED_TO. It never invents a similarity judgment and never touches a fact that already has any link (of any type):
# Dry-run (default) — report what would be linked
openclaw hybrid-mem maintenance graph-link-enrichment
# Actually create the links, bounded per run
openclaw hybrid-mem maintenance graph-link-enrichment --apply --limit 500
Run it repeatedly (e.g. as a periodic maintenance job) to work through the backlog — each run scans the oldest still-orphaned facts first, so it makes steady forward progress. This intentionally does not use embedding similarity; pair it with autoLink/memory_link for broader coverage.
Performance
Zero-LLM Traversal
Graph traversal uses SQLite joins and indexes — no embedding calls. Typical performance:
- 1-hop traversal: ~1-5ms
- 2-hop traversal: ~5-20ms
- 3-hop traversal: ~20-50ms
Compare this to vector search (50-200ms per embedding call + search).
Scaling
- 10,000 facts, 20,000 links: Traversal remains under 50ms for depth=2
- 100,000 facts, 200,000 links: Consider reducing
maxTraversalDepthto 1 - Foreign key cascades: When a fact is deleted, all its links are automatically removed
Use Cases
1. Causal Reasoning
Question: “Why did the Nibe heat pump integration fail?”
Traditional vector search: Finds “Nibe integration error” but misses the root cause.
Graph-enhanced recall:
- Vector search finds “Nibe integration error in Home Assistant”
- Traverse
CAUSED_BYlinks → “Zigbee coordinator firmware outdated” - Traverse
DEPENDS_ONlinks → “Nibe S1155 requires firmware 2.0.4+”
Result: Complete causal chain without asking the LLM “why?”.
2. Project Documentation
Question: “What components does Project X have?”
Setup:
memory_link({
sourceFact: "project-x-id",
targetFact: "frontend-id",
linkType: "PART_OF"
});
memory_link({
sourceFact: "project-x-id",
targetFact: "backend-id",
linkType: "PART_OF"
});
Recall:
memory_recall({ query: "Project X" });
// Returns: Project X + all PART_OF components
3. Decision History
Question: “What superseded the old API key decision?”
Setup:
memory_link({
sourceFact: "new-api-key-id",
targetFact: "old-api-key-id",
linkType: "SUPERSEDES"
});
Recall:
memory_recall({ query: "API key" });
// Returns: Both API keys, with SUPERSEDES link visible in memory_graph
4. Auto-Discovery of Related Topics
Scenario: User asks about “authentication” but the system has related facts about “OAuth”, “JWT”, “session management”.
Setup: Enable autoLink with autoLinkMinScore: 0.7.
Result: When storing “OAuth configuration for API”, the system auto-links to “JWT token expiry” and “session timeout settings”. Later recall of “authentication” finds all related facts via graph traversal.
Best Practices
1. Start Manual, Then Auto-Link
- Phase 1: Disable
autoLink, manually create high-quality links usingmemory_link - Phase 2: Once you have 50+ facts, enable
autoLink: trueto discover missed relationships - Phase 3: Review auto-links with
memory_graph, delete weak ones, strengthen strong ones
2. Use Specific Link Types
Don’t default to RELATED_TO for everything:
- Use
CAUSED_BYfor causal chains (debugging, decisions) - Use
PART_OFfor hierarchical structures (projects, systems) - Use
SUPERSEDESfor versioned facts (credentials, configs) - Use
DEPENDS_ONfor technical dependencies (libraries, features) - Use
RELATED_TOonly when no stronger relationship exists
3. Prune Weak Links
After auto-linking runs for a while:
- Query for low-strength links:
SELECT * FROM memory_links WHERE strength < 0.6 - Review each one — delete if not useful
- This improves traversal quality and speed
4. Depth vs. Precision
- depth=1: Fast, precise, use for interactive queries
- depth=2: Balanced, good default for most cases
- depth=3: Slower, returns many facts, use only for exploratory queries
Comparison to Competitors
Zep / Graphiti
Similarities:
- Typed relationships (Node-Edge-Node triplets)
- Bi-temporal edges (creation time, validity time)
- Episode nodes for provenance
Differences:
- OpenClaw Hybrid: SQLite-based, local, zero-cost, privacy-first
- Graphiti: Neo4j/cloud-based, managed service, API-priced
Mem0
Similarities:
- Dual storage (vector + graph)
- Auto-extraction of entities and relationships
- Hybrid retrieval (vector narrows, graph enriches)
Differences:
- OpenClaw Hybrid: Manual + auto links, SQLite recursive CTEs
- Mem0: LLM-based extraction, Neo4j/Kuzu backends
MAGMA
Inspiration:
- Multi-graph concept: MAGMA uses four orthogonal graphs (semantic, temporal, causal, entity)
- Future enhancement: Consider adding
link_categoryfield tomemory_linksto support multiple graph types
Troubleshooting
Auto-Linking Not Working
Check:
graph.enabledandgraph.autoLinkare bothtruegraph.autoLinkMinScoreis not too high (try 0.6 instead of 0.8)- Vector embeddings are working (check
memory_storelogs)
Debug:
-- Check if links were created
SELECT COUNT(*) FROM memory_links WHERE link_type = 'RELATED_TO';
-- Check auto-link strengths
SELECT strength, COUNT(*) FROM memory_links
WHERE link_type = 'RELATED_TO'
GROUP BY ROUND(strength, 1);
/graph dashboard or /api/graph looks disconnected (#2126)
If memory_health reports thousands of links but the graph explorer looks sparse or edge-free, it’s likely a sampling artifact, not a broken graph: /api/graph defaults to mode=connected, which samples recent facts plus their direct link-neighbors. Pass ?mode=recent to see the old recency-only sample, or check the response’s coverage block (totalExplicitLinks, selectedEdges, orphanNodesInView) to distinguish “the sample is small” from “the graph really is that sparse” — see Backfilling orphan links if it’s the latter.
Graph Traversal Slow
Solutions:
- Reduce
maxTraversalDepthfrom 2 to 1 - Prune weak links (strength < 0.5)
- Vacuum database:
sqlite3 facts.db 'VACUUM' - Check index usage:
EXPLAIN QUERY PLAN SELECT ...
Too Many Results in Recall
Solutions:
- Lower
maxTraversalDepth - Disable
graph.useInRecallfor specific queries - Filter by
link_type(future enhancement)
Future Enhancements
1. Link Categories (MAGMA-style)
Add link_category to distinguish graph types:
semantic(RELATED_TO)causal(CAUSED_BY, DEPENDS_ON)temporal(SUPERSEDES)hierarchical(PART_OF)
2. Query-Specific Traversal
Allow memory_recall to accept linkTypes filter:
memory_recall({
query: "Nibe error",
linkTypes: ["CAUSED_BY", "DEPENDS_ON"] // Only follow causal links
});
3. Link Metadata
Add created_by (user vs. auto) and notes fields to memory_links for audit trails.
4. Bi-Temporal Edges (Zep-style)
Add valid_at / invalid_at columns to support superseded links without deletion.
References
- Graph memory: Graph-Based Spreading Activation (Zero-LLM Recall)
- Zep Graphiti: https://github.com/getzep/graphiti, arXiv:2501.13956
- Mem0: https://docs.mem0.ai/features/graph-memory
- MAGMA: Multi-Graph Architecture, arXiv:2601.03236
- SQLite Recursive CTEs: https://www.sqlite.org/lang_with.html
Quick Reference
Tools
| Tool | Purpose | Key Parameters |
|---|---|---|
memory_link | Create typed link | sourceFact, targetFact, linkType, strength |
memory_graph | Explore connections | factId, depth |
memory_directory | Contacts + org-centric views | operation: list_contacts | org_view; name_prefix, org_name, limit |
memory_recall | Hybrid search | query (graph traversal automatic if enabled) |
Link Types
SUPERSEDES,CAUSED_BY,PART_OF,RELATED_TO,DEPENDS_ON
Config Keys
graph.enabled,graph.autoLink,graph.autoLinkMinScore,graph.autoLinkLimit,graph.maxTraversalDepth,graph.useInRecallcontacts.requireSurname— see Contact profile enrichment vs. NER
Related docs
- README — Project overview and all docs
- MEMORY-GRAPH-APP.md — Interactive real-time visualization of this graph
- FEATURES.md — Categories, decay, and other fact features
- ARCHITECTURE.md — System architecture overview
- CLI-REFERENCE.md — All CLI commands
- CONFIGURATION.md — Graph config settings (
graph.enabled,graph.autoLink, etc.)