Most AI assistants suffer from context amnesia — they forget everything between sessions. Even with context windows expanding to 200K+ tokens, you eventually hit limits, and long conversations become unwieldy.
The core issues:
- Session isolation — Each new chat starts blank
- Context bloat — Long histories slow down inference
- No semantic retrieval — Even with context, finding relevant past info is slow
- No relationship tracking — "Who is X?" requires re-reading entire history
We use two complementary stores that solve different retrieval problems:
Purpose: Structural/relational memory. Answers "who", "what", "when" queries.
- Stores typed entities:
decision,preference,episode,config,entity - Tracks relationships via
links_toedges - Keyword + graph traversal queries
- Collection-based namespace separation
Example query:
curl -s -X POST http://localhost:8766/query -H "Content-Type: application/json" \
-d '{"query": "user preferences", "collection": "brain", "limit": 10}'Purpose: Semantic/vector search. Answers "vibe", "similar", "context" queries.
- Dense vector embeddings (1024-dim via nomic-embed-text)
- Semantic similarity search
- Collection namespacing for isolation
Example query:
curl -s -X POST http://localhost:8765/query -H "Content-Type: application/json" \
-d '{"query": "what did we discuss about crypto?", "collection": "agent_memory", "limit": 5}'| Question Type | Best Store | Why |
|---|---|---|
| "What's the user's timezone?" | GraphDB | Exact keyword match + entity relation |
| "Show recent decisions" | GraphDB | Typed filtering + sort by date |
| "What did we talk about last week?" | Qdrant | Semantic similarity finds related concepts |
| "Find similar conversations to X" | Qdrant | Vector similarity |
Key insight: GraphDB is fast for known-unknowns (you know what you're looking for). Qdrant is fast for unknown-unknowns (you know the vibe but not the keywords).
We extract memories via two parallel mechanisms:
Triggers on every message:sent event from brain agents:
User message → Agent response → Hook fires → qwen3.5:397b-cloud analysis → Insert to both stores
- Uses a smaller model for fast LLM extraction
- Captures: decisions, preferences, episodes, entities
- Content-hash deduplication (sha256 of first 500 chars)
Runs every 3 hours:
- Reads last extraction position from state file
- Scans session transcripts for new content
- Spawns subagent for deep analysis
- Inserts to both stores
- Updates state file
Why both? Real-time catches immediate memories; batch sweep catches anything the hook missed (edge cases, multi-turn context, etc.)
We dedupe at insertion time using content hashing:
content_hash = sha256(text[:500]).hexdigest()If a memory with the same hash exists, insertion returns duplicate_skipped. This prevents the same fact from bloating both stores while allowing re-extraction of modified/updated information.
| Store | Count | Status |
|---|---|---|
| GraphDB | ~380 memories | 🟢 |
| Qdrant | ~306 vectors | 🟢 |
This architecture is not perfect. Here are the known problems:
Memories accumulate forever. We need:
- TTL-based pruning for ephemeral info
- Importance scoring to keep high-value memories
- "Forget" mechanism for outdated facts
Impact: Database grows unbounded. Old/wrong preferences persist indefinitely.
If we learn "user prefers X" then later "user prefers Y":
- Both memories exist
- No automatic update/invalidation
- Requires manual dedup
Impact: Conflicting facts coexist. Query results may return outdated info.
qwen3.5:397b-cloudis fast but not perfect- May miss subtle preferences or over-extract noise
- No quality scoring on extracted memories
Impact: Garbage in, garbage out. No grounding validation.
- Memories are extracted per-message
- No mechanism to link "Episode A" → "Episode B" causally
- Timeline reconstruction requires post-processing
Impact: Can't answer "what led to this decision?" without manual reconstruction.
We had to lower indexing_threshold from 10,000 → 100 because vectors weren't being indexed at small scale.
Impact: Default config is tuned for large-scale production, not small personal deployments.
- Both stores run locally
- No multi-tenancy support
- Would need Qdrant multi-tenant or separate instances per user for production
Impact: Can't serve multiple users without duplicating infrastructure.
- Both stores return results by similarity/date
- No combined scoring (e.g., recency + relevance + importance)
- No learning from user feedback on retrieval quality
Impact: May surface irrelevant old memories over relevant new ones.
- Using
nomic-embed-textfor all vectors - Switching models requires full re-embedding
- No A/B testing of embedding quality
Impact: Stuck with initial model choice. Can't experiment with better embeddings.
- Raw text stored as-is
- Long conversations generate many memories
- No summarization or hierarchical memory
Impact: Storage grows linearly with conversation length.
- Real-time hook processes each message independently
- Context from multi-turn exchanges may be lost
- Batch sweep helps but isn't real-time
Impact: May miss context that spans multiple messages.
┌─────────────────────────────────────────────────────────────┐
│ MEMORY EXTRACTION │
├─────────────────────────────────────────────────────────────┤
│ │
│ User Message ──▶ Agent ──▶ Response ──▶ Hook Fires │
│ │ │
│ ▼ │
│ ┌─────────────────┐ │
│ │ qwen3.5:397b │ │
│ │ LLM Extraction │ │
│ └────────┬────────┘ │
│ │ │
│ ┌───────────────────┴───────────────┐│
│ ▼ ▼│
│ ┌─────────────────┐ ┌───────────┐│
│ │ GraphDB │ │ Qdrant ││
│ │ (port 8766) │ │(port 8765)││
│ │ │ │ ││
│ │ • Entity graphs │ │ • Vectors ││
│ │ • Relationships │ │ • Semantic││
│ │ • Keyword search │ │ search ││
│ └─────────────────┘ └───────────┘│
│ │
└─────────────────────────────────────────────────────────────┘
- "Remember that time we..." — Qdrant finds similar contexts
- "What's the user's timezone?" — GraphDB keyword match
- "What projects are active?" — GraphDB entity filtering
- "Summarize our discussions about X" — Qdrant semantic retrieval
This architecture gives the agent persistent, queryable memory across sessions without bloating context windows. The dual-store approach means we get both precise retrieval (GraphDB) and fuzzy semantic search (Qdrant).
But the gaps are real and numerous. This is a v1 architecture — functional but with clear evolution paths:
Near-term fixes:
- Memory expiration (TTL + importance scoring)
- Conflict resolution (prefer newer, invalidate outdated)
- Retrieval ranking (recency + relevance + importance)
Long-term evolution:
- Hierarchical memory (summarization layers)
- Multi-tenancy (Qdrant tenant mode)
- Embedding model swap system
- Grounding validation (fact-check extraction)
Last updated: 2026-03-27