Skip to content

Instantly share code, notes, and snippets.

@p3nj
Created March 27, 2026 03:25
Show Gist options
  • Select an option

  • Save p3nj/91733edfa165777a7981905634b90454 to your computer and use it in GitHub Desktop.

Select an option

Save p3nj/91733edfa165777a7981905634b90454 to your computer and use it in GitHub Desktop.

AI Agent Memory Architecture

The Problem

Most AI assistants suffer from context amnesia — they forget everything between sessions. Even with context windows expanding to 200K+ tokens, you eventually hit limits, and long conversations become unwieldy.

The core issues:

  1. Session isolation — Each new chat starts blank
  2. Context bloat — Long histories slow down inference
  3. No semantic retrieval — Even with context, finding relevant past info is slow
  4. No relationship tracking — "Who is X?" requires re-reading entire history

Our Solution: Dual-Store Memory

We use two complementary stores that solve different retrieval problems:

GraphDB (Neo4j Wrapper) — Port 8766

Purpose: Structural/relational memory. Answers "who", "what", "when" queries.

  • Stores typed entities: decision, preference, episode, config, entity
  • Tracks relationships via links_to edges
  • Keyword + graph traversal queries
  • Collection-based namespace separation

Example query:

curl -s -X POST http://localhost:8766/query -H "Content-Type: application/json" \
  -d '{"query": "user preferences", "collection": "brain", "limit": 10}'

Qdrant — Port 8765/6333

Purpose: Semantic/vector search. Answers "vibe", "similar", "context" queries.

  • Dense vector embeddings (1024-dim via nomic-embed-text)
  • Semantic similarity search
  • Collection namespacing for isolation

Example query:

curl -s -X POST http://localhost:8765/query -H "Content-Type: application/json" \
  -d '{"query": "what did we discuss about crypto?", "collection": "agent_memory", "limit": 5}'

Why Both?

Question Type Best Store Why
"What's the user's timezone?" GraphDB Exact keyword match + entity relation
"Show recent decisions" GraphDB Typed filtering + sort by date
"What did we talk about last week?" Qdrant Semantic similarity finds related concepts
"Find similar conversations to X" Qdrant Vector similarity

Key insight: GraphDB is fast for known-unknowns (you know what you're looking for). Qdrant is fast for unknown-unknowns (you know the vibe but not the keywords).

Memory Extraction Pipeline

We extract memories via two parallel mechanisms:

1. Real-Time Hook (memory-logger)

Triggers on every message:sent event from brain agents:

User message → Agent response → Hook fires → qwen3.5:397b-cloud analysis → Insert to both stores
  • Uses a smaller model for fast LLM extraction
  • Captures: decisions, preferences, episodes, entities
  • Content-hash deduplication (sha256 of first 500 chars)

2. Batch Sweep (Cron)

Runs every 3 hours:

  1. Reads last extraction position from state file
  2. Scans session transcripts for new content
  3. Spawns subagent for deep analysis
  4. Inserts to both stores
  5. Updates state file

Why both? Real-time catches immediate memories; batch sweep catches anything the hook missed (edge cases, multi-turn context, etc.)

Deduplication Strategy

We dedupe at insertion time using content hashing:

content_hash = sha256(text[:500]).hexdigest()

If a memory with the same hash exists, insertion returns duplicate_skipped. This prevents the same fact from bloating both stores while allowing re-extraction of modified/updated information.

Current Stats

Store Count Status
GraphDB ~380 memories 🟢
Qdrant ~306 vectors 🟢

Gaps & Limitations

This architecture is not perfect. Here are the known problems:

1. No Memory Expiration

Memories accumulate forever. We need:

  • TTL-based pruning for ephemeral info
  • Importance scoring to keep high-value memories
  • "Forget" mechanism for outdated facts

Impact: Database grows unbounded. Old/wrong preferences persist indefinitely.

2. No Conflict Resolution

If we learn "user prefers X" then later "user prefers Y":

  • Both memories exist
  • No automatic update/invalidation
  • Requires manual dedup

Impact: Conflicting facts coexist. Query results may return outdated info.

3. Extraction Quality Depends on LLM

  • qwen3.5:397b-cloud is fast but not perfect
  • May miss subtle preferences or over-extract noise
  • No quality scoring on extracted memories

Impact: Garbage in, garbage out. No grounding validation.

4. No Cross-Session Linking

  • Memories are extracted per-message
  • No mechanism to link "Episode A" → "Episode B" causally
  • Timeline reconstruction requires post-processing

Impact: Can't answer "what led to this decision?" without manual reconstruction.

5. Qdrant Indexing Threshold

We had to lower indexing_threshold from 10,000 → 100 because vectors weren't being indexed at small scale.

Impact: Default config is tuned for large-scale production, not small personal deployments.

6. Single-Tenant Architecture

  • Both stores run locally
  • No multi-tenancy support
  • Would need Qdrant multi-tenant or separate instances per user for production

Impact: Can't serve multiple users without duplicating infrastructure.

7. No Retrieval Ranking

  • Both stores return results by similarity/date
  • No combined scoring (e.g., recency + relevance + importance)
  • No learning from user feedback on retrieval quality

Impact: May surface irrelevant old memories over relevant new ones.

8. Embedding Model Locked In

  • Using nomic-embed-text for all vectors
  • Switching models requires full re-embedding
  • No A/B testing of embedding quality

Impact: Stuck with initial model choice. Can't experiment with better embeddings.

9. No Memory Compression

  • Raw text stored as-is
  • Long conversations generate many memories
  • No summarization or hierarchical memory

Impact: Storage grows linearly with conversation length.

10. Extraction Hook is Single-Message

  • Real-time hook processes each message independently
  • Context from multi-turn exchanges may be lost
  • Batch sweep helps but isn't real-time

Impact: May miss context that spans multiple messages.


Architecture Diagram

┌─────────────────────────────────────────────────────────────┐
│                     MEMORY EXTRACTION                        │
├─────────────────────────────────────────────────────────────┤
│                                                              │
│   User Message ──▶ Agent ──▶ Response ──▶ Hook Fires        │
│                                              │               │
│                                              ▼               │
│                                    ┌─────────────────┐       │
│                                    │  qwen3.5:397b   │       │
│                                    │  LLM Extraction │       │
│                                    └────────┬────────┘       │
│                                             │                │
│                         ┌───────────────────┴───────────────┐│
│                         ▼                                   ▼│
│               ┌─────────────────┐              ┌───────────┐│
│               │    GraphDB       │              │  Qdrant   ││
│               │   (port 8766)    │              │(port 8765)││
│               │                  │              │           ││
│               │ • Entity graphs  │              │ • Vectors ││
│               │ • Relationships  │              │ • Semantic││
│               │ • Keyword search │              │   search  ││
│               └─────────────────┘              └───────────┘│
│                                                              │
└─────────────────────────────────────────────────────────────┘

What This Enables

  • "Remember that time we..." — Qdrant finds similar contexts
  • "What's the user's timezone?" — GraphDB keyword match
  • "What projects are active?" — GraphDB entity filtering
  • "Summarize our discussions about X" — Qdrant semantic retrieval

This architecture gives the agent persistent, queryable memory across sessions without bloating context windows. The dual-store approach means we get both precise retrieval (GraphDB) and fuzzy semantic search (Qdrant).

But the gaps are real and numerous. This is a v1 architecture — functional but with clear evolution paths:

Near-term fixes:

  • Memory expiration (TTL + importance scoring)
  • Conflict resolution (prefer newer, invalidate outdated)
  • Retrieval ranking (recency + relevance + importance)

Long-term evolution:

  • Hierarchical memory (summarization layers)
  • Multi-tenancy (Qdrant tenant mode)
  • Embedding model swap system
  • Grounding validation (fact-check extraction)

Last updated: 2026-03-27

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment