Sandbox
@savantskie/persistent-ai-memory

Local memory system for Copilot and MCP assistants

This repo gives AI assistants a persistent memory layer backed by SQLite and embeddings. It can store memories, search them semantically, track conversations, and expose the data through MCP or direct Python use.

235 stars29 forksPythonUpdated 4mo ago
Who it's for

Builders who want their agent to keep usable memory across chats, tools, and projects.

What it delivers

You can stop re-explaining context and let your agent retrieve relevant memory from past work.

What it does

Persistent memory storage

Stores long-term memories in SQLite so they survive across sessions.

Semantic search

Uses embeddings to find related memories by meaning instead of exact words.

Conversation tracking

Keeps chat history and links it to context for later retrieval.

MCP server

Exposes memory operations through Model Context Protocol for compatible assistants.

OpenWebUI short-term memory

Injects and filters relevant memory during chat with OpenWebUI.

Tool call logging

Records tool usage so the system can analyze patterns and behavior.

Health checks and tests

Includes scripts and tests to verify the database, embeddings, and integration flow.

How to get it

  1. 1Run
    # Linux/macOS
    pip install git+https://github.com/savantskie/persistent-ai-memory.git
    
    # Windows (same command, just use Command Prompt or PowerShell)
    pip install git+https://github.com/savantskie/persistent-ai-memory.git
  2. 2Run
    python tests/test_health_check.py
  3. 3Expected output
    [✓] Imported ai_memory_core
    [✓] Found embedding_config.json
    [✓] System health check passed
    [✓] All health checks passed! System is ready to use.

README

Persistent AI Memory System v1.5.0

License: MIT Python 3.8+ Release

🌟 Community Call to Action: Have you made improvements or additions to this system? Submit a pull request! Every contributor will be properly credited in the final product.

GITHUB LINK - https://github.com/savantskie/persistent-ai-memory.git


🆕 What's New in v1.5.0 (March 28, 2026)

Major Architectural Rewrite: OpenWebUI-Native Integration

  • OpenWebUI-first design - AI Memory System now deeply integrated into OpenWebUI via plugin (primary deployment method)
  • Advanced short-term memory - sophisticated memory extraction, filtering, and injection for chat conversations
  • User ID & Model ID isolation - strict multi-tenant support with configurable enforcement for security and tracking
  • Complete system portability - all hardcoded paths replaced with environment variables (works anywhere)
  • Generic class names - removed all Friday-specific branding (FridayMemorySystem → AIMemorySystem)
  • Production-ready - enhanced error handling, validation, and logging throughout

Upgrade from v1.1.0: See CHANGELOG.md for migration guide.


📚 Documentation Guide

Choose your starting point:

I want to...Read thisTime
Get started quicklyREDDIT_QUICKSTART.md5 min
Install the systemINSTALL.md10 min
Understand configurationCONFIGURATION.md15 min
Check system healthTESTING.md10 min
Use the APIAPI.md20 min
Deploy to productionDEPLOYMENT.md15 min
Fix a problemTROUBLESHOOTING.mdvaries
See examplesexamples/README.md15 min

🚀 Quick Start (30 seconds)

Installation

# Linux/macOS
pip install git+https://github.com/savantskie/persistent-ai-memory.git

# Windows (same command, just use Command Prompt or PowerShell)
pip install git+https://github.com/savantskie/persistent-ai-memory.git

First Validation

python tests/test_health_check.py

Expected output:

[✓] Imported ai_memory_core
[✓] Found embedding_config.json
[✓] System health check passed
[✓] All health checks passed! System is ready to use.

💡 What This System Does

Persistent AI Memory provides sophisticated memory management for AI assistants:

  • 📝 OpenWebUI Short-Term Memory Plugin - Intelligent memory extraction and injection directly in chat conversations
  • 🧠 Persistent Memory Storage - SQLite databases for structured, searchable long-term memories
  • 🔍 Semantic Search - Vector embeddings for intelligent memory retrieval and relevance scoring
  • 💬 Conversation Tracking - Multi-platform conversation history capture with context linking
  • 🎯 Smart Memory Filtering - Advanced blacklist/whitelist and relevance scoring to inject only what matters
  • 🧮 Tool Call Logging - Track and analyze AI tool usage patterns and performance
  • 🔄 Self-Reflection - AI insights into its own behavior and memory patterns
  • 📱 Multi-Platform Support - Works with OpenWebUI (primary), LM Studio, VS Code, and any MCP-compatible assistant
  • 🎨 MCP Server - Standard Model Context Protocol for cross-platform integration

⚙️ System Architecture

Five Specialized Databases

~/.ai_memory/
├── conversations.db      # Chat messages and conversation history
├── ai_memories.db       # Curated long-term memories
├── schedule.db          # Appointments and reminders
├── mcp_tool_calls.db    # Tool usage logs and reflections
└── vscode_project.db    # Development session context

Configuration Files

~/.ai_memory/
├── embedding_config.json   # Embedding provider setup
└── memory_config.json      # Memory system defaults

🎯 Core Features

Memory Operations

  • store_memory() - Save important information persistently
  • search_memories() - Find memories using semantic search
  • list_recent_memories() - Get recent memories without searching

Conversation Tracking

  • store_conversation() - Store user/assistant messages
  • search_conversations() - Search through conversation history
  • get_conversation_history() - Retrieve chronological conversations

Tool Integration

  • log_tool_call() - Record MCP tool invocations
  • get_tool_call_history() - Analyze tool usage patterns
  • reflect_on_tool_usage() - Get AI insights on tool patterns

System Health

  • get_system_health() - Check databases, embeddings, providers
  • built-in health check - python tests/test_health_check.py

🔌 Embedding Providers

Choose your embedding service:

ProviderSpeedQualityCost
Ollama (local)⚡⚡⭐⭐⭐FREE
LM Studio (local)⭐⭐⭐⭐FREE
OpenAI (cloud)⚡⚡⭐⭐⭐⭐⭐$$$

See CONFIGURATION.md for setup instructions for each provider.


� Important: User ID & Model ID Requirements

All memory operations require user_id and model_id parameters for data isolation and tracking.

This ensures:

  • Multi-user safety - Each user's memories are completely isolated
  • Model tracking - Different AI models can maintain separate memories
  • Audit trail - All operations are traceable to the user and model

Configuration Options

By default, user_id and model_id are required. You can change this in memory_config.json:

{
  "tool_requirements": {
    "require_user_id": true,
    "require_model_id": true,
    "default_user_id": "default_user",
    "default_model_id": "default_model"
  }
}
  • require_user_id/require_model_id: true → Strict mode (recommended for production, security-focused, or multi-user systems)
  • require_user_id/require_model_id: false → Use defaults instead (simpler for single-user/single-model setups)

For AI Assistants: Auto-Fill in System Prompt

To make your AI automatically provide these values, add this to its system prompt:

When using memory system tools (store_memory, search_memories, etc.), 
ALWAYS include these parameters:
- user_id='your_user_identifier' (e.g., 'nate_user_1')
- model_id='your_model_name' (e.g., 'llama-2:7b' or 'gpt-4')

If the actual values are unknown, use safe defaults:
- user_id='default_user'
- model_id='default_model'

This isolates memories per user and tracks which AI model generated each memory.

Examples

With user_id and model_id:

# Memories are stored with full isolation
await system.store_memory(
    "User likes Python", 
    user_id="alice", 
    model_id="gpt-4"
)

# Search returns only this user's memories for this model
results = await system.search_memories(
    "programming", 
    user_id="alice", 
    model_id="gpt-4"
)

Without strict requirements (if disabled):

# Uses defaults from memory_config.json
await system.store_memory("User likes Python")  # user_id="default_user", model_id="default_model"

See API.md for complete parameter documentation.


�🔄 Integration Methods (Choose One)

1. OpenWebUI Plugin (Recommended)

Primary deployment method - Deep integration for sophisticated memory management:

  • Deploy ai_memory_short_term.py as an OpenWebUI Function
  • Automatically extracts memories from conversations
  • Intelligently injects relevant memories before AI response
  • Configurable memory scoring, filtering, and injection preferences
  • No additional setup required beyond copying file into OpenWebUI Functions editor

Installation:

  1. In OpenWebUI: Settings → Functions → +New Function
  2. Paste entire ai_memory_short_term.py file
  3. Set trigger to Inlet (runs before model response)
  4. Configure memory preferences via function settings

2. MCP Server (Alternative Platforms)

Use with any MCP-compatible AI assistant (Claude, custom integrations, etc.):

# Via mcpo
python -m ai_memory_mcp_server

# Or make streamable for OpenWebUI's alternative integration
# (OpenWebUI supports both plugin and streamable MCP methods)

3. Standalone Library (Custom Implementations)

Use memory capabilities directly in your Python code:

from ai_memory_core import AIMemorySystem
system = AIMemorySystem()
await system.store_memory("Important information", user_id="user1", model_id="model1")
results = await system.search_memories("query", user_id="user1", model_id="model1")

🛠️ Development & Examples

Ready-to-use examples:

python examples/basic_usage.py          # Store and search memories
python examples/advanced_usage.py       # Conversation tracking and tool logging
python examples/performance_tests.py    # Benchmark operations

Full API reference: API.md


📖 Learning Resources


� System Sophistication

This is a significantly enhanced version of traditional memory systems:

FeatureTraditionalAI Memory System
Memory ExtractionManual/StaticLLM-powered intelligent extraction
FilteringSimple keyword matchingMulti-layer semantic + relevance scoring
Memory InjectionAll available memoriesSmart filtering - only inject relevant
Duplicate PreventionText matchingEmbedding-based semantic deduplication
Importance ScoringNot trackedDynamic importance analysis
Memory NormalizationN/AAutomatic format standardization
Context AwarenessLimitedFull conversation context integration
Tool IntegrationBasic loggingDeep reflection and pattern analysis
Error HandlingMinimalComprehensive validation and recovery
PerformanceN/AOptimized with async operations

Result: An AI assistant that truly learns from and adapts to your preferences over time.


�🤝 Contributing

We welcome contributions! See CONTRIBUTORS.md for:

  • Development setup instructions
  • How to run tests
  • Code style guidelines
  • Contribution process

📄 License

MIT License - Feel free to use this in your own AI projects!

See LICENSE for details.


🙏 Acknowledgments

This project represents a unique collaboration:

  • @savantskie - Project vision, architecture, testing
  • GitHub Copilot - Core implementation and system design
  • ChatGPT - Architectural guidance and insights

Special thanks to the AI and open-source communities for inspiration and support.


📞 Need Help?

  1. Start with: TESTING.md → Run health check
  2. Then check: TROUBLESHOOTING.md → Find your issue
  3. Or visit: COMMUNITY.md → Get help from community
  4. Or open: GitHub Issues

⭐ If this project helps you build better AI assistants, please give it a star!

Built with determination, debugged with patience, designed for the future of AI.

Files in the repo

Repository payload41 top-level entries
  • examples
  • scripts
  • tests
  • .env.example
  • .gitignore
  • ai_memory_core.py
  • ai_memory_mcp_server.py
  • ai_memory_normalization_migration.py
  • ai_memory_short_term.py
  • API.md
  • check_db.py
  • COMMIT_MESSAGE_v1.1.0.txt
  • COMMUNITY.md
  • CONFIGURATION.md
  • CONTRIBUTORS.md
  • core_identity.py
  • database_maintenance.py
  • database_maintenance.py.bak
  • DEPLOYMENT.md
  • embedding_config.json
  • GITHUB_GUIDE.md
  • install.bat
  • INSTALL.md
  • install.sh
  • INSTALLATION_COMPLETE.md
  • KOBOLDCPP_INTEGRATION.md
  • LICENSE
  • memory_config.json
  • port_manager.py
  • PROJECT_STRUCTURE.md
  • README.md
  • REDDIT_QUICKSTART.md
  • requirements.txt
  • settings.py
  • setup.py
  • start_maintenance_service.bat
  • start_maintenance_service.ps1
  • tag_manager.py
  • TESTING.md
  • TROUBLESHOOTING.md
  • utils.py

Discussion (0)

Ask about usage, or say what you built with it

Sign in to join the discussion.

No comments yet. Be the first to say what this is good for.

More connectors

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

43k

Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code

14k
okf-memory/
okf-agent-memory

Git-native persistent memory for AI coding agents. Implements Google OKF v0.2 with sub-300µs in-memory BM25 search, embedded MCP server, and progressive disclosure. Slashes token bloat by 80% with zero external databases or dependencies. Built in pure Go.

547
tirth8205/
code-review-graph

Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-repo workflows.

31k
2akouwu/
reverify

Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI.

1.1k
t8y2/dbxConnectors

20 MB lightweight cross-platform database client for 90+ databases, including MySQL, PostgreSQL, SQLite, Redis, MongoDB, DuckDB, SQL Server, and Dameng. Built-in AI, MCP Server, CLI, desktop and Docker. | 轻量级跨平台数据库管理工具,支持 MySQL、PostgreSQL、SQLite、Redis、MongoDB、达梦等 90+ 数据库,提供桌面端、Docker、CLI、内置 AI 助手和 MCP Server。

19k