Persistent semantic memory for AI agents with intelligent decision-making.
π Created by: TheProdSDE
Most AI memory systems simply retrieve and inject past context into every prompt. This leads to:
- π° Higher token costs - Unnecessary context in every query
- π Inconsistent responses - No validation of stale or incorrect memories
- β±οΈ Poor performance - Always processing memory, even when irrelevant
- π€ No intelligence - Memory is treated as a dumb cache
Agent Memory is a decision layer that intelligently chooses when and how to use memory:
flowchart TD
A[User Query] --> B[Resolve Memory]
B --> C[Decision Engine]
C -->|High confidence match| D[π Replay: Return stored answer]
C -->|Moderate match| E[π Restore: Inject as context]
C -->|Needs validation| F[β
Verify: Validate before reuse]
C -->|No match| G[β None: Answer from scratch]
style D fill:#0d47a1,color:#ffffff
style E fill:#e65100,color:#ffffff
style F fill:#1b5e20,color:#ffffff
style G fill:#b71c1c,color:#ffffff
Benefits:
- β Response consistency - Reuse proven answers
- β Lower token usage - Only inject when beneficial
- β Faster responses - Instant replay for repeated queries
- β Better long-term behavior - Agents learn when to trust memory
flowchart LR
subgraph Input["Input Layer"]
Q[User Query]
end
subgraph Retrieval["Retrieval Layer"]
BM25[BM25 Keyword Search]
Vector[Vector Semantic Search]
Fusion[Reciprocal Rank Fusion]
end
subgraph Decision["Decision Layer"]
Policy[Scoring Policy]
Engine[Decision Engine]
end
subgraph Storage["Storage Layer"]
Chroma[ChromaDB]
DB[(Local Persistence)]
end
subgraph Output["Output Layer"]
Replay[Replay Action]
Restore[Restore Action]
Verify[Verify Action]
None[None Action]
end
Q --> BM25
Q --> Vector
BM25 --> Fusion
Vector --> Fusion
Fusion --> Policy
Policy --> Engine
Engine --> Storage
Storage --> Engine
Engine --> Replay
Engine --> Restore
Engine --> Verify
Engine --> None
style Input fill:#4a148c,color:#ffffff
style Retrieval fill:#0d47a1,color:#ffffff
style Decision fill:#e65100,color:#ffffff
style Storage fill:#1b5e20,color:#ffffff
style Output fill:#b71c1c,color:#ffffff
- Query Input: User query enters the system
- Hybrid Retrieval: BM25 (keyword) + Vector (semantic) search with RRF fusion
- Policy Scoring: Multi-factor scoring (semantic + recency + confidence + usage)
- Decision Engine: Intelligently selects the best action
- Action Execution: Returns appropriate response based on decision
Memory is not automatically injected. Each query results in one of four actions:
| Action | Behavior | Use Case |
|---|---|---|
| Replay | Return previous answer | Exact or near-identical queries |
| Restore | Inject memory as context | Similar queries needing adaptation |
| Verify | Validate before reuse | Facts, workflows, tool outputs |
| None | Ignore memory | Unrelated queries |
flowchart TD
A[Query] --> B[BM25 Search]
A --> C[Vector Search]
B --> D[Reciprocal Rank Fusion]
C --> D
D --> E[Policy Reranking]
E --> F[Top K Results]
style B fill:#f57f17,color:#ffffff
style C fill:#2e7d32,color:#ffffff
style D fill:#1565c0,color:#ffffff
style E fill:#6a1b9a,color:#ffffff
Policy scoring considers:
- π Semantic similarity (70% weight)
- π Recency (15% weight)
- β Confidence score (20% weight)
- π Usage frequency (10% weight)
Store memories with type and scope for better organization:
Memory Types:
conversation- Chat historyfact- Verifiable informationworkflow- Step-by-step processesdocument- Long-form contenttool_output- API/tool responsescode- Code snippetssummary- Consolidated memoriespreference- User preferences
Scopes:
session- Current conversationuser- User-specificproject- Project-specificworkspace- Workspace-wideteam- Team-sharedglobal- Application-wide
Automatic expiration with flexible TTL:
# Absolute time
memory.remember(query, response, ttl="30d") # 30 days
memory.remember(query, response, ttl="2h") # 2 hours
# Relative time
memory.remember(query, response, ttl=3600) # 1 hour in secondsFull transparency into decision-making:
decision = memory.resolve(query)
print(decision) # Decision object
print(decision.explain()) # Detailed score breakdown# From PyPI
pip install agent-memory-sdk
# From source (development)
git clone https://github.com/TheProdSDE/agent-memory.git
cd agent-memory
pip install -e ".[dev]"from agent_memory import Memory, MemoryAction, MemoryType
# Initialize memory
memory = Memory(persist_dir=".agent_memory")
# Store a memory
memory.remember(
query="How do I reset my password?",
response="Go to Settings β Security β Reset Password and follow the email link.",
type=MemoryType.CONVERSATION,
tags=["auth", "faq"],
confidence=0.95
)
# Store a fact that requires verification
memory.remember(
query="Current API rate limit",
response="1000 requests/minute per API key.",
type=MemoryType.FACT,
requires_verification=True
)
# Resolve a query
decision = memory.resolve("How do I reset my password?")
# Handle the decision
match decision.action:
case MemoryAction.REPLAY:
print(f"Replaying: {decision.response}")
case MemoryAction.RESTORE:
context = memory.format_restore_context(decision)
print(f"Context: {context}")
# Use with your LLM: llm(query, context=context)
case MemoryAction.VERIFY:
print(f"Verify: {decision.memory.response}")
# Validate with tools before reuse
case MemoryAction.NONE:
print("No relevant memory - answer from scratch")# Memory management
memory.remember(query, response, *, type, scope, tags, confidence, ttl, metadata)
memory.get(memory_id)
memory.list(limit=100, offset=0, *, scope, include_archived, type)
memory.forget(memory_id)
memory.archive(memory_id)
# Query and resolve
decision = memory.resolve(query, *, mode, top_k, scope, enable_verify)
# Maintenance
memory.cleanup(delete=False) # Mark expired as expired
memory.cleanup(delete=True) # Delete expired
memory.consolidate(similarity_threshold=0.95) # Merge duplicates
memory.stats() # Get usage statisticsclass MemoryDecision:
action: MemoryAction # REPLAY, RESTORE, VERIFY, NONE
confidence: float # 0.0 - 1.0
query: str # Original query
reason: str # Human-readable reason
reasons: list[str] # Detailed reasons
response: str | None # For REPLAY action
memory: MemoryEntry | None # For REPLAY/VERIFY
context: list[RetrievalResult] # For RESTORE/VERIFY
def explain(self) -> str: # Detailed score breakdown
return "..."Expose Agent Memory as MCP tools for Cursor, VS Code, and other MCP-compatible agents.
Add to ~/.cursor/mcp.json:
{
"mcpServers": {
"agent-memory": {
"command": "agent-memory-mcp",
"env": {
"AGENT_MEMORY_DIR": "~/.agent_memory"
}
}
}
}| Tool | Description |
|---|---|
remember_memory |
Store a query/response pair |
resolve_memory |
Retrieve and decide action |
list_memories |
List with pagination |
get_memory |
Fetch single memory |
forget_memory |
Delete memory |
archive_memory |
Archive memory |
consolidate_memories |
Merge duplicates |
{
"mcpServers": {
"agent-memory": {
"command": "docker",
"args": [
"run", "--rm", "-i",
"-v", "agent_memory_data:/home/appuser/.agent_memory",
"ghcr.io/theprodsde/agent-memory:latest",
"agent-memory-mcp"
]
}
}
}# Show help
agent-memory --help
# Store a memory
agent-memory remember "query" "response" \
--type conversation \
--scope user \
--ttl 30d
# Resolve a query
agent-memory resolve "query" --explain
# Show statistics
agent-memory stats
# Cleanup expired memories
agent-memory cleanup --delete
# Run benchmark
agent-memory benchmark --seed --repeat 3
# Run evaluation
agent-memory eval --datasets ./benchmarks/datasets# Quick benchmark with default queries
agent-memory benchmark
# With seeded data from eval datasets
agent-memory benchmark --seed --repeat 3
# Custom baseline comparison
agent-memory benchmark --baseline-ms 500# Run all datasets
agent-memory eval
# Specific dataset directory
agent-memory eval --datasets ./benchmarks/datasetsIncluded Datasets:
coding_agent.json- Code-related queriescustomer_support.json- Support scenariosresearch_agent.json- Research workflows
The project uses automated CI/CD via GitHub Actions. Releases are triggered by pushing a version tag:
# Create and push a version tag (triggers full release pipeline)
git tag v0.1.3
git push origin v0.1.3When you push a tag matching v* (e.g., v0.1.3, v1.0.0, v2.0.0-beta.1):
| Step | Description |
|---|---|
| 1οΈβ£ Test | Runs tests on Python 3.10, 3.11, 3.12, 3.13 |
| 2οΈβ£ Benchmark | Runs performance benchmarks |
| 3οΈβ£ Docker | Builds and tests multi-stage Docker image |
| 4οΈβ£ Publish | Builds package β Publishes to PyPI β Creates GitHub Release |
| Artifact | Location |
|---|---|
| PyPI Package | pip install agent-memory-sdk==0.1.3 |
| GitHub Release | https://github.com/theprodsde/agent-memory/releases/tag/v0.1.3 |
| Docker Image | ghcr.io/theprodsde/agent-memory:v0.1.3 (if configured) |
| Source Archives | Auto-attached to GitHub Release |
Use Semantic Versioning with optional pre-release suffixes:
v1.0.0- Stable releasev1.0.1- Patch releasev1.1.0- Minor releasev2.0.0- Major releasev1.0.0-alpha.1- Alpha pre-releasev1.0.0-beta.2- Beta pre-releasev1.0.0-rc.1- Release candidate
- PyPI Token - Stored as
PYPI_API_TOKENin GitHub repository secrets - GitHub Token - Automatically provided as
GITHUB_TOKEN - Branch Protection - Recommended: require PR reviews before merging to main
# 1. Ensure you're on main with latest changes
git checkout main
git pull origin main
# 2. Create version tag
git tag v0.1.3
# 3. Push tag (triggers CI/CD)
git push origin v0.1.3
# 4. Monitor workflow
# https://github.com/theprodsde/agent-memory/actions# Delete local tag
git tag -d v0.1.3
# Delete remote tag (also deletes GitHub Release)
git push origin --delete v0.1.3
# Note: PyPI packages CANNOT be deleted, only yanked
# twine yank agent-memory 0.1.3- Hybrid retrieval (BM25 + Vector + RRF fusion)
- Decision engine (replay / restore / verify / none)
decision.explain()observability- Benchmark & evaluation datasets
- TTL + memory states + cleanup
- CLI (remember, resolve, stats, benchmark, eval)
- MCP support for Cursor and other clients
- Comprehensive documentation
- All tests passing
- CI/CD pipeline
- Docker support
| Feature | Status | ETA |
|---|---|---|
| Async API | β Completed | v0.1.0-alpha |
| SQLite backend | β Completed | v0.1.0-alpha |
| Redis backend | π Planned | v0.2.0 |
| Postgres backend | π Planned | v0.3.0 |
| FastAPI server + dashboard | π Planned | v0.3.0 |
| Memory graph | π Planned | v0.4.0 |
| Confidence learning | π Planned | v0.4.0 |
| Multi-agent support | π Planned | v0.5.0 |
| Component | Technology |
|---|---|
| Language | Python 3.10+ |
| Storage | ChromaDB (local embeddings) |
| Retrieval | BM25 + Vector Search + RRF |
| Interface | MCP (Model Context Protocol) |
| CLI | argparse |
| Testing | pytest + pytest-asyncio |
| Linting | ruff |
| Type Checking | mypy |
| CI/CD | GitHub Actions |
| Container | Docker + docker-compose |
No API keys required - Everything runs locally!
- Getting Started - Installation and basic usage
- Architecture - Deep dive into the system design
- Memory Model - Understanding memory types and states
- Policies - Customizing scoring and decision logic
- Benchmarks - Performance metrics and evaluation
- FAQ - Common questions and troubleshooting
Contributions are welcome! Please follow these steps:
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Make your changes
- Run tests (
python -m pytest tests/) - Run linting (
ruff check agent_memory/ tests/) - Commit your changes (
git commit -m 'Add amazing feature') - Push to the branch (
git push origin feature/amazing-feature) - Open a Pull Request
# Clone the repository
git clone https://github.com/TheProdSDE/agent-memory.git
cd agent-memory
# Create virtual environment
python -m venv .venv
source .venv/bin/activate # or .venv\Scripts\activate on Windows
# Install in development mode
pip install -e ".[dev]"
# Install pre-commit hooks
pip install pre-commit
pre-commit install
# Run tests
make test
# Run all checks
make checkThis project is licensed under the MIT License - see the LICENSE file for details.
- ChromaDB - Vector database
- Rank-BM25 - BM25 implementation
- MCP - Model Context Protocol
- FastMCP - MCP server framework
- Issues: GitHub Issues
- Discussions: GitHub Discussions
- Email: theprodsde@gmail.com
Agent Memory helps agents decide: Replay β Restore β Verify β Ignore
Built with β€οΈ by TheProdSDE