Releases: uncoalesced/Peridot
Release list
v1.5.4
Full Changelog: v1.5.3...v1.5.4
Peridot v1.5.3 - ZAT-SCS
Peridot - Changelog
Engineered by uncoalesced
[v1.5.3-ZAT_SCS] - 2026-07-18
Name: Peridot v1.5.3 STABLE - Zero-Overhead Active Telemetry and Speculative Context Streaming
Telemetry & Sensory Processing (Predictive Core)
- High-Frequency Telemetry Daemon: Shipped a multi-threaded sensory loop running at 10Hz to continuously calculate real-time user interaction probability P(I_t).
- Asynchronous Keystroke Monitor: Integrated an isolated pynput global hook using a sliding-window array (max capacity 10) to map typing interval acceleration f(C).
- Non-Blocking Acoustic Envelope Tracker: Leveraged a sounddevice InputStream capturing raw acoustic buffers to track root-mean-square (RMS) ambient density g(A).
- State Decay-Acceleration Model: Engineered the dynamic temporal decay equation: P(I_t) = min(1.0, P(I_{t-1}) * e^(-lambda * dt) + w_key * f(C) + w_aud * g(A)), triggering speculative preemption at the critical theta threshold of 0.65.
Sovereign GPU Orchestration & Compute Partitioning
- CUDA MPS Dynamic Slicing: Implemented an active hardware governor that programmatically throttles background volunteer compute (Folding@Home) to 10% SM occupancy when P(I_t) reaches threshold, and suspends it to 0% SM during inference.
- Platform-Agnostic Process Spawning: Implemented fallback mocking on Windows NT environments, while securely managing Linux pipeline subprocess allocation via shell=False and CREATE_NO_WINDOW suppression.
- UVM Memory Weight Pre-mapping: Exported GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 during speculative states to pre-map active LLM weights into the physical GPU page tables during typing, eliminating model load spikes.
Context Streaming & Inference Bypassing
- Asynchronous Context Prefetcher: Built a thread-wrapped loopback REST engine targeting local llama-server slots to pre-load pre-computed KV cache contexts via requests.post to /slots/0/restore.
- Loopback IO Block Exemption: Wrapped REST handshakes in robust ConnectionError and Timeout handlers, ensuring loopback socket failures never freeze the high-frequency sensory loop.
- Prefill-Phase Bypass: Patched the main server /ask Flask routing to detect KernelState.SPECULATIVE_PREPARED. Bypasses the VRAM purge latency path entirely and triggers immediate generation from the staged KV cache, dropping Time-to-First-Token (TTFT) to token generation limits.
Repository Hygiene & Refactoring
- Telemetry Modularization: Segmented core_system/telemetry.py into a fully typed core_system/telemetry/ package containing config, orchestration, and client subdirectories.
- Stability Ledger Preservation: Ported the legacy StabilityLedger class into the new package constructor (init.py), preserving the global namespace for external module imports.
What's Changed
- chore(deps): bump the pip group across 1 directory with 3 updates by @dependabot[bot] in #5
- chore(deps): bump torch from 2.8.0 to 2.12.0 in the pip group across 1 directory by @dependabot[bot] in #6
- chore(deps): bump idna from 3.15 to 3.18 by @dependabot[bot] in #7
- chore(deps): bump msgpack from 1.1.2 to 1.2.0 by @dependabot[bot] in #8
- chore(deps): bump flask-cors from 6.0.2 to 6.0.5 by @dependabot[bot] in #9
- chore(deps): bump pyspellchecker from 0.8.1 to 0.9.0 by @dependabot[bot] in #10
- chore(deps): bump safetensors from 0.7.0 to 0.8.0 by @dependabot[bot] in #15
- chore(deps): bump msgpack from 1.2.0 to 1.2.1 in the pip group across 1 directory by @dependabot[bot] in #16
- chore(deps): bump click from 8.4.0 to 8.4.2 by @dependabot[bot] in #18
- chore(deps): bump faiss-cpu from 1.8.0 to 1.14.3 by @dependabot[bot] in #17
- chore(deps): bump fsspec from 2026.4.0 to 2026.6.0 by @dependabot[bot] in #14
- chore(deps): bump pip-audit from 2.10.0 to 2.10.1 by @dependabot[bot] in #12
- chore(deps): bump torchaudio from 2.6.0 to 2.11.0 by @dependabot[bot] in #11
New Contributors
- @dependabot[bot] made their first contribution in #5
Full Changelog: v1.5.2...v1.5.3
What's Changed
- chore(deps): bump the pip group across 1 directory with 3 updates by @dependabot[bot] in #5
- chore(deps): bump torch from 2.8.0 to 2.12.0 in the pip group across 1 directory by @dependabot[bot] in #6
- chore(deps): bump idna from 3.15 to 3.18 by @dependabot[bot] in #7
- chore(deps): bump msgpack from 1.1.2 to 1.2.0 by @dependabot[bot] in #8
- chore(deps): bump flask-cors from 6.0.2 to 6.0.5 by @dependabot[bot] in #9
- chore(deps): bump pyspellchecker from 0.8.1 to 0.9.0 by @dependabot[bot] in #10
- chore(deps): bump safetensors from 0.7.0 to 0.8.0 by @dependabot[bot] in #15
- chore(deps): bump msgpack from 1.2.0 to 1.2.1 in the pip group across 1 directory by @dependabot[bot] in #16
- chore(deps): bump click from 8.4.0 to 8.4.2 by @dependabot[bot] in #18
- chore(deps): bump faiss-cpu from 1.8.0 to 1.14.3 by @dependabot[bot] in #17
- chore(deps): bump fsspec from 2026.4.0 to 2026.6.0 by @dependabot[bot] in #14
- chore(deps): bump pip-audit from 2.10.0 to 2.10.1 by @dependabot[bot] in #12
- chore(deps): bump torchaudio from 2.6.0 to 2.11.0 by @dependabot[bot] in #11
New Contributors
- @dependabot[bot] made their first contribution in #5
Full Changelog: v1.5.2...v1.5.3
Peridot v1.5.2 - Sanity & Polish Milestone
Peridot - Changelog
Engineered by uncoalesced
[v1.5.2-STABLE] - 2026-06-08
Name: Peridot v1.5.2 STABLE - Sanity & Polish Milestone
Bug Fixes & Stability
- Dynamic VRAM Splitting Heuristics: Rewrote the
_calculate_gpu_layerslogic inconfig.pyto allow partial tensor splitting into system RAM when the model exceeds 75% of total VRAM, preventing catastrophic OOM crashes on borderline 8GB hardware. - FAISS ABI C-Extension Crash: Upgraded
faiss-cputo a newer binary wheel to resolve a fatalValueError: input not a numpy arraymismatch between Numpy 2.x and the older SWIG wrappers on Windows. - Context Window Overflow (400 Bad Request): Engineered an Auto-Truncation Engine in
server.pythat clamps Semantic Memory blocks to 8000 characters and dynamically purges older chat history if the required generation headroom falls below 128 tokens, permanently eliminating Llama.cpp crashes from dense RAG retrievals. - Kinetic Scrolling Polish: Overhauled
ui.pyto calculate text box heights dynamically and applied 144Hz sub-pixel smooth scrolling for the chat matrix.
Peridot v1.5.1
Peridot - Changelog
Engineered by uncoalesced
Peridot v1.5.1 is the biggest update yet! A major step forward in Peridot’s security posture, hardware adaptability and long term usability. This release hardens the local perimeter (removing high risk deserialization and injection paths), introduces intelligent VRAM aware auto-scaling for smoother performance across GPUs, and delivers persistent multi-session conversational memory so Peridot can support real workflows without losing context between runs. Really delighted with how this update has come together and the stability achieved across this version.
[v1.5.1-STABLE] - 2026-06-08
Name: Peridot v1.5.1 STABLE - Security Hardening, Hardware Auto-Scaling & Multi-Session Memory
Security & Compliance (Phase 1: Critical Hotfixes)
- API Key Rotation: Eliminated hardcoded default key (
08101954).config.pynow generates a cryptographically securesecrets.token_hex(32)on first boot and persists to.envviapython-dotenv. Each deployment receives a unique 256-bit key. - Pickle RCE Remediation:
core_system/memory/vault.py- Replacedpickle.load()/dump()withjson.load()/dump()for metadata serialization (.metafile). Removes arbitrary code execution vector via malicious serialized metadata. - Shell Injection Hardening:
ui.py- Removedshell=Truefrom allsubprocess.check_output()calls tonvidia-smi. Arguments now passed as safe arrays (creationflags=0x08000000preserved for window suppression). - CORS Restriction:
server.py- RestrictedCORS(app)to exclusively allowhttp://127.0.0.1:5000andhttp://localhost:5000. Blocks cross-origin requests from external origins. - Multi-Session Conversational Memory: New
core_system/chat_ledger.py- SQLite-backed chat ledger with session CRUD, message logging, and sliding-window history retrieval (get_history(session_id, limit=6)). Integrated into/askroute: accepts optionalsession_id, auto-creates titled sessions, persists user/assistant turns, injects last 6 turns into prompt context, returnssession_idfor client continuity.
Hardware Auto-Scaling (Phase 2: VRAM Detection)
- Dynamic VRAM Detection:
config.py- Added_detect_total_vram_mb()usingnvidia-smi --query-gpu=memory.total(safe subprocess, no shell). Detects total GPU VRAM at module load. - Auto GPU_LAYERS Calculation: If model size < 75% of total VRAM ->
GPU_LAYERS=99(full offload). Else ->GPU_LAYERS=20(partial). CPU-only fallback:GPU_LAYERS=0. - Auto CONTEXT_LENGTH Calculation: >10GB VRAM -> 8192 tokens. <=10GB VRAM -> 4096 tokens. CPU-only -> 2048 tokens.
- UI Integration:
ui.pynow importsTOTAL_VRAM_GBfromconfiginstead of hardcoded8.0or duplicatenvidia-smicall. Model compatibility ratings in dropdown use dynamic VRAM envelope.
RAG Ingestion Overhaul (Phase 3: PyMuPDF & Metadata Tagging)
- Layout-Preserving PDF Extraction:
core_system/memory/vault.py- Replacedfitz.get_text("text")withfitz.get_text("text", sort=True)to preserve visual layout geometry of multi-column tables (e.g., Balance Sheets). - Provenance Tagging:
core_system/memory/vault.pyEvery text chunk now strictly prepends[SOURCE DOC: {filename}]before embedding, enabling the LLM to cite exact documentary sources during generation.
v1.5.0-STABLE: Split Tensor Allocation & Sovereign Kernel Architecture
Peridot - Changelog
Engineered by uncoalesced
[v1.5.0-STABLE] - 2026-06-03
Name: Peridot v1.5.0-STABLE: Split Tensor Allocation & Sovereign Kernel Architecture
Core Engine Architecture & Hardware Arbitration
-
14B Model Pivot & Split Tensor Allocation: Bypassed the planned 7B tier and transitioned the core inference weights to the high logic Qwen2.5-14B-Instruct-Q4_K_M to permanently resolve RAG hallucinations.
GPU_LAYERSinconfig.pywas adjusted to 20, safely splitting the 14B parameter load between the 8GB RTX 5050 GPU and Ryzen 7 CPU RAM. -
Structural Watchdog Hardening & VRAM Purge: Completely upgraded the
_execute_vram_purgemethod inside the VRAM State Machine. Integrated direct NVIDIA driver querying viapynvmlto calculate actual reclaimed physical bytes before allocating tensors, bypassing unreliable OS cache metrics. -
FSM Panic Tuning & Hardware Ceiling: Enforced a strict 7500MB FSM hard ceiling for display driver preservation. Lowered the VRAM purge safety threshold from 1.5GB to 200MB, preventing recursive 503 errors under the new 14B load. If the GPU fails to clear thresholds within a 2.0-second timeout, the FSM instantly trips a KERNEL PANIC.
-
System Prompt Hardening: Re-engineered
build_system_promptto intercept RLHF conversational tropes. Injected explicit constraints forcing the model to refuse queries outside of provided RAG contexts, neutralizing knowledge-bleed defects.
Interface & Operator UX
-
Multi-Tab Notebook Migration: Replaced the legacy single-buffer UI with a
ttk.Notebookframework, introducing three sectors:-
CHAT MATRIX (Text Generation)
-
KERNEL VAULT (Live RAG tracker)
-
SETTINGS (hardware configuration)
-
-
Control Console UI & Live Telemetry: Deployed an industrial, low-overhead administrative dashboard fetching from the secured
/telemetry/stabilityendpoint. Displays live FSM states, dynamic system health scores, total inferences, and panic counts. -
Research Core Toggles: Integrated UI control switches mapping to internal
/telemetry/enableand/research/disableHTTP pathways, allowing operators to authorize or suspend Folding@Home cycles directly from the interface. -
Hardware-Aware Model Swapper: Engineered a dynamic directory scanner that evaluates local
.gguffile sizes against the 8GB RTX 5050 VRAM envelope. Assigns runtime compatibility ratings ([HIGH],[MEDIUM],[LOW/CRITICAL]) and supports GUI hot-swapping throughconfig.pyrewriting. -
144Hz Kinetic Scrolling: Replaced Tkinter's default scroll behavior with a custom 5ms (200Hz) sub-pixel velocity decay loop for high refresh rate rendering.
Ingestion & RAG Pipeline
-
Standalone Command-Line Ingestion (
ingest_vault.py): Added an isolated CLI script to parse, chunk, embed, and commit files directly to the FAISS L2 database without touching the GUI. Embeddings are generated strictly through the CPU-bound Aether-Route. -
Binary PDF & Deep Semantic Search: Upgraded ingestion to dynamically decode binary PDF text layers via PyPDF2. Increased FAISS retrieval depth from
top_k=3totop_k=6for denser multi-document context injection. -
Sliding Window Chunking: Replaced the legacy double-newline chunker with a strict character-clamped fragmentation system (<800 characters) to improve vector precision and reduce dilution.
-
Staging Cleanup Automation: Processed files are automatically relocated from
input/toinput/processed/to prevent recursive re-ingestion and maintain a clean archive.
Fixes, Optimizations & Repository Hygiene
-
AVX2 Matrix Restoration: Reverted the Python environment to a stable NumPy 1.x baseline to resolve the fatal C-extension
_ARRAY_APIcrash during vector initialization. -
Live Buffer Search & UI Extraction: Added a
Ctrl+Freal-time search overlay with highlight support and a_copy_to_clipboardfunction for instant Markdown extraction. -
Global Instantiation Fix: Patched a catastrophic
NameErrorcrash loop by explicitly instantiatingPeridotProductionKernel()in the global scope before engine boot execution. -
Git Integrity & Cleanup: Updated
.gitignoreto block transient FAISS binary files (aether_cold_storage.db) from entering version control. Executed a cleanup sweep locking configuration, UI, and backend changes into the stableorigin/maintree.
"""
Full Changelog: v1.5...v1.5
v1.4.0-STABLE: The TurboQuant Update
Peridot — Changelog
Engineered by uncoalesced
[v1.4.0-STABLE] - 2026-05-14
Name: Peridot v1.4.0 STABLE — TurboQuant Architecture & Sovereign Runtime Finalization
Core Engine Architecture (TurboQuant)
- Deprecated Legacy K-Quants: Purged the default
Llama-3-8B-Instruct (Q4_K_M)baseline due to unacceptable memory bus saturation (~6.6GB VRAM footprint) on 8GB hardware. - Integrated Importance Matrix (I-Quant) Support: Shifted the primary inference engine to natively support
IQ3_XXSand FP4 execution paths. Vaporized ~1.5GB of VRAM overhead while increasing deep-reasoning inference throughput. - Dual-Profile Bootstrapping: Hardcoded two primary runtime profiles inside
config.pyfor dynamic loading:- Deep Thinker Profile:
Llama-3-8B-Instruct (IQ3_XXS)achieving 60.5 t/s at ~4.5GB VRAM. - Agile / Daily Driver Profile:
Qwen 2.5 3B (Q4_K_M)achieving 101.9 t/s at ~2.7GB VRAM.
- Deep Thinker Profile:
- Thermal & Context Limits: Locked the baseline context window to 8192 tokens and dropped the default engine temperature to 0.1 to enforce strict, hallucination-resistant RAG document citation behavior.
- Sliding Context Preservation: Retained the lightweight sliding conversational window internally to preserve RAM stability during prolonged execution sessions.
System Initialization & Security Perimeter
- Setup Wizard Overhaul (
setup.py): Rewrote the installation pipeline into a hardware-aware deployment interface. The wizard now actively profiles GPU VRAM pools and dynamically recommends runtime profiles to prevent Out-Of-Memory (OOM) deployment failures. - Engine Tuning Interface: Injected a dedicated initialization-stage tuning layer allowing operators to explicitly choose between:
- Deep Thinker (maximum reasoning depth)
- Agile / Daily Driver (maximum throughput)
- Manual Matrix Override: Added advanced profile bypass logic exposing raw runtime selection for unsupported or experimental hardware deployments.
- Cryptographic Handshake Integration: Completely abandoned the legacy static
config.jsonauthentication paradigm. System initialization is now locked behind a securely generated.envfile containing a localized 16-byte hexAPI_KEY. - Air-Gap Enforcement: The setup wizard now automatically injects:
into the environment to permanently sever unauthorized HuggingFace telemetry and outbound network synchronization.
HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 - AGPL-3.0 Migration: Upgraded the Peridot kernel licensing structure from MIT to AGPL-3.0 to preserve sovereign-source transparency across derivative deployments and hosted modifications.
State-Machine & Medical Handoff (Folding@Home)
- Zero-Latency Interrupt Protocol: Finalized the WebSocket interrupt architecture. When a prompt hits the API, the Peridot kernel dispatches the Folding@Home pause payload in ~21ms and fully purges the VRAM allocation buffer in under 510ms.
- Aggressive Idle Return: Reduced the
RESEARCH_IDLE_THRESHOLDto 30 seconds to maximize distributed medical research contribution when the user is not actively generating tokens. - Aether-Route CPU Offloading: Hardcoded the semantic embedding engine (
all-MiniLM-L6-v2) to execute strictly on CPU/RAM resources (e.g., Ryzen 7 DDR5 memory footprint), preserving 100% of GPU VRAM for inference and Folding@Home transitions. - Persistent Research Arbitration: Refined VRAM ownership logic to maintain deterministic hardware handoffs without requiring inference engine restarts.
Performance
- TurboQuant Throughput Validation: Established the new stable benchmark baseline:
Llama-3-8B-Instruct (IQ3_XXS)→ 60.5 tokens/secQwen 2.5 3B (Q4_K_M)→ 101.9 tokens/sec
- Reduced VRAM Saturation: Lowered active inference VRAM consumption from ~6.6GB to ~4.5GB under Deep Thinker mode.
- Improved Tensor Allocation Stability: Reduced CUDA allocation pressure during sustained inference + Folding@Home coexistence.
- Enhanced Low-VRAM Runtime Reliability: Optimized execution stability for systems operating below the 8GB VRAM threshold while preserving CPU-only fallback capability.
Architecture
- Aether-Route v1.4: Expanded the routing layer with:
- CPU-isolated semantic embedding
- deterministic VRAM preservation
- improved telemetry-aware execution
- hardware-aware RAG arbitration
- Inference Pipeline Refinement: Refactored orchestration boundaries between:
- embedding execution
- tensor generation
- VRAM arbitration
- telemetry polling
- Hardware-Aware Runtime Scaling: Improved dynamic runtime behavior on:
- constrained VRAM systems
- Ryzen AI processors
- CPU-only deployments
- multitasking inference environments
Security
- Offline Enforcement Hardening: Strengthened sovereign telemetry suppression by enforcing offline execution during setup initialization rather than post-launch configuration.
- Expanded Authentication Isolation: Refined
.envhandling to eliminate residual static credential dependencies. - Runtime Boundary Preservation: Hardened subsystem isolation between:
- telemetry
- inference
- RAG execution
- GhostLogger auditing
- Folding@Home orchestration
Changed
- README Overhaul: Completely rebuilt repository documentation around the v1.4 STABLE runtime architecture, including:
- TurboQuant execution profiles
- sovereign network topology
- VRAM allocation diagrams
- hardware handoff visualization
- Aether-Route topology mapping
- Benchmark Visualization Infrastructure: Added dedicated benchmark illustrations and engineering diagrams for performance validation and architecture transparency.
- Stable Release Transition: Removed beta-stage medical research warnings and finalized the sovereign runtime stack as the official v1.4 STABLE baseline.
Peridot v1.3.2-beta
Peridot - Changelog
Engineered by uncoalesced
[v1.3.2-beta] - 2026-05-13
Name: Peridot v1.3.2 - Memory Deduplication, Meta-Citations & Sovereign Telemetry
Security
- Environment-Level Cryptography: Migrated the
API_KEYcompletely out of Python source code and into a localized.envfile. Established.env.exampleand locked.gitignoreto prevent automated scraping of the host's cryptographic handshake. - Client-Server Handshake Hardening: Patched
core.pyto securely transmit explicitAuthorization: Bearerheaders during both standard inference and system shutdown operations.
Added
- Hash-Based Memory Deduplication: Upgraded
vector_store.pywith a persistentregistry.jsontracking system. It now calculates SHA-256 hashes of all ingested files to prevent redundant vector embeddings and save CPU cycles. - Explicit Source Citations: Engineered the RAG Context-Injection loop in
server.pyto dynamically tag semantic blocks with[SOURCE: filename]. The LLM is now structurally instructed to cite its specific documentary sources during generation. - Automated Ingestion Runner: Shipped
index_all.py, a dedicated ingestion script that automatically scans theinput/zone, extracts text, checks the deduplication registry, and commits new data to the FAISS index.
Changed
- Sovereign Telemetry Override: Forced the
sentence-transformersandhuggingface_hublibraries into strict offline mode via global environment variables (HF_HUB_OFFLINE=1,TRANSFORMERS_OFFLINE=1). This permanently silences network-check warnings and maintains a true air-gapped architecture. - Context Search Depth: Increased the
vector_storeretrieval depth (top_k=3) to feed the LLM denser contextual clusters for more accurate multi-source answers. - Configuration Bootstrap: Rewrote
config.pyto prioritizeload_dotenv()before any secondary module imports, ensuring environment variables govern the entire kernel boot sequence.
Fixed
- 403 Forbidden Handshake Failure: Resolved a critical client-server desynchronization bug where
config.pywas generating conflictingsecrets.token_hex(16)keys for independent processes. The key is now statically anchored to the.envfile. - Flask Payload Rejection: Fixed silent failures in the Neural Link by explicitly forcing the
"Content-Type": "application/json"header inrequests.post()calls originating fromcore.py.
[v1.3.1-beta] - 2026-05-11
Name: Peridot v1.3.1 - Aether-Route Architecture & RAG Synchronization
Added
- Aether-Route (CPU Semantic Router): Engineered a high-efficiency routing layer that offloads vectorization and intent classification to the Ryzen 7 CPU. This preserves VRAM on 4GB/6GB hardware by keeping the embedding matrix strictly in system RAM.
- Split-Payload Architecture: Implemented a decoupled communication protocol between
core.pyandserver.py. The system now transmits an isolated query for semantic mapping and a full prompt for LLM ingestion, preventing chat history from corrupting vector search accuracy. - Unified Aether-Audit: Integrated all RAG subsystems into the ghost auditing module, providing real-time telemetry on VRAM states, routing latency, and inference speeds.
Changed
- Server-Side Cache Centralization: Stripped the duplicate L1 memory cache from
core.pyand centralized all caching logic withinserver.pyto eliminate client-server race conditions and redundant CPU cycles. - Contextual Injection Logic: Updated the RAG pipeline to prepend L2 Vault findings into the system instruction block rather than the user prompt, improving the LLM's adherence to retrieved document data.
- CLI Ingestion Interface: Overhauled
vault.pywith a standalone CLI entry point, enabling batch ingestion of theinput/directory viapython core_system/memory/vault.py ingest.
Fixed
- L1 Cache Signature Collision: Resolved fatal
TypeErrorcrashes where the Router was passing incorrect argument counts toEphemeralCache.add()andEphemeralCache.search(). - GhostLogger Formatting Bug: Fixed a system-wide crash caused by improper string formatting (
TypeError: not all arguments converted) inside the ghost auditing calls. - The "Thermodynamics Loop": Corrected a logic flaw where the router was embedding the entire conversational buffer, leading to 95%+ semantic overlap and causing the system to get stuck repeating previous cached answers.
- Pathing Import Failures: Injected absolute root discovery (
sys.path.insert) intovault.pyandmain.pyto resolveModuleNotFoundErrorwhen running scripts from different terminal directories. - Windows File-Locking: Implemented robust garbage collection and context management in the ingestion pipeline to ensure PDF file pointers are released immediately after vectorization.
Optimized
- Ryzen 7 250 AI Alignment: Optimized the
all-MiniLM-L6-v2embedding engine to utilize Ryzen multi-core efficiency, reducing query vectorization latency to <30ms. - Inference Telemetry: Refined the benchmarking output to provide real-time tokens-per-second (tps) metrics and precise VRAM delta tracking during Aether-Route execution.
[v1.3.0-beta] - 2026-03-30
Name: Peridot v1.3.0 - Dual-Tier Memory Engine & Sterile RAG Architecture
Added
- Dual-Tier Memory Engine: Implemented Layer 1 Ephemeral RAM Cache for instant query interception and Layer 2 FAISS Persistent Vault for local PDF knowledge retrieval.
- Sterile RAG Extraction (
_ask_ai_isolated): Engineered an isolated inference pipeline incore.pyto bypass standard conversational memory, strictly preventing context poisoning and hallucinations when querying the L2 Vault. - Command Routing Subsystem: Deployed
CommandRouterto safely isolate system operations from standard LLM inference. - Dynamic Ingestion Command: Added the
ingestcommand to the router, allowing users to trigger PyMuPDF extraction and SentenceTransformer vectorization on theinput/directory directly from the UI. - Silent Forensic Auditing (GhostLogger): Implemented a non-blocking background logger in
core_system/audit.pythat writes to a 1MB rotating file without polluting the terminal UI. - API Authentication Middleware: Secured the Neural Engine by implementing
@require_authwith Bearer token validation across all operational endpoints. - Comprehensive Benchmark Suite: Engineered custom stress-testing scripts for cold start metrics, VRAM handoff latency, memory stability, and L2 semantic search speed.
Changed
- Vault Intercept Logic: Stripped the automatic PDF database search from the default conversational loop. The Vault is now strictly gated behind the explicit
vault [query]command to preserve VRAM and conversation fluidity. - FAISS Semantic Threshold: Relaxed the L2 distance threshold in
vault.pyfrom 1.5 to 1.85 to allow shorter, highly specific queries to successfully match with longer document chunks. - API Payload Structure: Updated the core inference endpoint from
/chatto/askand modified the required JSON payload key from"prompt"to"command"to align with the v1.3 architecture. - Version String: Bumped system designation from v1.2.1 BETA to v1.3 STABLE.
Fixed
- Context Poisoning Hallucinations: Resolved the bug where the LLM would blend previous chat history with RAG extraction data by forcing sterile prompt injection.
- Windows File-Locking Bug ([WinError 32]): Fixed ingestion crashes by implementing strict context managers (
with fitz.open(...) as doc:) and forced Python garbage collection to release OS-level file pointers after vectorization. - Hardware Architecture Collisions: Hardcoded the Vault Embedding Engine (
all-MiniLM-L6-v2) to run exclusively on the CPU, preventingsm_120architecture clashes with the GPU during LLM inference. - Benchmark Timeout Failures: Rewrote the benchmarking suite to target the correct v1.3 endpoints, inject the required API keys, and account for the 8-billion parameter model load times during cold starts.
- Logger Attribute Error: Patched an upstream integration bug by aliasing the deprecated
.record()method to the native.info()method insidesetup_ghost_logger.
[v1.2.2-beta] - 2026-03-14
Name: Peridot v1.2.2 - Empirical Benchmarking & Security Upgrades
Security
- RAM-Only Authentication (CWE-312 Mitigation): Completely removed disk-based API key storage (
auth.token). The kernel now generates ephemeral cryptographic keys in RAM viaos.environthat evaporate upon shutdown. - Application-Layer Input Sanitization: Implemented a pre-inference regex filter to destroy malicious code injection attempts (e.g., XSS payloads,
os.systemexecution) before they reach the LLM. - Strict Path Traversal Blacklist: The kernel now explicitly blocks attempts to read sensitive system directories (e.g.,
C:\Windows\System32,/etc/) and cryptographic material (e.g.,.ssh/id_rsa,.env). - Subprocess Command Whitelisting: Hardcoded the Medical Research (Folding@Home) WebSocket integration to strictly accept only
pause,unpause,finish, andshutdowndirectives to prevent arbitrary command injection. - Timing-Attack Resistance: Upgraded API authentication in
server.pyto usesecrets.compare_digest()for Bearer token validation, preventing cryptographic timing attacks. - API Rate Limiting: Enforced a strict 60 requests/minute limit per local IP address to mitigate local Denial-of-Service (DoS) and script-kiddie spam.
- Constitution Fallback: If
constitution.jsonis missing or corrupted, the system safely defaults to a zero-trust state (allow_file_read: False).
Added
- Automated Penetration Testing: Shipped `tests/se...
Peridot v1.3.1 - Aether-Route Architecture & RAG Synchronization
Peridot — Changelog
Engineered by uncoalesced
[v1.3.1-beta] - 2026-05-11
Name: Peridot v1.3.1 - Aether-Route Architecture & RAG Synchronization
Added
- Aether-Route (CPU Semantic Router): Engineered a high-efficiency routing layer that offloads vectorization and intent classification to the Ryzen 7 CPU. This preserves VRAM on 4GB/6GB hardware by keeping the embedding matrix strictly in system RAM.
- Split-Payload Architecture: Implemented a decoupled communication protocol between
core.pyandserver.py. The system now transmits an isolated query for semantic mapping and a full prompt for LLM ingestion, preventing chat history from corrupting vector search accuracy. - Unified Aether-Audit: Integrated all RAG subsystems into the ghost auditing module, providing real-time telemetry on VRAM states, routing latency, and inference speeds.
Changed
- Server-Side Cache Centralization: Stripped the duplicate L1 memory cache from
core.pyand centralized all caching logic withinserver.pyto eliminate client-server race conditions and redundant CPU cycles. - Contextual Injection Logic: Updated the RAG pipeline to prepend L2 Vault findings into the system instruction block rather than the user prompt, improving the LLM's adherence to retrieved document data.
- CLI Ingestion Interface: Overhauled
vault.pywith a standalone CLI entry point, enabling batch ingestion of theinput/directory viapython core_system/memory/vault.py ingest.
Fixed
- L1 Cache Signature Collision: Resolved fatal
TypeErrorcrashes where the Router was passing incorrect argument counts toEphemeralCache.add()andEphemeralCache.search(). - GhostLogger Formatting Bug: Fixed a system-wide crash caused by improper string formatting (
TypeError: not all arguments converted) inside the ghost auditing calls. - The "Thermodynamics Loop": Corrected a logic flaw where the router was embedding the entire conversational buffer, leading to 95%+ semantic overlap and causing the system to get stuck repeating previous cached answers.
- Pathing Import Failures: Injected absolute root discovery (
sys.path.insert) intovault.pyandmain.pyto resolveModuleNotFoundErrorwhen running scripts from different terminal directories. - Windows File-Locking: Implemented robust garbage collection and context management in the ingestion pipeline to ensure PDF file pointers are released immediately after vectorization.
Optimized
- Ryzen 7 250 AI Alignment: Optimized the
all-MiniLM-L6-v2embedding engine to utilize Ryzen multi-core efficiency, reducing query vectorization latency to <30ms. - Inference Telemetry: Refined the benchmarking output to provide real-time tokens-per-second (tps) metrics and precise VRAM delta tracking during Aether-Route execution.
[v1.3.0-beta] - 2026-03-30
Name: Peridot v1.3.0 - Dual-Tier Memory Engine & Sterile RAG Architecture
Added
- Dual-Tier Memory Engine: Implemented Layer 1 Ephemeral RAM Cache for instant query interception and Layer 2 FAISS Persistent Vault for local PDF knowledge retrieval.
- Sterile RAG Extraction (
_ask_ai_isolated): Engineered an isolated inference pipeline incore.pyto bypass standard conversational memory, strictly preventing context poisoning and hallucinations when querying the L2 Vault. - Command Routing Subsystem: Deployed
CommandRouterto safely isolate system operations from standard LLM inference. - Dynamic Ingestion Command: Added the
ingestcommand to the router, allowing users to trigger PyMuPDF extraction and SentenceTransformer vectorization on theinput/directory directly from the UI. - Silent Forensic Auditing (GhostLogger): Implemented a non-blocking background logger in
core_system/audit.pythat writes to a 1MB rotating file without polluting the terminal UI. - API Authentication Middleware: Secured the Neural Engine by implementing
@require_authwith Bearer token validation across all operational endpoints. - Comprehensive Benchmark Suite: Engineered custom stress-testing scripts for cold start metrics, VRAM handoff latency, memory stability, and L2 semantic search speed.
Changed
- Vault Intercept Logic: Stripped the automatic PDF database search from the default conversational loop. The Vault is now strictly gated behind the explicit
vault [query]command to preserve VRAM and conversation fluidity. - FAISS Semantic Threshold: Relaxed the L2 distance threshold in
vault.pyfrom 1.5 to 1.85 to allow shorter, highly specific queries to successfully match with longer document chunks. - API Payload Structure: Updated the core inference endpoint from
/chatto/askand modified the required JSON payload key from"prompt"to"command"to align with the v1.3 architecture. - Version String: Bumped system designation from v1.2.1 BETA to v1.3 STABLE.
Fixed
- Context Poisoning Hallucinations: Resolved the bug where the LLM would blend previous chat history with RAG extraction data by forcing sterile prompt injection.
- Windows File-Locking Bug ([WinError 32]): Fixed ingestion crashes by implementing strict context managers (
with fitz.open(...) as doc:) and forced Python garbage collection to release OS-level file pointers after vectorization. - Hardware Architecture Collisions: Hardcoded the Vault Embedding Engine (
all-MiniLM-L6-v2) to run exclusively on the CPU, preventingsm_120architecture clashes with the GPU during LLM inference. - Benchmark Timeout Failures: Rewrote the benchmarking suite to target the correct v1.3 endpoints, inject the required API keys, and account for the 8-billion parameter model load times during cold starts.
- Logger Attribute Error: Patched an upstream integration bug by aliasing the deprecated
.record()method to the native.info()method insidesetup_ghost_logger.
[v1.2.2-beta] - 2026-03-14
Name: Peridot v1.2.2 - Empirical Benchmarking & Security Upgrades
Security
- RAM-Only Authentication (CWE-312 Mitigation): Completely removed disk-based API key storage (
auth.token). The kernel now generates ephemeral cryptographic keys in RAM viaos.environthat evaporate upon shutdown. - Application-Layer Input Sanitization: Implemented a pre-inference regex filter to destroy malicious code injection attempts (e.g., XSS payloads,
os.systemexecution) before they reach the LLM. - Strict Path Traversal Blacklist: The kernel now explicitly blocks attempts to read sensitive system directories (e.g.,
C:\Windows\System32,/etc/) and cryptographic material (e.g.,.ssh/id_rsa,.env). - Subprocess Command Whitelisting: Hardcoded the Medical Research (Folding@Home) WebSocket integration to strictly accept only
pause,unpause,finish, andshutdowndirectives to prevent arbitrary command injection. - Timing-Attack Resistance: Upgraded API authentication in
server.pyto usesecrets.compare_digest()for Bearer token validation, preventing cryptographic timing attacks. - API Rate Limiting: Enforced a strict 60 requests/minute limit per local IP address to mitigate local Denial-of-Service (DoS) and script-kiddie spam.
- Constitution Fallback: If
constitution.jsonis missing or corrupted, the system safely defaults to a zero-trust state (allow_file_read: False).
Added
- Automated Penetration Testing: Shipped
tests/security_tests.py, an automated Red Team suite to barrage the local kernel and verify the containment field holds against actual payloads. - Empirical Benchmarking Suite: Added
benchmarks/vram_test.pyandbenchmarks/inference_test.pyto measure precise hardware metrics rather than relying on estimates. - Security & Benchmark Policies: Published
SECURITY.mddetailing the threat model and responsible disclosure, alongsideBENCHMARKING.mdfor community hardware testing. - GhostLogger: Integrated a zero-latency, asynchronous JSONL telemetry logger (
logs/ghost_audit.jsonl) for tracking system state changes without blocking the main OS loop. - Security Logger: Added a dedicated forensic logger (
logs/security.log) to quietly record all blocked file accesses, rejected inputs, and authentication failures.
Changed
- Verified Hardware Metrics: Replaced the estimated README hardware claims with empirical data tested on an RTX 5050: 6.55ms VRAM hot-swap latency and 45-55 t/s Llama-3 8B inference speed.
Fixed
- CodeQL CWE-312 Vulnerability: Permanently patched clear-text storage of sensitive information by migrating the API key entirely to ephemeral RAM.
[v1.2.1-beta] - 2026-03-10
Name: Peridot v1.2.1 - Security Changes, Patched Memory Leaks & Secured Command Routing
Security
- Localhost API Authentication: Implemented dynamic API key generation (
auth.token) and strict Bearer token authentication across all Flask endpoints to prevent unauthorized local processes from hijacking GPU resources. - Hardened System Directives: Updated the core AI system prompt to establish a hard boundary against OS-level destruction. Peridot now explicitly refuses commands that attempt to delete system files, compromise host OS integrity, or exfiltrate sensitive data over the network, while maintaining uncensored operation for standard tasks.
Added
- True Hardware Telemetry: Integrated
pynvml(vianvidia-ml-py) into the VRAM State Machine to provide real-time NVIDIA GPU memory tracking and reporting. - Server Health Polling: Added a
/healthendpoint toserver.pyto allow client processes to safely verify engine readiness before mounting the interface.
Changed
- WebSocket VRAM Hot-Swaps: Completely removed legacy CLI subprocess polling for Folding@Home. The VRAM State Machine now communicates directly w...