-
-
Notifications
You must be signed in to change notification settings - Fork 68
Development Guide
This page covers local development setup, Docker builds, testing, and contributing to GPT Home.
| Tool | Version | Purpose |
|---|---|---|
| Python | 3.11 | Backend runtime |
| Node.js | 18.x | Frontend build |
| Docker | 24.x+ | Containerization |
| Docker Compose | 2.x | Service orchestration |
| Git | 2.x+ | Version control |
The easiest way to develop is using the Docker dev profile with hot reload:
# Clone repository
git clone https://github.com/judahpaul16/gpt-home.git
cd gpt-home
# Copy environment template
cp .env.example .env
# Edit .env with your API key
nano .env
# Start development environment
COMPOSE_PROFILES=dev docker compose up
# Access:
# - Web Interface: http://localhost (nginx routes to web-dev)
# - FastAPI Backend: http://localhost:8000 (direct access)
# - API Docs: http://localhost:8000/docsChanges to files in ./src/frontend are automatically hot-reloaded via volume mount.
For testing the production build locally:
# Production is default (or explicitly set)
docker compose up -d
# or: COMPOSE_PROFILES=prod docker compose up -d
docker compose logs -fgpt-home/
βββ .env # Environment configuration (create from .env.example)
βββ .env.example # Environment template
βββ docker-compose.yml # Service definitions with profiles
βββ compose/ # Dockerfiles for each service
β βββ app/
β β βββ Dockerfile # Voice assistant + FastAPI backend
β βββ web/
β β βββ Dockerfile # Production: multi-stage React build β nginx
β β βββ Dockerfile.dev # Development: React dev server (hot reload)
β βββ spotify/
β βββ Dockerfile # Spotify Connect + Avahi
βββ contrib/
β βββ setup.sh # Automated Pi setup script
β βββ nginx.conf # Nginx reverse proxy config
β βββ alarm.wav # Alarm sound file
βββ src/
β βββ app.py # Main voice assistant entry
β βββ backend.py # FastAPI application
β βββ common.py # Shared utilities
β βββ routes.py # Action router
β βββ actions.py # Legacy action functions (weather, Spotify, etc.)
β βββ audio_capture.py # Real-time audio capture for waveform visualization
β βββ audio_activity.py # Unified audio activity detection (Strategy pattern)
β βββ settings.json # Runtime configuration
β βββ requirements.txt # Python dependencies
β βββ agent/ # LangGraph agent
β β βββ __init__.py
β β βββ core.py # GPTHomeAgent class
β β βββ config.py # AgentConfig
β β βββ state.py # AgentState schema
β βββ display/ # Display management (HDMI/SPI/I2C)
β β βββ __init__.py
β β βββ base.py # BaseDisplay class and DisplayMode enum
β β βββ detection.py # Auto-detect connected displays
β β βββ factory.py # Display factory pattern
β β βββ manager.py # DisplayManager singleton (core logic)
β β βββ animations.py # Animation utilities
β β βββ integration.py # Tool context parsing for display
β β βββ palette.py # Color palette, ScrollingText, easing
β β βββ renderers.py # Shared rendering utilities
β β βββ spotify.py # Spotify now playing display loop
β β βββ weather.py # Weather rendering + data fetching
β β βββ multi.py # Multi-display support with mirroring
β β βββ modes/ # Display mode implementations
β β β βββ __init__.py
β β β βββ clock.py # Clock mode loop
β β β βββ gallery.py # Gallery mode loop
β β β βββ waveform.py # Waveform mode loop
β β β βββ weather.py # Weather mode loop
β β β βββ screensaver.py # Screensaver implementations
β β βββ drivers/ # Display driver implementations
β β βββ kmsdrm.py # KMS/DRM driver (HDMI displays)
β β βββ fbdev.py # Framebuffer driver (PiScreen displays)
β β βββ st7789.py # ST7789 SPI display driver
β β βββ i2c.py # SSD1306 I2C display
β βββ memory/ # Memory management
β β βββ __init__.py
β β βββ manager.py # MemoryManager
β β βββ store.py # Store factory
β βββ tools/ # LangChain tools
β β βββ __init__.py
β β βββ registry.py # Tool registry
β β βββ weather.py
β β βββ spotify.py
β β βββ lights.py
β β βββ calendar.py
β β βββ alarm.py
β βββ waveform/ # Waveform visualization system
β β βββ __init__.py
β β βββ interfaces.py # WaveformObserver, RenderStrategy ABCs
β β βββ mediator.py # WaveformMediator singleton
β β βββ strategies.py # VoiceGated, AlwaysOn, I2CDisplay strategies
β β βββ observers.py # FullDisplay and I2C observers
β βββ frontend/ # React web interface
β βββ package.json
β βββ src/
β β βββ App.tsx
β β βββ components/
β βββ build/ # Production build
βββ screenshots/ # Documentation images
GPT Home uses Docker Compose profiles to separate development and production services:
| Profile | Services Started | Use Case |
|---|---|---|
prod (default) |
db, nginx, backend, frontend, spotify
|
Production deployment |
dev |
db, nginx, backend, frontend-dev, spotify
|
Local development with React hot reload |
Note:
- The
backendservice runs a single uvicorn process that starts FastAPI (backend.py) which imports and runs the voice assistant (app.py) as a background task. This allows direct method calls between components instead of HTTP.- The
frontendservice (prod) serves pre-built React static files via nginx on port 80.- The
frontend-devservice (dev) runs React dev server on port 80 with a network aliasfrontendso nginx routing works unchanged.nginxroutes API requests tobackend:8000and static requests tofrontend:80(which resolves tofrontend-devin dev mode).
# Development
COMPOSE_PROFILES=dev docker compose up
# Production (default)
docker compose up -d
# Build specific profile
docker compose build
# View logs
docker compose logs -f backend# Create virtual environment
python3.11 -m venv venv
source venv/bin/activate # Linux/Mac
# or: venv\Scripts\activate # Windows
# Install dependencies
pip install -r src/requirements.txt
# Set environment variables
export LITELLM_API_KEY=sk-your-key
export DATABASE_URL=postgresql://gpt_home:gpt_home_secret@localhost:5432/gpt_home
# Start PostgreSQL (requires Docker or local install)
docker run -d --name gpt-home-db \
-e POSTGRES_DB=gpt_home \
-e POSTGRES_USER=gpt_home \
-e POSTGRES_PASSWORD=gpt_home_secret \
-p 5432:5432 \
pgvector/pgvector:0.8.1-pg18-trixie
# Run the backend
cd src
python -m uvicorn backend:app --host 0.0.0.0 --port 8000 --reloadcd src/frontend
# Install dependencies
npm install
# Development server (with hot reload)
npm start
# Production build
npm run build# This requires hardware (microphone/speaker) or mock inputs
cd src
python app.py# Build for local architecture
docker compose build
# Build for Raspberry Pi (ARM64)
docker buildx build --platform linux/arm64 -t judahpaul/gpt-home:latest .
# Build with no cache
docker compose build --no-cache# Start services
docker compose up -d
# View logs
docker compose logs -f backend
# Execute commands in container
docker compose exec backend bash
# Restart after code changes
docker compose restart backend
# Stop all services
docker compose down
# Stop and remove volumes
docker compose down -v# Shell access
docker compose exec backend bash
# Python REPL
docker compose exec backend /env/bin/python
# Check container status
docker compose ps
# View service logs
docker compose logs -f backend
docker compose logs -f spotify
# Test agent directly
docker compose exec backend /env/bin/python -c "
import asyncio
from routes import action_router
async def test():
response = await action_router('What time is it?')
print(response)
asyncio.run(test())
"# Backend tests (when available)
cd src
python -m pytest tests/ -v
# Frontend tests
cd src/frontend
npm testTest Weather Tool:
curl -X POST http://localhost:8000/weather \
-H "Content-Type: application/json" \
-d '{"location": "New York"}'Test Agent:
import asyncio
from routes import action_router
async def test_queries():
queries = [
"What's the weather?",
"Set an alarm for 5 minutes",
"What's on my calendar?",
]
for q in queries:
print(f"\nQuery: {q}")
response = await action_router(q)
print(f"Response: {response}")
asyncio.run(test_queries())- Follow PEP 8 guidelines
- Use type hints
- Document functions with docstrings
async def weather_tool(query: str) -> str:
"""Get current weather or forecast for a location.
Args:
query: Weather query like "weather in New York"
Returns:
Weather information string
"""
...- Use TypeScript strict mode
- Follow React best practices
- Use functional components with hooks
interface Props {
title: string;
onSubmit: (value: string) => void;
}
const Component: React.FC<Props> = ({ title, onSubmit }) => {
// ...
};# Python (if configured)
pip install flake8 black
black src/
flake8 src/
# Frontend
cd src/frontend
npm run lint- Create tool file:
# src/tools/my_tool.py
from langchain_core.tools import tool
from .env_utils import get_env
@tool
async def my_tool(query: str) -> str:
"""Description for the LLM.
Args:
query: What the user wants
Returns:
Result string
"""
api_key = get_env("MY_API_KEY")
if not api_key:
return "My tool is not configured."
# Implementation
return "Result"- Register in init.py:
# src/tools/__init__.py
from .my_tool import my_tool
__all__ = [..., "my_tool"]- Add to registry:
# src/tools/registry.py
from .my_tool import my_tool
registry.register(
my_tool,
ToolMetadata(
name="my_tool",
description="Does something",
category="productivity",
requires_api_key=True,
api_key_env_var="MY_API_KEY"
)
)# src/backend.py
@app.post("/my-endpoint")
async def my_endpoint(request: Request):
try:
data = await request.json()
# Process request
result = {"success": True, "data": "..."}
return JSONResponse(content=result)
except Exception as e:
return JSONResponse(
content={"error": str(e)},
status_code=500
)// src/frontend/src/components/MyComponent.tsx
import React, { useState, useEffect } from 'react';
interface MyComponentProps {
title: string;
}
export const MyComponent: React.FC<MyComponentProps> = ({ title }) => {
const [data, setData] = useState<string>('');
useEffect(() => {
// Fetch data
fetch('/my-endpoint', { method: 'POST' })
.then(res => res.json())
.then(setData);
}, []);
return (
<div className="my-component">
<h2>{title}</h2>
<p>{data}</p>
</div>
);
};GPT Home uses several design patterns:
The display manager uses a flag-based pattern for thread-to-async communication. This is the recommended approach when synchronous code (audio capture threads) needs to signal async code (display render loops).
# display/manager.py
import threading
_activity_pending = threading.Event()
def signal_activity() -> None:
"""Thread-safe. Called from any context to signal user activity."""
_activity_pending.set()
def check_and_clear_activity() -> bool:
"""Called from async loops to check and clear the flag."""
if _activity_pending.is_set():
_activity_pending.clear()
return True
return FalseUsage in render loops:
async def _screensaver_loop(self) -> None:
while self._screensaver_active:
if check_and_clear_activity():
await self._deactivate_screensaver()
break
# ... render frameWhy this pattern:
-
threading.Eventis thread-safe by design - No need to manage event loop references across threads
- Async loops check the flag on each iteration (fast polling)
- Avoids
call_soon_threadsafe()complexity and race conditions
_agent_instance: Optional[GPTHomeAgent] = None
async def _get_agent() -> GPTHomeAgent:
global _agent_instance
if _agent_instance is None:
_agent_instance = GPTHomeAgent(...)
return _agent_instanceasync def create_memory_store(database_url: Optional[str]) -> BaseStore:
if database_url:
return AsyncPostgresStore.from_conn_string(database_url)
return InMemoryStore()config = AgentConfig.builder() \
.with_model("gpt-4o") \
.with_temperature(0.5) \
.build()class ToolRegistry:
_tools: dict[str, Tool] = {}
def register(self, tool: Tool, metadata: ToolMetadata):
self._tools[tool.name] = tool- Fork the repository
- Create a feature branch:
git checkout -b feature/my-feature - Make changes and test
- Commit with descriptive messages:
git commit -m "Add my feature" - Push to your fork:
git push origin feature/my-feature - Open a Pull Request
Follow conventional commits:
feat: add new weather provider
fix: resolve memory leak in agent
docs: update API reference
refactor: simplify tool registry
test: add tests for calendar tool
- Code follows style guidelines
- Tests pass (if applicable)
- Documentation updated
- No sensitive data committed
- PR description explains changes
Agent not responding after wake word:
-
Enable DEBUG filter in Event Logs and check VAD output:
DEBUG:src.common:VAD: dB=-52.3, threshold=-55.0, duration=1.20s, peak=0.012, has_speech=FalseIf
has_speech=False, audio is being rejected before reaching STT. -
Check if speech chunks are being detected:
DEBUG:src.common:Speech at chunk 15, rms=1200, thresh=200, baseline=50If no speech chunks appear, the mic gain is too low.
-
Fix options:
- Settings β Audio β Mic Gain: Increase to 50-80%
- Settings β Audio β Voice Detection Threshold: Lower to -55 or -60 dB
- Settings β Audio β Pause Threshold: Increase to 1.5-2.0s if being cut off mid-sentence
-
Verify microphone is working:
arecord -l # List capture devices arecord -D plughw:2,0 -f cd test.wav # Test recording
-
Check debug audio file:
docker cp gpt-home-backend-1:/tmp/debug_audio.wav ./debug_audio.wav aplay debug_audio.wav # Listen to what was captured
Database connection issues:
# Test connection
docker compose exec db psql -U gpt_home -d gpt_home -c "SELECT 1"Tool not being selected:
# Check tool registration
from tools import get_all_tools
tools = get_all_tools()
print([t.name for t in tools])GPT Home uses Python's standard logging module. The root logger is configured at DEBUG level with two handlers β a FlushingFileHandler writing to events.log and a StreamHandler writing to stdout. Both streams show identical content. Third-party loggers (httpx, litellm, etc.) are suppressed to WARNING/ERROR.
import logging
logger = logging.getLogger(__name__)
logger.debug("Internal plumbing β device selection, per-cycle details")
logger.info("Significant events β voice detected, keyword matched, mode changes")
logger.warning("Recoverable issues β no display found, fallback used")
logger.error("Failures β device open failed, API error")Use logger.debug() for anything that fires per-cycle or is internal plumbing. Use logger.info() only for significant events the user would want to see in the Event Logs page. The web UI Event Logs page supports filtering by level (DEBUG, INFO, WARNING, ERROR).
Enable tracing to debug agent behavior:
# .env
LANGCHAIN_TRACING_V2=true
LANGCHAIN_API_KEY=lsv2_...
LANGCHAIN_PROJECT=gpt-home-devView traces at smith.langchain.com
- Update version in relevant files
- Update CHANGELOG (if exists)
- Create git tag:
git tag v1.2.3 - Push tag:
git push origin v1.2.3
The GitHub Actions workflow handles:
- Building multi-arch images (ARM64)
- Pushing to Docker Hub
- Tagging with version and
latest
# .github/workflows/workflow.yml
- name: Build and push
uses: docker/build-push-action@v5
with:
platforms: linux/arm64
push: true
tags: judahpaul/gpt-home:latest- Review the Architecture for system design
- Explore Tools Reference for tool examples
- Check Configuration for all options