Skip to content

Development Guide

Judah Paul edited this page Mar 2, 2026 · 13 revisions

πŸ’» Development Guide

This page covers local development setup, Docker builds, testing, and contributing to GPT Home.

Development Environment

Prerequisites

Tool Version Purpose
Python 3.11 Backend runtime
Node.js 18.x Frontend build
Docker 24.x+ Containerization
Docker Compose 2.x Service orchestration
Git 2.x+ Version control

Quick Start (Docker Dev Profile)

The easiest way to develop is using the Docker dev profile with hot reload:

# Clone repository
git clone https://github.com/judahpaul16/gpt-home.git
cd gpt-home

# Copy environment template
cp .env.example .env

# Edit .env with your API key
nano .env

# Start development environment
COMPOSE_PROFILES=dev docker compose up

# Access:
# - Web Interface: http://localhost (nginx routes to web-dev)
# - FastAPI Backend: http://localhost:8000 (direct access)
# - API Docs: http://localhost:8000/docs

Changes to files in ./src/frontend are automatically hot-reloaded via volume mount.

Production Profile

For testing the production build locally:

# Production is default (or explicitly set)
docker compose up -d
# or: COMPOSE_PROFILES=prod docker compose up -d

docker compose logs -f

Project Structure

gpt-home/
β”œβ”€β”€ .env                    # Environment configuration (create from .env.example)
β”œβ”€β”€ .env.example            # Environment template
β”œβ”€β”€ docker-compose.yml      # Service definitions with profiles
β”œβ”€β”€ compose/                # Dockerfiles for each service
β”‚   β”œβ”€β”€ app/
β”‚   β”‚   └── Dockerfile      # Voice assistant + FastAPI backend
β”‚   β”œβ”€β”€ web/
β”‚   β”‚   β”œβ”€β”€ Dockerfile      # Production: multi-stage React build β†’ nginx
β”‚   β”‚   └── Dockerfile.dev  # Development: React dev server (hot reload)
β”‚   └── spotify/
β”‚       └── Dockerfile      # Spotify Connect + Avahi
β”œβ”€β”€ contrib/
β”‚   β”œβ”€β”€ setup.sh            # Automated Pi setup script
β”‚   β”œβ”€β”€ nginx.conf          # Nginx reverse proxy config
β”‚   └── alarm.wav           # Alarm sound file
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ app.py              # Main voice assistant entry
β”‚   β”œβ”€β”€ backend.py          # FastAPI application
β”‚   β”œβ”€β”€ common.py           # Shared utilities
β”‚   β”œβ”€β”€ routes.py           # Action router
β”‚   β”œβ”€β”€ actions.py          # Legacy action functions (weather, Spotify, etc.)
β”‚   β”œβ”€β”€ audio_capture.py    # Real-time audio capture for waveform visualization
β”‚   β”œβ”€β”€ audio_activity.py   # Unified audio activity detection (Strategy pattern)
β”‚   β”œβ”€β”€ settings.json       # Runtime configuration
β”‚   β”œβ”€β”€ requirements.txt    # Python dependencies
β”‚   β”œβ”€β”€ agent/              # LangGraph agent
β”‚   β”‚   β”œβ”€β”€ __init__.py
β”‚   β”‚   β”œβ”€β”€ core.py         # GPTHomeAgent class
β”‚   β”‚   β”œβ”€β”€ config.py       # AgentConfig
β”‚   β”‚   └── state.py        # AgentState schema
β”‚   β”œβ”€β”€ display/            # Display management (HDMI/SPI/I2C)
β”‚   β”‚   β”œβ”€β”€ __init__.py
β”‚   β”‚   β”œβ”€β”€ base.py         # BaseDisplay class and DisplayMode enum
β”‚   β”‚   β”œβ”€β”€ detection.py    # Auto-detect connected displays
β”‚   β”‚   β”œβ”€β”€ factory.py      # Display factory pattern
β”‚   β”‚   β”œβ”€β”€ manager.py      # DisplayManager singleton (core logic)
β”‚   β”‚   β”œβ”€β”€ animations.py   # Animation utilities
β”‚   β”‚   β”œβ”€β”€ integration.py  # Tool context parsing for display
β”‚   β”‚   β”œβ”€β”€ palette.py      # Color palette, ScrollingText, easing
β”‚   β”‚   β”œβ”€β”€ renderers.py    # Shared rendering utilities
β”‚   β”‚   β”œβ”€β”€ spotify.py      # Spotify now playing display loop
β”‚   β”‚   β”œβ”€β”€ weather.py      # Weather rendering + data fetching
β”‚   β”‚   β”œβ”€β”€ multi.py        # Multi-display support with mirroring
β”‚   β”‚   β”œβ”€β”€ modes/          # Display mode implementations
β”‚   β”‚   β”‚   β”œβ”€β”€ __init__.py
β”‚   β”‚   β”‚   β”œβ”€β”€ clock.py    # Clock mode loop
β”‚   β”‚   β”‚   β”œβ”€β”€ gallery.py  # Gallery mode loop
β”‚   β”‚   β”‚   β”œβ”€β”€ waveform.py # Waveform mode loop
β”‚   β”‚   β”‚   β”œβ”€β”€ weather.py  # Weather mode loop
β”‚   β”‚   β”‚   └── screensaver.py # Screensaver implementations
β”‚   β”‚   └── drivers/        # Display driver implementations
β”‚   β”‚       β”œβ”€β”€ kmsdrm.py   # KMS/DRM driver (HDMI displays)
β”‚   β”‚       β”œβ”€β”€ fbdev.py    # Framebuffer driver (PiScreen displays)
β”‚   β”‚       β”œβ”€β”€ st7789.py   # ST7789 SPI display driver
β”‚   β”‚       └── i2c.py      # SSD1306 I2C display
β”‚   β”œβ”€β”€ memory/             # Memory management
β”‚   β”‚   β”œβ”€β”€ __init__.py
β”‚   β”‚   β”œβ”€β”€ manager.py      # MemoryManager
β”‚   β”‚   └── store.py        # Store factory
β”‚   β”œβ”€β”€ tools/              # LangChain tools
β”‚   β”‚   β”œβ”€β”€ __init__.py
β”‚   β”‚   β”œβ”€β”€ registry.py     # Tool registry
β”‚   β”‚   β”œβ”€β”€ weather.py
β”‚   β”‚   β”œβ”€β”€ spotify.py
β”‚   β”‚   β”œβ”€β”€ lights.py
β”‚   β”‚   β”œβ”€β”€ calendar.py
β”‚   β”‚   └── alarm.py
β”‚   β”œβ”€β”€ waveform/           # Waveform visualization system
β”‚   β”‚   β”œβ”€β”€ __init__.py
β”‚   β”‚   β”œβ”€β”€ interfaces.py   # WaveformObserver, RenderStrategy ABCs
β”‚   β”‚   β”œβ”€β”€ mediator.py     # WaveformMediator singleton
β”‚   β”‚   β”œβ”€β”€ strategies.py   # VoiceGated, AlwaysOn, I2CDisplay strategies
β”‚   β”‚   └── observers.py    # FullDisplay and I2C observers
β”‚   └── frontend/           # React web interface
β”‚       β”œβ”€β”€ package.json
β”‚       β”œβ”€β”€ src/
β”‚       β”‚   β”œβ”€β”€ App.tsx
β”‚       β”‚   └── components/
β”‚       └── build/          # Production build
└── screenshots/            # Documentation images

Docker Compose Profiles

GPT Home uses Docker Compose profiles to separate development and production services:

Profile Services Started Use Case
prod (default) db, nginx, backend, frontend, spotify Production deployment
dev db, nginx, backend, frontend-dev, spotify Local development with React hot reload

Note:

  • The backend service runs a single uvicorn process that starts FastAPI (backend.py) which imports and runs the voice assistant (app.py) as a background task. This allows direct method calls between components instead of HTTP.
  • The frontend service (prod) serves pre-built React static files via nginx on port 80.
  • The frontend-dev service (dev) runs React dev server on port 80 with a network alias frontend so nginx routing works unchanged.
  • nginx routes API requests to backend:8000 and static requests to frontend:80 (which resolves to frontend-dev in dev mode).
# Development
COMPOSE_PROFILES=dev docker compose up

# Production (default)
docker compose up -d

# Build specific profile
docker compose build

# View logs
docker compose logs -f backend

Local Development (Without Docker)

Backend Setup

# Create virtual environment
python3.11 -m venv venv
source venv/bin/activate  # Linux/Mac
# or: venv\Scripts\activate  # Windows

# Install dependencies
pip install -r src/requirements.txt

# Set environment variables
export LITELLM_API_KEY=sk-your-key
export DATABASE_URL=postgresql://gpt_home:gpt_home_secret@localhost:5432/gpt_home

# Start PostgreSQL (requires Docker or local install)
docker run -d --name gpt-home-db \
  -e POSTGRES_DB=gpt_home \
  -e POSTGRES_USER=gpt_home \
  -e POSTGRES_PASSWORD=gpt_home_secret \
  -p 5432:5432 \
  pgvector/pgvector:0.8.1-pg18-trixie

# Run the backend
cd src
python -m uvicorn backend:app --host 0.0.0.0 --port 8000 --reload

Frontend Setup

cd src/frontend

# Install dependencies
npm install

# Development server (with hot reload)
npm start

# Production build
npm run build

Running the Voice Assistant

# This requires hardware (microphone/speaker) or mock inputs
cd src
python app.py

Docker Development

Building the Image

# Build for local architecture
docker compose build

# Build for Raspberry Pi (ARM64)
docker buildx build --platform linux/arm64 -t judahpaul/gpt-home:latest .

# Build with no cache
docker compose build --no-cache

Development with Docker

# Start services
docker compose up -d

# View logs
docker compose logs -f backend

# Execute commands in container
docker compose exec backend bash

# Restart after code changes
docker compose restart backend

# Stop all services
docker compose down

# Stop and remove volumes
docker compose down -v

Debugging in Container

# Shell access
docker compose exec backend bash

# Python REPL
docker compose exec backend /env/bin/python

# Check container status
docker compose ps

# View service logs
docker compose logs -f backend
docker compose logs -f spotify

# Test agent directly
docker compose exec backend /env/bin/python -c "
import asyncio
from routes import action_router

async def test():
    response = await action_router('What time is it?')
    print(response)

asyncio.run(test())
"

Testing

Running Tests

# Backend tests (when available)
cd src
python -m pytest tests/ -v

# Frontend tests
cd src/frontend
npm test

Manual Testing

Test Weather Tool:

curl -X POST http://localhost:8000/weather \
  -H "Content-Type: application/json" \
  -d '{"location": "New York"}'

Test Agent:

import asyncio
from routes import action_router

async def test_queries():
    queries = [
        "What's the weather?",
        "Set an alarm for 5 minutes",
        "What's on my calendar?",
    ]
    for q in queries:
        print(f"\nQuery: {q}")
        response = await action_router(q)
        print(f"Response: {response}")

asyncio.run(test_queries())

Code Style

Python

  • Follow PEP 8 guidelines
  • Use type hints
  • Document functions with docstrings
async def weather_tool(query: str) -> str:
    """Get current weather or forecast for a location.
    
    Args:
        query: Weather query like "weather in New York"
    
    Returns:
        Weather information string
    """
    ...

TypeScript (Frontend)

  • Use TypeScript strict mode
  • Follow React best practices
  • Use functional components with hooks
interface Props {
  title: string;
  onSubmit: (value: string) => void;
}

const Component: React.FC<Props> = ({ title, onSubmit }) => {
  // ...
};

Linting

# Python (if configured)
pip install flake8 black
black src/
flake8 src/

# Frontend
cd src/frontend
npm run lint

Adding New Features

Adding a New Tool

  1. Create tool file:
# src/tools/my_tool.py
from langchain_core.tools import tool
from .env_utils import get_env

@tool
async def my_tool(query: str) -> str:
    """Description for the LLM.
    
    Args:
        query: What the user wants
    
    Returns:
        Result string
    """
    api_key = get_env("MY_API_KEY")
    if not api_key:
        return "My tool is not configured."
    
    # Implementation
    return "Result"
  1. Register in init.py:
# src/tools/__init__.py
from .my_tool import my_tool

__all__ = [..., "my_tool"]
  1. Add to registry:
# src/tools/registry.py
from .my_tool import my_tool

registry.register(
    my_tool,
    ToolMetadata(
        name="my_tool",
        description="Does something",
        category="productivity",
        requires_api_key=True,
        api_key_env_var="MY_API_KEY"
    )
)

Adding a New API Endpoint

# src/backend.py

@app.post("/my-endpoint")
async def my_endpoint(request: Request):
    try:
        data = await request.json()
        # Process request
        result = {"success": True, "data": "..."}
        return JSONResponse(content=result)
    except Exception as e:
        return JSONResponse(
            content={"error": str(e)},
            status_code=500
        )

Adding a Frontend Component

// src/frontend/src/components/MyComponent.tsx
import React, { useState, useEffect } from 'react';

interface MyComponentProps {
  title: string;
}

export const MyComponent: React.FC<MyComponentProps> = ({ title }) => {
  const [data, setData] = useState<string>('');

  useEffect(() => {
    // Fetch data
    fetch('/my-endpoint', { method: 'POST' })
      .then(res => res.json())
      .then(setData);
  }, []);

  return (
    <div className="my-component">
      <h2>{title}</h2>
      <p>{data}</p>
    </div>
  );
};

Architecture Patterns

GPT Home uses several design patterns:

Thread-Safe Activity Signaling

The display manager uses a flag-based pattern for thread-to-async communication. This is the recommended approach when synchronous code (audio capture threads) needs to signal async code (display render loops).

# display/manager.py
import threading

_activity_pending = threading.Event()

def signal_activity() -> None:
    """Thread-safe. Called from any context to signal user activity."""
    _activity_pending.set()

def check_and_clear_activity() -> bool:
    """Called from async loops to check and clear the flag."""
    if _activity_pending.is_set():
        _activity_pending.clear()
        return True
    return False

Usage in render loops:

async def _screensaver_loop(self) -> None:
    while self._screensaver_active:
        if check_and_clear_activity():
            await self._deactivate_screensaver()
            break
        # ... render frame

Why this pattern:

  • threading.Event is thread-safe by design
  • No need to manage event loop references across threads
  • Async loops check the flag on each iteration (fast polling)
  • Avoids call_soon_threadsafe() complexity and race conditions

Singleton (Agent Instance)

_agent_instance: Optional[GPTHomeAgent] = None

async def _get_agent() -> GPTHomeAgent:
    global _agent_instance
    if _agent_instance is None:
        _agent_instance = GPTHomeAgent(...)
    return _agent_instance

Factory (Store Creation)

async def create_memory_store(database_url: Optional[str]) -> BaseStore:
    if database_url:
        return AsyncPostgresStore.from_conn_string(database_url)
    return InMemoryStore()

Builder (Configuration)

config = AgentConfig.builder() \
    .with_model("gpt-4o") \
    .with_temperature(0.5) \
    .build()

Registry (Tools)

class ToolRegistry:
    _tools: dict[str, Tool] = {}
    
    def register(self, tool: Tool, metadata: ToolMetadata):
        self._tools[tool.name] = tool

Contributing

Workflow

  1. Fork the repository
  2. Create a feature branch: git checkout -b feature/my-feature
  3. Make changes and test
  4. Commit with descriptive messages: git commit -m "Add my feature"
  5. Push to your fork: git push origin feature/my-feature
  6. Open a Pull Request

Commit Messages

Follow conventional commits:

feat: add new weather provider
fix: resolve memory leak in agent
docs: update API reference
refactor: simplify tool registry
test: add tests for calendar tool

Pull Request Checklist

  • Code follows style guidelines
  • Tests pass (if applicable)
  • Documentation updated
  • No sensitive data committed
  • PR description explains changes

Debugging Tips

Common Issues

Agent not responding after wake word:

  1. Enable DEBUG filter in Event Logs and check VAD output:

    DEBUG:src.common:VAD: dB=-52.3, threshold=-55.0, duration=1.20s, peak=0.012, has_speech=False
    

    If has_speech=False, audio is being rejected before reaching STT.

  2. Check if speech chunks are being detected:

    DEBUG:src.common:Speech at chunk 15, rms=1200, thresh=200, baseline=50
    

    If no speech chunks appear, the mic gain is too low.

  3. Fix options:

    • Settings β†’ Audio β†’ Mic Gain: Increase to 50-80%
    • Settings β†’ Audio β†’ Voice Detection Threshold: Lower to -55 or -60 dB
    • Settings β†’ Audio β†’ Pause Threshold: Increase to 1.5-2.0s if being cut off mid-sentence
  4. Verify microphone is working:

    arecord -l                              # List capture devices
    arecord -D plughw:2,0 -f cd test.wav   # Test recording
  5. Check debug audio file:

    docker cp gpt-home-backend-1:/tmp/debug_audio.wav ./debug_audio.wav
    aplay debug_audio.wav  # Listen to what was captured

Database connection issues:

# Test connection
docker compose exec db psql -U gpt_home -d gpt_home -c "SELECT 1"

Tool not being selected:

# Check tool registration
from tools import get_all_tools
tools = get_all_tools()
print([t.name for t in tools])

Logging

GPT Home uses Python's standard logging module. The root logger is configured at DEBUG level with two handlers β€” a FlushingFileHandler writing to events.log and a StreamHandler writing to stdout. Both streams show identical content. Third-party loggers (httpx, litellm, etc.) are suppressed to WARNING/ERROR.

import logging
logger = logging.getLogger(__name__)

logger.debug("Internal plumbing β€” device selection, per-cycle details")
logger.info("Significant events β€” voice detected, keyword matched, mode changes")
logger.warning("Recoverable issues β€” no display found, fallback used")
logger.error("Failures β€” device open failed, API error")

Use logger.debug() for anything that fires per-cycle or is internal plumbing. Use logger.info() only for significant events the user would want to see in the Event Logs page. The web UI Event Logs page supports filtering by level (DEBUG, INFO, WARNING, ERROR).

Tracing with LangSmith

Enable tracing to debug agent behavior:

# .env
LANGCHAIN_TRACING_V2=true
LANGCHAIN_API_KEY=lsv2_...
LANGCHAIN_PROJECT=gpt-home-dev

View traces at smith.langchain.com


Release Process

Version Bumping

  1. Update version in relevant files
  2. Update CHANGELOG (if exists)
  3. Create git tag: git tag v1.2.3
  4. Push tag: git push origin v1.2.3

Docker Hub Release

The GitHub Actions workflow handles:

  • Building multi-arch images (ARM64)
  • Pushing to Docker Hub
  • Tagging with version and latest
# .github/workflows/workflow.yml
- name: Build and push
  uses: docker/build-push-action@v5
  with:
    platforms: linux/arm64
    push: true
    tags: judahpaul/gpt-home:latest

Resources

Documentation

Community


Next Steps

Clone this wiki locally