Skip to content

04 configuration voice notifications

Doug Beard edited this page Aug 20, 2025 · 2 revisions

APM Voice Notifications Configuration Guide

This guide covers the configuration and customization of the voice notification system in the Agentic Persona Mapping (APM) framework.

Voice System Overview

APM's voice notification system provides:

  • Persona-Specific Voices: Different audio characteristics for each persona
  • Context-Aware Notifications: Intelligent audio feedback based on actions
  • Multi-Platform Support: Works across Windows, macOS, and Linux
  • Customizable Settings: Adjustable speed, volume, and voice selection
  • Integration Points: Seamless integration with all APM personas

Voice System Architecture

Voice Trigger → Voice Script → TTS Engine → Audio Output
     ↓              ↓            ↓           ↓
APM Command    speakPersona.sh  System TTS  Speakers/Headphones

Components

  1. Voice Scripts: Shell scripts for each persona (speakPersona.sh)
  2. TTS Providers: Text-to-speech providers (system, piper, elevenlabs, discord, none)
  3. Configuration: Settings for voice characteristics and behavior
  4. Integration: Hooks into APM persona activation and responses

Voice Script Locations

All voice scripts are located in: .apm/agents/voice/

Default Voice Scripts

  • speakOrchestrator.sh - Coherence Orchestrator voice
  • speakDeveloper.sh - Developer persona voice
  • speakArchitect.sh - System Architect voice
  • speakDesignArchitect.sh - Design Architect voice
  • speakAnalyst.sh - Business Analyst voice
  • speakQa.sh - Quality Assurance voice
  • speakPm.sh - Project Manager voice
  • speakPo.sh - Product Owner voice
  • speakSm.sh - Scrum Master voice

Basic Voice Configuration

Environment Variables

Configure voice behavior through environment variables:

# Enable/disable voice notifications
export VOICE_NOTIFICATIONS_ENABLED="true"

# Select TTS provider
export TTS_PROVIDER="system"  # system, piper, elevenlabs, discord, none

# Control speech characteristics
export TTS_VOICE_SPEED="1.0"
export TTS_VOICE_VOLUME="0.8"

TTS Provider Configuration

Configure TTS providers through the TTS manager:

# Configure TTS provider
.apm/agents/scripts/configure-tts.sh

# Check available providers
ls .apm/agents/scripts/tts-providers/

# Test TTS system
.apm/agents/scripts/tts-manager.sh test

# Set provider directly
export TTS_PROVIDER="system"

Available TTS Providers

system (Default)

  • Description: Uses system default TTS (macOS say, Linux espeak)
  • Configuration: No additional setup required
  • Usage: Automatically detected and configured

piper

  • Description: Piper neural text-to-speech
  • Configuration: Requires Piper installation
  • Usage: High-quality neural voices

elevenlabs

  • Description: ElevenLabs cloud TTS service
  • Configuration: Requires API key
  • Usage: Premium cloud-based voices

discord

  • Description: Discord TTS integration
  • Configuration: Requires Discord bot setup
  • Usage: TTS through Discord channels

none

  • Description: Disables TTS output
  • Configuration: No setup required
  • Usage: Silent operation

## Platform-Specific Configuration

### macOS Configuration

macOS uses the built-in `say` command:

```bash
# Test system voices
say -v "?" | head -20

# Configure for APM
export TTS_PROVIDER="system"
export TTS_VOICE_SPEED="1.0"

# Test voice
say "Coherence APM voice test"

Available macOS Voices

# List all available voices
say -v "?"

# Test specific voices
say -v "Alex" "Coherence Orchestrator activated"
say -v "Samantha" "Developer ready for coding tasks"
say -v "Tom" "System architecture review complete"

Linux Configuration

Linux uses espeak by default:

# Install espeak (if not present)
sudo apt-get install espeak espeak-data

# Configure for APM
export TTS_PROVIDER="system"
export TTS_VOICE_SPEED="1.0"

# Test espeak
espeak "Coherence APM voice test"

Piper Configuration (Advanced)

# Install Piper for higher quality voices
# Follow Piper installation instructions
# Configure provider
export TTS_PROVIDER="piper"

Windows/WSL Configuration

For Windows Subsystem for Linux:

# Use espeak in WSL
sudo apt-get install espeak
export TTS_PROVIDER="system"

# Test voice
espeak "Coherence APM voice test"

Advanced Voice Configuration

Custom Voice Scripts

Create advanced voice scripts with persona-specific characteristics:

#!/bin/bash
# Enhanced voice script template

PERSONA="{{PERSONA_NAME}}"
MESSAGE="$1"
CONFIG_FILE="{{APM_ROOT}}/config/voice-config.json"

# Load persona-specific settings
if [ -f "$CONFIG_FILE" ]; then
    VOICE=$(jq -r ".persona_voices.$PERSONA.voice // \"default\"" "$CONFIG_FILE")
    RATE=$(jq -r ".persona_voices.$PERSONA.rate // 1.0" "$CONFIG_FILE")
    PITCH=$(jq -r ".persona_voices.$PERSONA.pitch // 0" "$CONFIG_FILE")
    VOLUME=$(jq -r ".persona_voices.$PERSONA.volume // 0.8" "$CONFIG_FILE")
else
    # Fallback to environment variables
    VOICE="${VOICE_PERSONA_${PERSONA^^}:-default}"
    RATE="${VOICE_SPEED:-1.0}"
    PITCH="${VOICE_PITCH:-0}"
    VOLUME="${VOICE_VOLUME:-0.8}"
fi

# Voice notification with persona characteristics
speak_with_characteristics() {
    local message="$1"
    
    case "${VOICE_ENGINE:-system}" in
        "system")
            if command -v say >/dev/null 2>&1; then
                # macOS with persona-specific voice
                say -v "$VOICE" -r $(echo "$RATE * 175" | bc) "$message"
            elif command -v espeak >/dev/null 2>&1; then
                # Linux espeak with persona characteristics
                espeak -v "$VOICE" -s $(echo "$RATE * 150" | bc) \
                       -a $(echo "$VOLUME * 100" | bc) \
                       -p $(echo "$PITCH + 50" | bc) "$message"
            fi
            ;;
        "espeak")
            espeak -v "${VOICE:-en}" -s $(echo "$RATE * 150" | bc) \
                   -a $(echo "$VOLUME * 100" | bc) \
                   -p $(echo "$PITCH + 50" | bc) "$message"
            ;;
        "festival")
            echo "$message" | festival --tts
            ;;
        "powershell")
            powershell.exe -Command "
                Add-Type -AssemblyName System.Speech
                \$synth = New-Object System.Speech.Synthesis.SpeechSynthesizer
                \$synth.Rate = $([int]($RATE * 0))
                \$synth.Volume = $([int]($VOLUME * 100))
                \$synth.Speak('$message')
            "
            ;;
    esac
}

# Context-aware messaging
add_context() {
    local base_message="$1"
    local timestamp=$(date "+%H:%M")
    
    # Add persona prefix and context
    echo "[$timestamp] [AP $PERSONA] $base_message"
}

# Main execution
if [ "${VOICE_NOTIFICATIONS_ENABLED:-true}" = "true" ]; then
    CONTEXTUAL_MESSAGE=$(add_context "$MESSAGE")
    speak_with_characteristics "$CONTEXTUAL_MESSAGE"
fi

# Always log the message
echo "[AP $PERSONA] $MESSAGE" >> {{APM_ROOT}}/logs/voice-notifications.log

Dynamic Voice Selection

Implement dynamic voice selection based on context:

#!/bin/bash
# Dynamic voice selection based on context

select_voice_for_context() {
    local context="$1"
    local persona="$2"
    
    case "$context" in
        "activation")
            # Use authoritative voice for activation
            echo "Alex"
            ;;
        "error")
            # Use clear, distinct voice for errors
            echo "Daniel"
            ;;
        "completion")
            # Use pleasant voice for completion
            echo "Samantha"
            ;;
        "handoff")
            # Use professional voice for handoffs
            echo "Victoria"
            ;;
        *)
            # Use persona-specific default
            jq -r ".persona_voices.$persona.voice // \"default\"" "$CONFIG_FILE"
            ;;
    esac
}

# Usage in voice scripts
CONTEXT="${2:-normal}"
SELECTED_VOICE=$(select_voice_for_context "$CONTEXT" "$PERSONA")

Notification Types and Triggers

Activation Notifications

Triggered when personas are activated:

# Example activation messages
"Coherence Orchestrator activated. Loading configuration and persona network."
"Developer ready. Current project: {{PROJECT_NAME}}. Sprint: {{SPRINT_NAME}}."
"System Architect online. Architecture review mode active."

Progress Notifications

Triggered during task execution:

# Example progress messages
"Task 1 of 4 completed. Proceeding with implementation."
"Code review in progress. 3 files reviewed, 2 remaining."
"Deployment pipeline initiated. Monitoring progress."

Completion Notifications

Triggered when tasks or sessions complete:

# Example completion messages
"Sprint planning complete. 12 stories groomed and prioritized."
"Code implementation finished. All tests passing."
"Architecture review complete. Documentation updated."

Error Notifications

Triggered when errors occur:

# Example error messages
"Build failed. Compilation errors detected in user authentication module."
"Test suite failed. 3 unit tests require attention."
"Deployment blocked. Infrastructure health check failed."

Handoff Notifications

Triggered during persona transitions:

# Example handoff messages
"Handoff to Developer initiated. Context preserved. Session transferred."
"Transitioning to QA Agent. Test requirements documented."
"Project Manager taking over. Sprint metrics updated."

Voice Customization Examples

Creating Persona-Specific Voice Profiles

Orchestrator Voice Profile

{
  "orchestrator": {
    "voice": "Alex",
    "rate": 1.0,
    "pitch": 0,
    "volume": 0.9,
    "characteristics": "authoritative, clear",
    "message_templates": {
      "activation": "Coherence Orchestrator online. Coordinating {{ACTIVE_AGENTS}} agents. System ready.",
      "delegation": "Delegating {{TASK_TYPE}} to {{TARGET_PERSONA}}. Context transferred.",
      "completion": "Coordination complete. {{TASKS_COMPLETED}} tasks finished successfully.",
      "error": "System issue detected. {{ERROR_TYPE}}. Initiating recovery procedures."
    }
  }
}

Developer Voice Profile

{
  "developer": {
    "voice": "Samantha",
    "rate": 1.1,
    "pitch": 5,
    "volume": 0.7,
    "characteristics": "technical, precise",
    "message_templates": {
      "activation": "Developer active. Loading {{PROJECT_NAME}} codebase. Ready for development tasks.",
      "progress": "{{FEATURE_NAME}} implementation {{PROGRESS_PERCENT}} complete. {{TESTS_STATUS}}.",
      "completion": "Feature implementation complete. {{LINES_OF_CODE}} lines added. All tests passing.",
      "error": "Development issue encountered. {{ERROR_DESCRIPTION}}. Debugging initiated."
    }
  }
}

Context-Aware Voice Messages

Create dynamic messages based on current context:

#!/bin/bash
# Context-aware voice messaging

generate_contextual_message() {
    local persona="$1"
    local action="$2"
    local context="$3"
    
    # Load project context
    local project_name=$(jq -r '.project.name // "Unknown"' {{APM_ROOT}}/project-context.json)
    local sprint_name=$(jq -r '.sprint.current // "Unknown"' {{APM_ROOT}}/project-context.json)
    local team_size=$(jq -r '.team.size // 0' {{APM_ROOT}}/project-context.json)
    
    # Generate contextual message
    case "$action" in
        "activation")
            echo "AP $persona activated for $project_name. Sprint: $sprint_name. Team size: $team_size."
            ;;
        "handoff")
            echo "Handoff from $persona to $context. Project context preserved."
            ;;
        "completion")
            echo "$persona task complete. $context. Updating project status."
            ;;
    esac
}

Voice Testing and Validation

Testing Framework

Create comprehensive voice testing:

#!/bin/bash
# Voice system testing framework

test_voice_system() {
    echo "=== APM Voice System Testing ==="
    
    # Test 1: Basic voice functionality
    echo "Testing basic voice functionality..."
    if {{APM_ROOT}}/agents/voice/speakOrchestrator.sh "Voice system test" >/dev/null 2>&1; then
        echo "✅ Basic voice functionality working"
    else
        echo "❌ Basic voice functionality failed"
        return 1
    fi
    
    # Test 2: All persona voices
    echo "Testing all persona voices..."
    local personas=("Orchestrator" "Developer" "Architect" "Analyst" "Qa" "Pm" "Po" "Sm")
    for persona in "${personas[@]}"; do
        if {{APM_ROOT}}/agents/voice/speak${persona}.sh "Testing $persona voice" >/dev/null 2>&1; then
            echo "✅ $persona voice working"
        else
            echo "❌ $persona voice failed"
        fi
    done
    
    # Test 3: Voice configuration
    echo "Testing voice configuration..."
    local config_file="{{APM_ROOT}}/config/voice-config.json"
    if [ -f "$config_file" ] && jq empty "$config_file" 2>/dev/null; then
        echo "✅ Voice configuration valid"
    else
        echo "❌ Voice configuration invalid or missing"
    fi
    
    # Test 4: Platform TTS engines
    echo "Testing platform TTS engines..."
    if command -v say >/dev/null 2>&1; then
        echo "✅ macOS say command available"
    elif command -v espeak >/dev/null 2>&1; then
        echo "✅ Linux espeak available"
    elif command -v festival >/dev/null 2>&1; then
        echo "✅ Linux festival available"
    else
        echo "❌ No TTS engine detected"
    fi
    
    echo "Voice system testing complete."
}

# Performance testing
test_voice_performance() {
    echo "=== Voice Performance Testing ==="
    
    local start_time=$(date +%s.%N)
    {{APM_ROOT}}/agents/voice/speakOrchestrator.sh "Performance test message"
    local end_time=$(date +%s.%N)
    
    local duration=$(echo "$end_time - $start_time" | bc)
    echo "Voice notification latency: ${duration}s"
    
    if (( $(echo "$duration < 2.0" | bc -l) )); then
        echo "✅ Voice performance acceptable"
    else
        echo "⚠️ Voice performance may be slow"
    fi
}

# Run tests
test_voice_system
test_voice_performance

Voice Quality Assessment

#!/bin/bash
# Voice quality assessment

assess_voice_quality() {
    local persona="$1"
    local test_messages=(
        "Simple test message"
        "Complex technical terminology: authentication, containerization, orchestration"
        "Numbers and metrics: 42 tests passed, 3.7 seconds execution time, 97% success rate"
        "Mixed content: The API returned HTTP 200 with JSON payload containing user data"
    )
    
    echo "Assessing voice quality for $persona..."
    
    for i in "${!test_messages[@]}"; do
        echo "Test $((i+1)): ${test_messages[$i]}"
        {{APM_ROOT}}/agents/voice/speak${persona^}.sh "${test_messages[$i]}"
        echo "Rate clarity (1-5): "
        read -r clarity_rating
        echo "Rate naturalness (1-5): "
        read -r naturalness_rating
        
        echo "Test $((i+1)) - Clarity: $clarity_rating, Naturalness: $naturalness_rating"
    done
}

# Usage
assess_voice_quality "Orchestrator"

Troubleshooting Voice Issues

Common Problems and Solutions

Issue: No audio output

Symptoms: Voice scripts execute but no sound Solutions:

# Test TTS provider directly
say "test"  # macOS
espeak "test"  # Linux

# Check TTS provider setting
echo $TTS_PROVIDER

# Test TTS manager
.apm/agents/scripts/tts-manager.sh test

# Verify voice notifications enabled
echo $VOICE_NOTIFICATIONS_ENABLED

Issue: Wrong voice or characteristics

Symptoms: Voice doesn't match persona configuration Solutions:

# Check voice configuration
cat {{APM_ROOT}}/config/voice-config.json

# Validate JSON configuration
jq empty {{APM_ROOT}}/config/voice-config.json

# Test voice selection
say -v "Alex" "Test message"  # macOS
espeak -v "en" "Test message"  # Linux

# Check environment variables
env | grep VOICE

Issue: Voice script permissions

Symptoms: "Permission denied" errors Solutions:

# Fix script permissions
chmod +x {{APM_ROOT}}/agents/voice/*.sh

# Check file ownership
ls -la {{APM_ROOT}}/agents/voice/

# Verify script syntax
bash -n {{APM_ROOT}}/agents/voice/speakOrchestrator.sh

Issue: Voice notifications too slow

Symptoms: Long delays before audio output Solutions:

# Optimize voice speed
export VOICE_SPEED="1.5"

# Use faster TTS engine
export VOICE_ENGINE="espeak"

# Disable complex message processing
export VOICE_SIMPLE_MESSAGES="true"

# Check system resources
top
ps aux | grep -E "(say|espeak|festival)"

Diagnostic Tools

Voice System Diagnostics

#!/bin/bash
# Voice system diagnostic tool

diagnose_voice_system() {
    echo "=== APM Voice System Diagnostics ==="
    
    echo "1. Environment Configuration:"
    echo "   VOICE_NOTIFICATIONS_ENABLED: ${VOICE_NOTIFICATIONS_ENABLED:-not set}"
    echo "   VOICE_ENGINE: ${VOICE_ENGINE:-not set}"
    echo "   VOICE_SPEED: ${VOICE_SPEED:-not set}"
    echo "   VOICE_VOLUME: ${VOICE_VOLUME:-not set}"
    
    echo -e "\n2. Platform Detection:"
    if [[ "$OSTYPE" == "darwin"* ]]; then
        echo "   Platform: macOS"
        echo "   Available voices: $(say -v "?" | wc -l) voices detected"
        echo "   Default voice: $(defaults read com.apple.speech.voice.prefs SelectedVoiceName 2>/dev/null || echo "System default")"
    elif [[ "$OSTYPE" == "linux"* ]]; then
        echo "   Platform: Linux"
        if command -v espeak >/dev/null 2>&1; then
            echo "   espeak: Available"
        fi
        if command -v festival >/dev/null 2>&1; then
            echo "   festival: Available"
        fi
    fi
    
    echo -e "\n3. Voice Scripts Status:"
    local voice_dir="{{APM_ROOT}}/agents/voice"
    if [ -d "$voice_dir" ]; then
        for script in "$voice_dir"/speak*.sh; do
            if [ -x "$script" ]; then
                echo "   ✅ $(basename "$script"): Executable"
            else
                echo "   ❌ $(basename "$script"): Not executable"
            fi
        done
    else
        echo "   ❌ Voice scripts directory not found"
    fi
    
    echo -e "\n4. Configuration Files:"
    local config_file="{{APM_ROOT}}/config/voice-config.json"
    if [ -f "$config_file" ]; then
        if jq empty "$config_file" 2>/dev/null; then
            echo "   ✅ voice-config.json: Valid"
        else
            echo "   ❌ voice-config.json: Invalid JSON"
        fi
    else
        echo "   ⚠️ voice-config.json: Not found (using defaults)"
    fi
    
    echo -e "\n5. Audio System Status:"
    if command -v pactl >/dev/null 2>&1; then
        if pactl info >/dev/null 2>&1; then
            echo "   ✅ PulseAudio: Active"
        else
            echo "   ❌ PulseAudio: Not responding"
        fi
    elif command -v osascript >/dev/null 2>&1; then
        local volume=$(osascript -e "output volume of (get volume settings)" 2>/dev/null)
        echo "   ✅ macOS Audio: Volume at $volume%"
    fi
    
    echo -e "\nDiagnostics complete."
}

diagnose_voice_system

Integration with APM Workflows

Session-Aware Voice Notifications

Integrate voice notifications with session management:

# Session-aware voice script
session_aware_speak() {
    local message="$1"
    local persona="$2"
    local session_file="{{APM_ROOT}}/session_notes/current_session.md"
    
    # Add session context to message
    if [ -f "$session_file" ]; then
        local session_title=$(grep "^# Session:" "$session_file" | cut -d' ' -f3-)
        message="[$session_title] $message"
    fi
    
    # Speak with session context
    speak_with_characteristics "$message"
    
    # Log to session notes
    if [ -f "$session_file" ]; then
        echo "[$(date '+%H:%M')] 🔊 $message" >> "$session_file"
    fi
}

Parallel Execution Voice Coordination

Coordinate voice notifications during parallel execution:

# Parallel execution voice coordinator
coordinate_parallel_voices() {
    local agents=("$@")
    local coordination_file="/tmp/apm_voice_coordination"
    
    # Create coordination file
    echo "parallel_execution_active" > "$coordination_file"
    
    # Stagger voice notifications to avoid overlap
    for i in "${!agents[@]}"; do
        sleep $(echo "scale=1; $i * 0.5" | bc)
        {{APM_ROOT}}/agents/voice/speak${agents[$i]^}.sh "Agent $((i+1)) of ${#agents[@]} active"
    done
    
    # Cleanup
    rm -f "$coordination_file"
}

Next Steps: After configuring voice notifications, review Path Configuration to complete your APM setup, or return to the Configuration Overview for other configuration options.

Clone this wiki locally