Skip to content

v0.3.0 — bidirectional iMessage middle-man

Choose a tag to compare

@heyfinal heyfinal released this 16 Jun 20:53
· 5 commits to main since this release

callsign v0.3 — REWORK after GPT-5.4 + DeepSeek V3.2 adversarial review

Verdict: REWORK from both reviewers. 40 raw findings, 19 unique critical/high-priority issues after dedup against original Open Risks.

CRITICAL fixes (must land before first run)

C1. launchd plist must use ABSOLUTE paths

~ is not shell-expanded by launchd. All plist paths (ProgramArguments, StandardOutPath, StandardErrorPath, WorkingDirectory) must be /Users/daniel/.... Generate the plist at install time via envsubst or Python f-string, not as a checked-in template with ~.

C2. processed_db check moved INSIDE the per-session lock

Bug: Plan checked processed_db.seen(msg.guid) before locks.acquire(). Two concurrent workers can both pass the check, then both dispatch.
Fix: Acquire lock first → check seen → dispatch → mark. Order is non-negotiable.

C3. Body delivery via stdin, NOT argv

Bugs: (a) argv has E2BIG limit (~256KB on macOS); long iMessages crash silently. (b) Bodies with leading -, NULs, or terminal control bytes get misparsed as flags. (c) shell injection risk into downstream tools.
Fix: claude and imsg-callsign both receive the body on stdin. Use argparse --body-from-stdin flag or just pipe. Add explicit -- argv separator.

C4. Staged processing states for exactly-once

Bug: Single boolean seen causes either (a) message loss if marked before send + crash, or (b) duplicates if marked after send + crash.
Fix: Three states in processed.db: received_at (set immediately on enqueue), dispatched_at (set after claude returns), reply_sent_at (set after final chunk delivered). On daemon restart, recover by scanning rows with received_at but no reply_sent_at and replaying from the appropriate stage.

C5. Lock scope must NOT cover the 600s claude call

Bug: Per-session flock held for 600s during claude --print means 4 hung sessions starve everyone.
Fix: Short lock (≤2s) only for processed_db check + state transitions. The claude subprocess runs OUTSIDE the lock. Re-acquire briefly to mark dispatched_at. This requires C4's staged states to be safe.

C6. Mac asleep at 06:00 → quiet-hours drain never fires

Bug: Plan assumed daemon wakes at 06:00 wall clock. macOS may be asleep; LaunchAgents don't wake the system.
Fix: Add a separate com.callsign.morning-drain.plist with StartCalendarInterval (Hour=6, Minute=0) AND WakeFromSleep=true in pmset config. Daemon also drains lazily on every inbound message check.

C7. KeepAlive needs exponential backoff

Bug: while true; do imsg watch; sleep 2; done hot-restarts if imsg watch immediately crashes.
Fix: Backoff: 2s → 5s → 15s → 60s → 300s cap, reset after 30s of stable run. Track via a small bash counter or move the loop into Python.

C8. Kill fallback_recipient as silent exfiltration default

Bug: Any routing/parse failure forwards content to a hardcoded number. Privacy footgun.
Fix: No silent fallback. On unresolved routing, write to ~/.callsign/dead-letter/<ts>.json and os.system("osascript -e 'display notification ...'") to surface to daniel locally. Optional explicit per-callsign notify_on_fail setting.

HIGH-priority fixes

H1. SQLite PRAGMA journal_mode=WAL set on init

Set in processed_db.connect() once at first use; verify with PRAGMA journal_mode;. Also set PRAGMA busy_timeout=5000; for transient lock contention.

H2. Per-chunk idempotency token

Each outbound chunk gets a stable token (e.g., <guid>:k/N). Recipient can dedupe on retry; daemon can resume mid-chunk on crash. Stored in processed.db as chunks_sent INT.

H3. Unicode/byte-safe chunking

3500 bytes of UTF-8, not chars. Use text.encode() and split on UTF-8 boundary + sentence boundary. Test with emoji, ZWJ, CJK.

H4. chat_id routing for group threads

Dispatcher must accept chat_id and prefer it over sender for group threads. Plumb through imsg watch --json schema.

H5. Path-safe UID for flock

Slugify session UID before using as filename. re.sub(r'[^a-zA-Z0-9_-]', '_', uid)[:64].

H6. Process-group kill on timeout

subprocess.Popen(..., start_new_session=True) then os.killpg(proc.pid, signal.SIGKILL) on timeout. Otherwise grandchildren survive.

H7. SIGTERM/SIGHUP signal handlers in daemon

Catch SIGTERM, finish current dispatch with short grace period (5s), flush DB, exit clean. Otherwise launchctl kickstart corrupts state.

H8. Session UID validation before dispatch

Before claude --session-id <UID>, verify: (a) UID exists in registry as active, (b) the PID associated with it (if any) is alive, (c) the session file exists in ~/.claude/projects/*/. Otherwise claude may silently start a NEW session and answer from wrong context.

H9. ~/.callsign/* mode 0700

Lock down directory permissions. Otherwise any local process can poison registry/processed.db.

H10. CALLSIGN env sanitization in imsg-callsign

Strip newlines + control chars from $CALLSIGN before using as prefix. Otherwise log injection / spoofed prefixes possible.

H11. Independent alert path

Failure alerts must NOT route through imsg send. Use osascript display notification or write to ~/.callsign/alerts.jsonl polled by the SessionStart hook. Otherwise failure of primary path silences alerts about that very failure.

Resolved / accepted risks (no fix needed)

  • R1 (orig) imsg watch --json flag: confirmed needs upstream verification. Action: gate v0.3 on real test of imsg watch --json BEFORE writing the daemon. If flag absent, patch imsg first.
  • R2 (orig) claude --session-id <UID> --print: confirmed needs verification. Action: write tests/probe_claude_resume.sh as gating test #1.
  • R7 (orig) --print streaming: same — verified before coding.
  • R8 (orig) claim race: addressed via INSERT INTO registry (callsign) VALUES (?) ON CONFLICT(callsign) DO NOTHING RETURNING *.
  • DeepSeek #14 hardcoded +14053151310 invalidity: subsumed by C8 (no silent fallback).
  • GPT-5.4 #11 poison-message FIFO stall: mitigate via per-message try/except + 3-retry then dead-letter; one bad message doesn't stall the queue.
  • GPT-5.4 #13 DST/TZ: use zoneinfo.ZoneInfo("America/Chicago") not naive datetime; handles DST.

Build order (gated by probes)

  1. Probe imsg watch --json — write 1 line of JSON to stdout from an inbound iMessage? Yes → proceed. No → patch imsg first.
  2. Probe claude --session-id <UID> --print <body-via-stdin> — does it resume the specified session, run one turn, exit with reply on stdout? Yes → proceed. No → fall back to inbox-polling-via-UserPromptSubmit-hook architecture (different design, escalate to daniel).
  3. Fix SessionStart hook (auto-claim, kill Python heredoc).
  4. Schema migration: add claimed_via, received_at/dispatched_at/reply_sent_at, enable WAL.
  5. Implement dispatcher.py per C2, C3, C4, C5, H6, H8.
  6. Implement inbox.py, processed_db.py, locks.py, config.py with hardening.
  7. Generate launchd plists at install time with absolute paths.
  8. Smoke test: callsign smoke-test injects fake message through full pipeline without iMessage.
  9. E2E: real iMessage round-trip from daniel's phone.
  10. Tag v0.3.0, GitHub release.

Cost of review

GPT-5.4: $0.019. DeepSeek V3.2: $0.0008. Total $0.020. Caught ~$200 worth of post-deploy debugging.