Shared spaces for people and bots.
Chatticus is a shared, collaborative space where people and named bots work together around common files, tools, and a system of authority and approvals — like an office, not a chat window.
The production product workspace is hey.chattic.us.
The public marketing site, chattic.us, is a separate
private repo (AnthusAI/Chattic.us-web) that depends on this repo's shared
UI kit and brand tokens; see the exports map in package.json
and Kanbus epic chatticus-3926bc for the public/private boundary.
v1 is personal: one household, one AWS account, as many named bots as we
want. Every record already carries a tenant_id so the same system can
later serve other people without a rewrite.
A bot is a persistent, named teammate. It keeps memory, preferences, and conversation. It is not a fresh chat session on every task.
A computer is a Linux workplace with a browser, a filesystem, and a
terminal. It belongs to a user, not to a bot. Every bot on that user
shares /workspace, browser sessions, and command-line credentials so they
can hand work off. Each bot gets its own screen so they can work in parallel.
Separate bots are not a security boundary.
Work prefers structured tools (APIs, MCP servers, connectors) when they exist. The computer's browser is the fallback for sites and apps that have no clean API.
A skill is how to do a task. A routine is when to run it (a schedule or an event). Approvals stop consequential actions: sending, purchasing, publishing, deleting, and production changes.
Closing your laptop does not stop work. The computer runs on a worker, not on the device in front of you.
A turn is one bot response: from a human message (or a routine) until the control plane commits a single bot message to the transcript.
There are no persistent sockets. The browser POSTs a message and then
reads server-sent events scoped to that one turn. Reconnecting is a new
SSE request that reads already-committed chunks after Last-Event-Id.
In-process queues are not the architecture.
The control plane is the AWS side that accepts HTTP, stores tenant state, and enqueues jobs. It does not run the model loop and it does not own a display. It holds no always-on process in the token path, so when nobody is working, nothing bills. That idle floor is the point: a quiet household (and later, quiet tenants) should not pay for capacity that is sitting empty. There is never a load balancer in front of the API; an hourly LB charge would recreate the floor this design avoids.
Pull workers register, heartbeat, and take jobs from a queue. The
control plane never SSHs into a garage Mac. A worker that has CPU and
network but no display, browser, or /workspace is computerless.
Text-only turns can finish there. The computer is summoned when a tool
needs it, not assumed at enqueue: reaching for a display, browser, or
/workspace escalates the same turn onto a computer-capable worker. No
second transcript.
See Architecture for routing, Messaging for the transcript and stream, and Design challenges for rejected alternatives.
Last updated: 2026-09-02. Git main and develop are aligned after
the v0.10 promote: principal enforcement (#7b4616), sign-out ending the SSO
session (#169), and the behavior-driven spec migration are on main and deployed
across all three named environments. Future work lands on develop and promotes
to main for release.
| Environment | ThinTurn API (enforcement) | Cognito Auth (sign-out) | Web bundle |
|---|---|---|---|
development (dev.chattic.us) |
Live — CloudFront → Lambda | Live — logoutUrls + /auth/signout-callback |
Live — product /chat at / (CF enabled; marketing moved to AnthusAI/Chattic.us-web, chatticus-3926bc) |
staging (staging.chattic.us) |
Live — Lambda URL | Live — logoutUrls + callback |
Deployed — S3 bundle staged, CF dark by design |
production (hey.chattic.us) |
Live — Lambda URL | Live — logoutUrls + callback |
Deployed — S3 bundle staged, CF dark by design |
Every org route requires a Cognito id_token or worker bearer on all three
environments; only /health and POST …/workers/register (invoke-key gated)
are open. Live-verified on each environment: /health 200, /me 403, org
route 403 with no credential.
A seeded owner (ryan@anth.us → org anthus, display name Anthus AI
Solutions, enabled) can sign in with Google on development and reach the
workspace. Operator org records are DynamoDB data, not CDK; see
Operator org seed.
- Phase 5 org-scoping (6c1a9b, 0814e6, ddf609) — in flight; will drop
user_idfrom bot/channel/computer identity - Org-scoped HTTP, worker bearer credentials, members CLI, budgets stack, turn recovery kernel — same as before
| Host | Role | Notes |
|---|---|---|
| chattic.us | Marketing (private AnthusAI/Chattic.us-web repo) |
Updates, Agent Zoo, and the Markus wiki at /wiki; same-origin /api to this repo's thin-turn front door |
| dev.chattic.us | Development product + API | Same-origin /api; product /chat at / |
| hey.chattic.us | Production product (planned) | Web CloudFront disabled (stack exists, dark) |
| staging.chattic.us | Staging (planned) | Web CloudFront disabled |
Staging and production thin-turn stacks may lag develop; their web front doors
stay dark until explicitly re-enabled and re-proven.
Live stack proof is manual: sign in at dev.chattic.us and send a message.
GitHub Deploy ThinTurn (development) and Deploy Web (development) are
manual (workflow_dispatch). GitHub Actions must not hit live AWS.
ChatticusDns, ChatticusAuth, ChatticusWeb* (development CF enabled; staging/
production CF disabled), ChatticusThinTurn*, ChatticusSnapshots, ChatticusComputers,
ChatticusBudgets. Never cdk deploy --all. Computers desiredCount stays 0.
Overnight gated-action, immutable approval binding, unbound-browser stops, computer-seam recovery, capability-gated readiness beyond waiting turns, and structured handoff journal events are tested in Gherkin/pytest but not wired into the live worker HTTP loop yet. See Approval spec and Organizations.
flowchart LR
Caller["Caller<br/>exercise script or HTTP"]
CF["CloudFront"]
FD["Front-door Lambda<br/>FastAPI, turn SSE"]
DDB[("DynamoDB: messages, chunks, roster")]
SQS["SQS turn jobs"]
W["Computerless worker Lambda"]
OA["OpenAI<br/>gpt-5.6-luna"]
Caller --> CF --> FD
FD --> DDB
FD --> SQS
SQS --> W
W --> OA
W -->|"POST turn chunks"| FD
FD -->|"poll after cursor"| DDB
sequenceDiagram
participant Caller
participant CF as CloudFront
participant FD as Front-door Lambda
participant DDB as DynamoDB
participant SQS as Turn queue
participant W as Computerless worker
participant OA as OpenAI
Caller->>CF: POST /channels/channel_id/messages
CF->>FD: origin
FD->>DDB: commit human message
FD->>SQS: enqueue turn
FD-->>Caller: turn_id
Caller->>CF: GET /turns/turn_id/stream
SQS->>W: deliver job
W->>OA: text-only completion
W->>FD: POST coalesced chunks
FD->>DDB: write chunk items with TTL
loop poll store
FD->>DDB: read after cursor
FD-->>Caller: SSE frames
end
FD->>DDB: commit one bot message at turn.completed
v1 is the product Next.js app at hey.chattic.us (production) talking to that same per-request control
plane, plus pull workers that can stay computerless or host the user's
Linux computer: local Docker when a Mac is on, Fargate ARM64 that scales
to zero when it is not. Workplace disk lives in S3 snapshots. EventBridge
will wake routines later. The model vendor is OpenAI; Amazon Bedrock may
follow.
Lambda is the right runtime for HTTP, auth callbacks, webhooks, routine wake-ups, a computerless model loop, and holding one turn's SSE stream. It is the wrong runtime for a workplace: it cannot hold a browser, a display, or a session-lifetime connection. The computer is one Docker image on Fargate, later stop/start EC2, or Docker on a Mac.
flowchart TB
Person["Person"]
Web["hey.chattic.us<br/>Next.js"]
subgraph control [Control plane per request]
FD["HTTP front door<br/>POST plus one-turn SSE"]
Sched["Scheduler / routing"]
DDB[("DynamoDB: bots, transcript, rules")]
SQS["SQS turn jobs"]
EB["EventBridge<br/>routines, worker starts, push"]
SM["Secrets Manager"]
end
subgraph workers [Pull workers]
CL["Computerless worker<br/>cpu, no display"]
Local["Garage Mac<br/>Docker computer"]
Far["Fargate ARM64 computer<br/>scale to 0"]
end
S3[("S3: snapshots, screenshots")]
LLM["OpenAI<br/>Bedrock later"]
Person --> Web --> FD
FD --> DDB
FD --> Sched
Sched --> SQS
EB --> SQS
SQS --> CL
SQS --> Local
SQS --> Far
CL --> LLM
Local --> LLM
Far --> LLM
CL -->|"POST chunks"| FD
Local -->|"POST chunks"| FD
Far -->|"POST chunks"| FD
Local --> S3
Far --> S3
FD --> SM
How a turn chooses a host:
- A message or routine creates a turn job with
tenant_id, required capabilities (cpu, and maybecomputer/browser/terminal), and an optionalcomputer_idpin. - Healthy workers pull. Prefer an already-warm cheap host (local Docker when the Mac is on), then a warm AWS computer, then a cold Fargate start.
- The model loop runs on the worker, as soon as it has network and memory. Display, Chromium, and snapshot hydrate gate only the actions that need them.
- Reaching for a computer tool escalates the same turn: the computerless
attempt appends the pending tool call and a computer-capable worker
continues. There is no
stop_computer; the workplace is shared by all of a user's bots. - At
turn.completedthe control plane commits one message row.
Compute and disk are separate. The computer is a computer_id. The
host is whichever Mac, Fargate task, or EC2 instance is running it.
Canonical workplace disk (/workspace and the browser profile) is an S3
snapshot. Hosts hydrate a local cache, run, and publish. Relocate is an
administrator action: point the next host at that snapshot. There is no live
container move. See Computer snapshots.
| State | Where it lives |
|---|---|
| Bot memory, skills, routines | DynamoDB |
| Transcript (append-only messages) | DynamoDB |
| In-flight turn chunks | DynamoDB, with a TTL; they expire after the turn commits |
/workspace and browser profile |
S3 snapshot; local volume / EBS / EFS as a host cache |
| Secrets | AWS Secrets Manager, never in the image |
Computer policy: prefer_local, aws_only, or local_only. Idle AWS
computers stop (EC2) or scale to 0 (Fargate). The snapshot stays.
Chattic.us/
features/ Shared Gherkin (product narrative)
python/ Control plane, computerless and computer workers, snapshot packer
web/ Marketing (`/`) and product workspace (`/chat`)
computer/ Linux computer image
infra/ AWS CDK
docs/ Product, architecture, design challenges, stack
v1 language for the product brain is Python. The web app is
TypeScript. Gherkin in features/ is the behavior spec. Root
package.json workspaces web and infra so one npm install at the
repo root installs all JavaScript dependencies (Node 22+).
Local JavaScript setup (Node 22+):
npm install
npm run build
npm run test
npm run lint
npm run devLocal quality gates (CI uses a fake OpenAI client; a live key is not required):
cd python
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
behave
pytest
black --check src ../features tests
ruff check src ../features testsblack and ruff versions are pinned in python/pyproject.toml so local
pip install -e ".[dev]" matches GitHub CI.
The deployed thin turn is exercised against a named cloud environment
(CloudFront on development only today), not against an in-process queue.
GitHub CI (behave, pytest) uses in-memory stores and moto. Live stack
proof is manual: sign in at dev.chattic.us and
send a message.
Watch one live conversation as a human (tokens on stdout, committed reply at the end). Workers use bearer tokens after registration. The product SPA uses Google sign-in on development; scripts still use invoke key + worker bearer until #7b4616 lands:
cd python
export CHATTICUS_INVOKE_KEY=... # or pass --invoke-key
python scripts/chatticus_chat.py --environment development \
--tenant-id anthus --user-id ryan --bot Luna --message "hello"The script resolves the front door from CHATTICUS_DEVELOPMENT_BASE_URL,
SSM, or CloudFormation. Omit
--message for an interactive prompt. --list-turns calls
GET /users/{user_id}/turns; --watch-turn reconnects with
Last-Event-ID on GET /turns/{id}/stream.
That resolves the front door from CHATTICUS_DEVELOPMENT_BASE_URL, SSM
/chatticus/development/thin-turn/cloudfront-url, or the
CloudFrontUrl output on stack ChatticusWeb (or ChatticusThinTurn
before the web stack exists). When AWS lookup fails, pass --base-url or
set the env var (see gitignored AGENTS.local.md). SQS queue checks
need that same AWS identity. Staging and production web CloudFront is
disabled; use thin-turn stack outputs or AGENTS.local.md when you mean
those environments. It does not scale Fargate. Do not run this from
GitHub Actions.
If Docker Desktop is running, the snapshot packer can be checked with
sh computer/test_relocate.sh.
AWS resources are CDK only (infra/). Do not create them in the console.
cdk deploy --all would touch ChatticusSnapshots,
ChatticusComputers, and every thin-turn environment; never do that.
Deploy one named stack:
Deploy development ThinTurn only:
cd infra
sh deploy-chatticus-thinturn-development.shThat fails closed without AWS identity and never passes --all. Staging
and production, when you mean those stacks:
npx cdk deploy ChatticusThinTurnStaging
npx cdk deploy ChatticusThinTurnProductionChatticusThinTurn is development. Staging and production are separate
stacks with their own DynamoDB, SQS, Lambda, and CloudFront (web CF
disabled on staging and production). Do not merge develop to main as
parking. Do not destroy the snapshot, computer, or budgets stacks.
Postgres in docker-compose.yml is unused (it predates DynamoDB).
This repository uses Kanbus. Prefer
kbs when it is on PATH. See CONTRIBUTING_AGENT.md.
kbs list
kbs create "Describe the work" --type taskAt every milestone (a slice merged to develop, a promote to main, a
deploy, a closed epic, or anything that changes what is true), launch a
sub-agent whose only job is housekeeping. That agent comments on the
Kanbus issues it touches, sets statuses to match reality, creates new
issues for significant work that appeared, and rewrites the "What is
live today" section of this README (and the next-up line) so a new reader
is not looking at last week's world. Do not skip that pass. Do not treat it
as a leftover for the implementation agent.
- Product
- Architecture
- Messages and the cloud API
- Design challenges
- Computer snapshots
- Stack
- Roadmap
- Approval spec
- Browser authority
- Threat model
- Tasks
- Feasibility tests
- AGENTS.md for coding agents working in this repo