An async, event-driven agent that monitors a Gmail inbox, scrapes regulatory filings from the Nova Scotia UARB portal, and replies with a ZIP of the documents and a structured metadata summary.
Note
Status: Vellum is currently offline. The hosted deployment has been paused to avoid the ongoing cost of keeping the Google Cloud infrastructure running. For now, Vellum is primarily a showcase of agentic capabilities and an exploration of what an end-to-end autonomous document retrieval workflow can look like.
See the demo for a walkthrough of the system in action.
The hosted Vellum instance is currently offline, so the public email interface is not available at this time.
Vellum can still be run locally or self-hosted using the deployment instructions below.
When running, Vellum accepts natural-language requests over email. Send a
message with [vellum] anywhere in the subject line, and provide the matter
number and requested document types in the subject or body.
| What you want | Subject line |
|---|---|
| A single document type | [vellum] Exhibits for M12205 |
| Multiple document types | [vellum] Exhibits and Transcripts for M12205 |
| Everything on a matter | [vellum] all documents for M12205 |
| Free-form | Subject: [vellum] filing request · Body: Can you pull Exhibits and Key Documents for M12205? |
- Exhibits
- Key Documents
- Other Documents
- Transcripts
- Recordings
When running, Vellum replies within a few minutes with a ZIP archive named
vellum_<matter>_<types>_<date>.zip, containing one subfolder per requested
document type (up to 10 files each) and a job_summary.json with structured
matter metadata: title, status, category, dates, tab counts, and a file manifest.
If a requested tab has no documents, or the matter number is not found, Vellum returns an error reply explaining what was attempted.
flowchart LR
G[Gmail Inbox] -->|Push notification| P[Pub/Sub]
P -->|webhook| W[FastAPI /gmail/webhook]
W -->|ParsedRequest| Q[asyncio.Queue]
Q --> WK[Worker pool N=2]
WK --> S[scraper - Playwright]
S --> PK[packager - ZIP + job_summary.json]
PK --> SM[summarizer - HTML body]
PK --> ST[GCS delivery link]
SM --> M[mailer - Resend reply]
M --> U[User inbox]
The webhook acknowledges Pub/Sub immediately and puts the job onto an
asyncio.Queue. A pool of worker coroutines runs the scrape-package-render-reply
pipeline, bounding concurrent Playwright sessions to avoid hammering the UARB site.
The architecture is designed to support an always-on hosted deployment, but the current hosted instance is paused.
-
Clone the repository and install dependencies:
uv sync uv run playwright install chromium
-
Create a Google Cloud project with the Gmail API, Pub/Sub API, Cloud Run API, and Secret Manager API enabled. Create a Pub/Sub topic and a push subscription pointed at
/gmail/webhook. Grantgmail-api-push@system.gserviceaccount.comtheroles/pubsub.publisherrole on the topic. -
Create OAuth2 Desktop credentials, download them as
credentials.json, and run the authorization flow once. The OAuth token needs Gmail read/watch access so Vellum can process inbound requests:uv run python scripts/authorize.py
-
Copy
.env.exampleto.local.envand fill in your values. -
Deploy to Cloud Run:
bash scripts/deploy.sh
The script builds the image, pushes it to Artifact Registry, stores
credentials.jsonandtoken.jsonin Secret Manager, deploys the service, and updates the Pub/Sub subscription endpoint.
| Variable | Description |
|---|---|
GMAIL_ADDRESS |
The mailbox Vellum watches and replies from. |
GMAIL_CREDENTIALS_PATH |
Path to the OAuth2 client credentials JSON. |
GMAIL_TOKEN_PATH |
Path to the cached OAuth token (default token.json). |
PUBSUB_TOPIC |
Fully qualified Pub/Sub topic for Gmail push. |
PUBSUB_SUBSCRIPTION |
Fully qualified Pub/Sub push subscription. |
OPENAI_API_KEY |
OpenAI API key used by the email parser. |
OPENAI_MODEL |
Parser model (default gpt-5.4-mini). |
MAX_CONCURRENT_WORKERS |
Worker coroutines / max concurrent browsers (default 2). |
MAX_DOCUMENTS |
Maximum documents downloaded per tab (default 10). |
DOWNLOAD_TIMEOUT_MS |
Per-download timeout in milliseconds (default 30000). |
SCRAPER_RETRY_ATTEMPTS |
Download retry attempts (default 3). |
SCRAPER_RETRY_BACKOFF_S |
Backoff between retries in seconds (default 2). |
PARSER_MAX_CALLS_PER_MINUTE |
Rate limit for parser LLM calls (default 20). |
UARB_BASE_URL |
UARB portal entry URL. |
SELECTOR_TIMEOUT_MS |
Explicit wait timeout for selectors (default 15000). |
SCRAPER_HEADLESS |
Run Chromium headless (default true). |
EMAIL_FROM / EMAIL_FROM_NAME |
Reply sender address and display name. |
RESEND_API_KEY |
Resend API key used for outbound replies. |
MAILER_SEND_RETRY_ATTEMPTS |
Gmail send attempts before surfacing delivery failure (default 3). |
MAILER_RETRY_BASE_DELAY_S |
Base backoff for retryable mailer errors without provider retry hints (default 2). |
MAILER_MAX_RETRY_DELAY_S |
Maximum delay between Gmail send retries in seconds (default 900). |
HOST / PORT |
FastAPI bind address (default 0.0.0.0:8000). |
- Listener (
agent/listener.py) receives the Pub/Sub push, decodes thehistoryId, fetches new messages via the Gmail API, and extracts the sender, subject, and body. - Parser (
agent/parser.py) ignores emails without the[vellum]subject tag, then uses a single LLM call to extract the matter number and document types. A sliding-window rate limiter caps API usage. - Scraper (
agent/scraper.py) drives Chromium against the FileMaker WebDirect portal: searches the matter, extracts header metadata, parses tab counts, scrolls paginated lists, and intercepts each download. Empty tabs are non-fatal. - Packager (
agent/packager.py) builds the ZIP with one subfolder per type and writesjob_summary.json. - Summarizer (
agent/summarizer.py) renders the HTML email body and subject line from the matter metadata. - Storage (
agent/storage.py) uploads the ZIP to Cloud Storage and produces a temporary download link. - Mailer (
agent/mailer.py) sends the reply through Resend with the download link in the HTML body.
These commands are intended for a self-hosted deployment.
Set these values for your deployment:
PROJECT_ID="your-google-cloud-project-id"
REGION="your-cloud-run-region"
SERVICE="vellum"Check which Cloud Run revision is live:
gcloud run services describe "${SERVICE}" \
--project="${PROJECT_ID}" \
--region="${REGION}" \
--format='table(status.latestReadyRevisionName,status.latestCreatedRevisionName,status.traffic[0].revisionName,status.traffic[0].percent)'Read the latest Vellum logs:
gcloud logging read "resource.type=\"cloud_run_revision\" AND resource.labels.service_name=\"${SERVICE}\"" \
--project="${PROJECT_ID}" \
--limit=80 \
--freshness=20m \
--format='table(timestamp,resource.labels.revision_name,jsonPayload.request_id,jsonPayload.step,jsonPayload.level,jsonPayload.message,jsonPayload.matter_number,jsonPayload.document_types,jsonPayload.error_type,jsonPayload.reason)'Follow one request after you have its request_id:
REQUEST_ID="paste-request-id-here"
gcloud logging read "resource.type=\"cloud_run_revision\" AND resource.labels.service_name=\"${SERVICE}\" AND jsonPayload.request_id=\"${REQUEST_ID}\"" \
--project="${PROJECT_ID}" \
--limit=200 \
--format='table(timestamp,jsonPayload.step,jsonPayload.level,jsonPayload.message,jsonPayload.downloaded,jsonPayload.download_url,jsonPayload.error_type,jsonPayload.reason)'Healthy delivery ends with:
mailer.send INFO email sent
queue.complete INFO job complete
Run the test suite:
uv run pytestRun the live scraper integration test against the UARB site:
VELLUM_LIVE_TESTS=1 uv run pytest tests/test_scraper.pyExercise the full pipeline without sending email:
uv run python main.py --dry-run --matter M12205 --types "Exhibits,Transcripts"
uv run python main.py --dry-run --matter M12205 --types all- Vector database ingestion after download, feeding
job_summary.jsonand the raw documents into a searchable regulatory index. - Multi-mailbox support.
- A dead-letter queue for failed jobs with automatic replay.
- Revisit hosted operation if Vellum proves useful enough to justify the infrastructure costs.