Website: highball.profoundry.us
The Highball runner: executes a repo's .highball/checks.yml rules and blocks
AI coding agents on failure (exit 2, the Claude Code hook contract) — "local
CI for AI agents". Enforcement is entirely local and needs no account, no
network, and no configuration beyond checks.yml.
Runs can optionally be reported to PostHog for team-wide visibility. That is the witness half, and it is deliberately somebody else's server: you point at your own PostHog project, so Highball never takes custody of your data. The runner never uploads code — only rule ids, pass/fail, durations, and a one-line summary of a failure. Reporting is always best-effort: an unreachable endpoint means checks still run and still block, they just aren't recorded.
Published releases: npm install --save-dev @profoundry-us/highball.
Always use the scoped name. The unscoped npm name highball belongs to
an unrelated package, so a bare npx highball — in a committed hook, a
README, or a one-off — is a single uninstalled checkout away from fetching
a stranger's code and running it.
From a local tarball (pre-release):
npm pack # in this repo → profoundry-us-highball-<v>.tgz
npm install --save-dev ../highball/profoundry-us-highball-<v>.tgzHighball is installed by the AI agent that will be checked by it. After installing the package, tell the repo's Claude Code agent:
Run
npx @profoundry-us/highball onboardand follow the instructions.
ONBOARDING.md (which that command prints) walks the agent through surveying the repo's real toolchain, scaffolding, writing rules that reflect what the repo already trusts, and verifying all four proofs — including that exit 2 actually blocks.
The pieces, for reference or manual setup:
npx @profoundry-us/highball init # scaffolds checks.yml + Claude Code hooksinit never overwrites an existing checks.yml and never edits an existing
.claude/settings.json (it prints the hook snippet to merge by hand). There
is no login step: the only credential Highball takes is a PostHog project
key, which is write-only by design and lives in committed config.
The fast hook matches Write|Edit|Bash and runs --fast --if-changed.
Agents in auto mode edit through Bash — sed, heredocs, scripts — so a hook
on the edit tools alone never fires for them (one repo here went two weeks
with zero fast runs that way while full runs kept landing). Matching Bash
too would run the fast rules after every command, most of them reads, so
--if-changed fingerprints the working tree (HEAD plus each dirty path's
size and mtime) and exits at once when nothing moved since the last run.
Older installs: change the matcher and add the flag by hand.
version: 1
project: my-app
# runs_limit: 25 # rows the MCP widget / list_runs return; HIGHBALL_RUNS_LIMIT overrides
# Optional. Omit the block entirely to keep runs in the local journal only.
reporting:
posthog:
host: https://us.i.posthog.com # EU cloud or self-hosted also fine
project_key: phc_xxx # write-only by design — committed
# Containerized toolchain? Declare the wrapper once and every rule runs
# through it; rules opt out with `exec: host`. Rule definitions stay
# environment-agnostic on purpose — the *where* is per-checkout config.
exec:
via: docker compose exec -T app
checks:
- id: unit-tests
name: Unit tests
run: bundle exec rspec spec # runs through exec.via
- id: js-syntax
name: Playback JS parses
run: node --check web/app.js
exec: host # host-side tool, opts out
fast: true # cheap → runs on every agent edit
- id: architecture-quality
name: Architecture & naming (AI)
rubric: .highball/packs/rails/rubrics/architecture.md
- id: coverage-ratchet
name: Coverage never decreases
todo: true # declared, tracked, not yet builtThe runner computes the branch's changed-file list once (it owns git) and
hands it to every rule via HIGHBALL_CHANGED_FILES — check scripts stay pure
analyzers and need no git in their execution context.
Three switches, differing in who they affect and — the part that usually decides it — when they take effect:
touch .highball/disabled # this checkout, gitignored, next runenabled: false # in checks.yml — committed, everyone, next runexport HIGHBALL_DISABLED=1 # your whole machine, next session restartAny of them makes highball run exit 0 immediately, with no rule executed, no
journal entry, and nothing reported. Hooks and rules stay exactly where they
are, so switching back on is a one-line revert.
| takes effect | scope | committed | |
|---|---|---|---|
.highball/disabled |
next run | this checkout | no — init gitignores it |
enabled: false |
next run | one repo, everyone who clones it | yes |
HIGHBALL_DISABLED |
next Claude Code restart, for hooks | every repo, your machine | no |
Reach for the marker file first. touch .highball/disabled while an agent
session is running and the very next hook run skips; delete it and the run
after that is armed again. Both directions are live because the file is
checked on every run and every hook invocation is a new process — there is no
state anywhere and nothing to restart. init writes a .highball/.gitignore
alongside it, so a switch-off in your checkout can't ride along in a commit.
Anything you write into the marker comes back as the reason on every run, which is what the person who finds the checks off next week actually needs:
$ echo "bisecting a flaky spec" > .highball/disabled
$ npx @profoundry-us/highball run --fast
highball: disabled by .highball/disabled (bisecting a flaky spec) — no checks runenabled: false is the committed counterpart, for a repo genuinely stepping
away from its checks. It lands in a diff, which is what you want when the
decision belongs to the team rather than to your afternoon.
HIGHBALL_DISABLED is an ordinary environment variable, so a hook sees
whatever value the Claude Code process had when it started — Claude Code
writes settings env entries into its environment at launch, and a shell
export after that never reaches an already-running session. That makes it
the right switch for a machine-wide default you rarely change, and the wrong
one for a mid-session toggle. Like the marker, it is read before
checks.yml, so both still work when the config is itself what's broken.
HIGHBALL_DISABLED=0, false, no, off and empty all mean not
disabled. Anything else disables. Plain truthiness would make =0 stop every
check in the repo, which is exactly the "is the guardrail live right now?"
doubt this switch exists to remove.
None of the three is silent. Each prints a line on every run, because a guardrail
that has quietly stopped guarding is worse than no guardrail — the next
person reads green and believes it. For the same reason enabled: accepts
only a real boolean: enabled: "false" and enabled: no are strings, and
rather than leaving checks quietly on, they fail loudly.
A rule with rubric: instead of run: is judged by headless Claude rather
than by a script: the runner bundles the changed files the rubric asks for,
applies the rubric, and turns the verdict into the same pass/fail contract
every other rule uses.
The rubric is markdown with optional YAML front matter, which carries the only language-specific part:
---
include: "**/*.rb" # default: every changed file
exclude: [db/, config/] # path prefixes
model: claude-haiku-4-5-20251001
---
A comment VIOLATES this rubric when it restates what the code already says.Three properties are enforced by the runner rather than left to each repo:
- Never on the fast path. Rubric rules are dropped from
--fastruns even if markedfast: true— otherwise you pay model latency on every edit. - Always host-side. The judge needs the
claudeCLI, soexec.vianever wraps it and noexec: hostannotation is required. - No evidence, no call. When nothing in the changed set matches
include, the rule passes without spawning the model at all.
Rubrics live with the opinions they express: a framework pack such as
@profoundry-us/highball-rails ships them, and the runner supplies the engine.
reporting.posthog sends runs to PostHog — the runner's only telemetry path.
A team that already runs PostHog needs no server for this, and a team that
doesn't can skip the block entirely and use the local journal. The project key
is write-only by design — it is the same key PostHog has you ship in browser
bundles, and all it can do is capture events — so committing it is safe:
no login step, no credentials file.
If you would still rather keep it out of the repo, set the environment
instead. HIGHBALL_POSTHOG_KEY (or POSTHOG_API_KEY) on its own turns
reporting on with no reporting: block at all, and HIGHBALL_POSTHOG_HOST
(or POSTHOG_HOST) points it at an EU or self-hosted instance; either wins
over committed config, which is also how CI redirects a run. One place that
reaches every repo's hooks is the env block of your user-level Claude Code
settings, ~/.claude/settings.json:
{ "env": { "HIGHBALL_POSTHOG_KEY": "phc_your_key" } }The trade-off is that reporting is then per machine rather than per repo: a teammate cloning the repo reports nothing until they set the key too.
The whole run leaves in ONE request to /batch/. PostHog events are
immutable, which suits a runner that already defers reporting to after the
checks finish.
Two event types per run:
| event | one per | key properties |
|---|---|---|
highball_run |
run | status, duration_ms, rules_passed/failed/todo, rules_run[], failed_rules[] |
highball_check |
rule result | rule_id, rule_name, status, duration_ms, summary, command |
Run context (project, branch, commit, trigger, session_key,
runner_version) is repeated on every event rather than joined at query time,
because PostHog has no join back to a run: a breakdown like "failure rate by
rule, on this branch only" needs branch on the check event itself.
rules_run[] is the denominator for any per-rule rate that spans runs whose
rulesets differ.
Log tails are never sent — multi-kilobyte blobs in event properties bloat the
column store and slow every query that touches it. Failures carry a one-line
summary; the full output stays in the local journal (below).
Volume is smaller than "runs on every edit" suggests, because agent edits
batch into turns. Measured across 842 real runs on one developer's machine
over 18 active days and 9 projects: a median of 22 runs/day, edit and stop
runs at close to 1.3:1, and ~15k events/month/dev at the model above.
Dashboard queries live in docs/posthog-queries.sql — rule cost vs. benefit, failure rate by week, fast-path latency, repo health, and todo debt. They are versioned rather than left as PostHog UI state, which drifts and cannot be reviewed.
Events are attributed to git config user.email, falling back to
user@hostname when git has no identity (CI images, fresh containers).
HIGHBALL_POSTHOG_DISTINCT_ID overrides it.
highball mcp serves the journal over MCP (stdio) with three tools —
list_runs, get_run, run_checks — and an
MCP Apps widget:
in hosts that render Apps (Claude Desktop and friends), asking about your
checks produces an interactive inline dashboard — click a run for per-rule
detail with expandable command output, re-run fast or full checks from a
button. In hosts without Apps support the same tools answer in plain text,
per the extension's graceful-degradation rule. Register it with the scoped
name — hosts spawn the server from an arbitrary directory, so it resolves
from the registry rather than a local install:
"highball": { "command": "npx", "args": ["-y", "@profoundry-us/highball", "mcp"] }The journal it reads is machine-global (~/.highball/runs/), so one
registration covers every repo on that machine — there is no per-repo MCP
setup.
Which project the widget shows is never guessed. Hosts spawn the server with
no useful working directory (Claude Desktop and Claude Code both use /), so
list_runs resolves the project from its project or dir argument, a
.highball/checks.yml at the server's cwd, or the client's MCP roots. When
none of those names a repo it returns the journaled projects for the widget
to offer as a picker, rather than showing whichever repo happened to run most
recently. An agent calling from inside a repo should pass dir.
list_runs returns the newest 25 runs, and the widget says so at the top.
Change it per call with limit, per machine with HIGHBALL_RUNS_LIMIT, or
per repo with runs_limit: in checks.yml.
The split is capability-driven, not guesswork: the server reads the
client's initialize capabilities (io.modelcontextprotocol/ui) — hosts
that render MCP Apps get a short text summary plus the widget; everything
else gets the full picture as aligned plain text. Widget development has
its own harness — npm run harness, open http://localhost:3777 — which
plays the host role against the real assets/dashboard.html and live
journal data, so widget edits are a reload away instead of a Claude
Desktop restart.
Every run appends to a local journal (~/.highball/runs/<project>.jsonl,
pruned to the last 200) — unconditionally, whether or not reporting is
configured. npx @profoundry-us/highball runs lists recent runs; adding a
number shows one run's detail with failure output, and --logs prints every
rule's captured output, GitHub-Actions-style.
The journal is the richer of the two records: it keeps the last 10KB per rule
pass or fail, while PostHog gets a one-line summary and no logs at all. So the
runner is fully self-sufficient with no reporting: block — PostHog adds
cross-developer trends, not visibility you'd otherwise lack.
Built-in generic rules (spec pairing, focused-spec detection, diff budgets)
land here in a future release, along with more per-framework starter packs
(highball-python, highball-go) carrying recommended check scripts. The
Rails pack (@profoundry-us/highball-rails)
and AI-judged rubric: rules have shipped.