Skip to content

fix(grok): include subagent session usage - #1193

Merged
robinebers merged 1 commit into
mainfrom
codex/grok-subagent-usage-1191
Sep 2, 2026
Merged

fix(grok): include subagent session usage#1193
robinebers merged 1 commit into
mainfrom
codex/grok-subagent-usage-1191

Conversation

@robinebers

@robinebers robinebers commented Sep 1, 2026

Copy link
Copy Markdown
Owner

TL;DR

Include Grok subagent, resumed, and forked session usage in local spend tiles and Usage Trend, while retaining event-ID/model deduplication. Fixes #1191.

What was happening

  • The scanner skipped every session marked subagent*, assuming the coordinator already included its usage.
  • Child work missing from coordinator turns disappeared from spend totals and could leave entire days without usage.
  • An unreadable session summary also excluded an otherwise valid usage ledger.

What this changes

  • Scan every durable updates.jsonl ledger regardless of session kind, without requiring summary.json.
  • Preserve existing event-ID/model deduplication across parent and child logs.
  • Add regression coverage for child-only days, subagent/resume/fork sessions, replayed multi-model events, and malformed summaries.
  • Update the Grok provider documentation to describe child-session accounting.

Heads-up

  • Deduplication uses event IDs and models, not matching token totals: separate turns can have identical totals.
  • Weekly billing metrics and metric layout defaults are unchanged.

Tests

  • swift test --filter Grok compiled successfully; the local Xcode beta test runner initially could not locate Sparkle.
  • After linking the existing Sparkle artifact into the local build's PackageFrameworks directory, swift test --skip-build --filter Grok passed all 50 tests.
  • CONFIG=debug ./script/build_and_run.sh verify rebuilt, signed, and launched the dev app successfully.
  • git diff --check passed.

Note

Medium Risk
Changes Grok spend aggregation logic users rely on for cost visibility; dedup by event ID may miss edge cases where IDs differ for the same work.

Overview
Grok local spend tiles no longer exclude subagent, fork, and resume sessions based on summary.json session_kind. The scanner now ingests every updates.jsonl under $GROK_HOME/sessions, so usage that only appears in child ledgers shows up in Today / Yesterday / Last 30 Days.

Double-counting is still avoided by the existing event ID + model dedup when the same turn_completed is copied across parent and child logs. Sessions with missing or malformed summary.json are counted again instead of being dropped entirely.

Tests and Grok provider docs were updated for child-only days, replayed events, and summary-independent reads.

Reviewed by Cursor Bugbot for commit dca9625. Bugbot is set up for automated code reviews on this repo. Configure here.

@robinebers
robinebers merged commit bb94841 into main Sep 2, 2026
5 of 6 checks passed
@robinebers
robinebers deleted the codex/grok-subagent-usage-1191 branch September 2, 2026 09:42
mstallone added a commit to mstallone/runway that referenced this pull request Sep 5, 2026
## TL;DR

Selective re-implementation of the OpenUsage commits since Runway #111
that still apply: Claude Desktop's account-prefixed token caches, Claude
spend tiles without an OAuth login, Grok subagent session ledgers, Codex
Business Premium, Cursor's Models/Other Models labels, and new model
rates.

## What was happening

- Upstream has moved on since #111. Each new OpenUsage commit was
reviewed against Runway's architecture, existing ports, and the "don't
bloat" bar.
- Recent Claude Desktop builds store tokens under `acct:<user>|<legacy
key>` (openusage robinebers#1212). Runway expected the client UUID first and
skipped those entries, so a Desktop-only login showed Not logged in.
- Claude returned a hard authentication error before scanning local logs
when no OAuth login existed (openusage robinebers#1138), so API-key gateway users
lost Today/Yesterday/Last 30 Days.
- Grok's scanner skipped every `subagent*` session (openusage robinebers#1193).
Child work that the coordinator turn did not include disappeared from
spend.
- Codex's `self_serve_business_prolite` entitlement rendered as "Self
Serve Business Prolite" (openusage robinebers#1194).
- Cursor's dashboard now calls the two model pools **Cursor Models** and
**Other Models**; Runway still said Auto Usage / API Usage (openusage
robinebers#1134, labels only).
- GPT-6 Astra, Gemini 3.8 Flash, Fable 5.1, GLM 5.3, and Grok Bot CSV
slugs had no supplement entries, so those rows tripped the
unpriced-model warning.

## What this changes

- Desktop cache selection strips the `acct:<user>|` prefix, keeps only
the signed-in account (from `lastKnownAccountUuid`), and lets a scoped
tombstone suppress the matching legacy V1 alias.
- Unauthenticated Claude refreshes still scan local logs. Spend tiles
render under the existing Not logged in notice when those logs contain
usage; an empty machine stays a hard error card.
- Grok scans every durable `updates.jsonl` ledger. Prompt-id dedup still
drops forked parent replays. `summary.json` is no longer required to
keep a ledger.
- Codex maps `self_serve_business_prolite` to **Business Premium**.
- Cursor widget IDs are unchanged. Titles/labels become Cursor Models /
Other Models to match Cursor's dashboard.
- Pricing supplement: GPT-6 Astra (OpenAI card, 2× fast), Gemini 3.8
Flash (Cursor table, $3.50 output), Fable 5.1, GLM 5.3, and `grok-bot-*`
→ Grok 4.6.

## Heads-up

Reviewed and **not** ported, with reasons:

- **openusage robinebers#1116 / robinebers#1127 / robinebers#1185** (analytics ping, PostHog) — Runway
removed analytics in #9.
- **openusage robinebers#1111 / robinebers#1136 / robinebers#1106** (scroll / Settings lag / SVG
parse) — Runway already has `ReorderFrameStore`, parsed-once
`ProviderMark`, and the rebuilt popover path. Taking their patch would
duplicate that work.
- **openusage robinebers#1137 / robinebers#1165 / robinebers#1141** (Codex Session default, Fable
order) — Runway already hides Codex Session by default and already
places Fable directly below Weekly. Layout defaults stay an owner
decision.
- **openusage robinebers#1134 Grok Bot meter** — new Cursor metric. AGENTS.md
requires owner confirmation of the four defaults before adding it; this
PR only takes the dashboard label rename and the `grok-bot-*` pricing
aliases.
- **openusage robinebers#1139** (Antigravity local spend) — new scanner, protobuf
decoder, and new metrics. Too large for this wave and needs the same
default-placement call.
- **openusage robinebers#1195** (OpenCode Codex OAuth attribution) — new scanner
sharing Codex request pricing. Real feature, own follow-up; folding it
in here would bloat the PR.
- **openusage robinebers#1164** (Claude multi-account) — Runway already discovers
Claude homes and gives each account its own card.
- **openusage robinebers#1177** (Codex fallback pricing Settings) — extra Settings
surface; earlier port waves skipped extra reset/settings chrome for the
same reason.
- **openusage robinebers#1179** (dead pin ID remap) — OpenUsage layout keys and
old Antigravity IDs. Runway installs never held those keys (different
defaults domain), and schema v3/v4 are already used for the beta-channel
and telemetry retirements.
- **openusage robinebers#1172** (bound log memory) — Runway already rejects
non-finite / overflowing token counts at the parse boundary instead of
clamping them.
- **openusage robinebers#1167 / robinebers#1016** (sub-1% "Not started", untouched pacing) —
already in Runway (`used <= 0`, `Pace.evaluate` returns nil when
unused).
- **openusage robinebers#1128** (Sparkle 2.9.6) — still a relevant bump; leaving
it to Dependabot rather than mixing a package-resolution change into
this accuracy PR.
- **openusage robinebers#1170 / robinebers#1159 / robinebers#1143 / robinebers#1163** (contribution policy,
screenshot assets, test-suite cleanup) — not user-facing on Runway, or
would churn tests without changing behavior.
- **openusage robinebers#1196** (legacy Codex iCloud identity) — Runway's sync
identity path is already fork-specific.

Gemini 3.8 Flash output is **$3.50**, from [Cursor's
table](https://cursor.com/docs/models-and-pricing.md), not upstream's
$3.75 Google API rate. That matches how Runway priced Gemini 3.7.

## Tests

- Desktop: prefixed key for the signed-in account wins; foreign `acct:`
keys are ignored; a scoped V2 tombstone suppresses the V1 alias;
`load()` reads `lastKnownAccountUuid`.
- Claude: no credentials plus local logs → spend tiles and Not logged
in, not an error card. Empty machine still errors.
- Grok: subagent ledger is included; fork replay of a shared prompt
still counts once.
- Codex: `self_serve_business_prolite` → Business Premium, weekly-only
window.
- Pricing: Astra / 3.8 Flash / Fable 5.1 / GLM 5.3 / grok-bot slugs and
router labels resolve. `testEveryAliasCanonicalResolves` covers the new
rules.
- Cursor mapper tests updated to the new labels; widget IDs unchanged.

`swift test --filter
"ClaudeDesktopAuthStoreTests|ClaudeProviderTests|CursorProviderTests|CursorUsageSummaryTests|GrokLogUsageScannerTests|CodexUsageMapperTests|PricingBundledResourceTests|LayoutStoreTests"`
— 179 tests, 1 skipped, 0 failures.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Grok spend tiles skip subagent session ledgers even when the parent turn does not include child usage

1 participant