UN-3883 [FEAT] Cut dashboard metrics cron DB load: narrower source windows, monthly from daily, two new indexes - #2276
Conversation
…from the daily tier (#2255) * UN-3973 Derive monthly metrics from the daily tier, narrow source window to 2 days The dashboard aggregation widened its DAY-granularity query to the first of the previous month so monthly buckets could be summed in Python from the same rows. Every run re-read 32-62 days of source data per metric, per org, 96 times a day. Monthly is now rolled up from event_metrics_daily in one statement for all orgs, so the source queries only need the daily window. That window drops to 2 days, sized against the measured worst created_at -> terminal-status lag of ~2h. A once-daily 7-day pass reruns the same task at a wider bound to repair gaps left by cron downtime. The active-org prefilter is decoupled from the daily window and pinned at 7 days: metrics filtered on another column (hitl_completions on approved_at) can land for an org whose executions are older than the source window. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 Address Sonar and Greptile review findings Sonar: - S117: rename apps.get_model() locals in 0004 to snake_case - S3776: cut _run_aggregation cognitive complexity from 22 by hoisting the static metric config tables to module level and extracting the per-org body, the active-org prefilter and the result shape into helpers Greptile: - Monthly rows in the rebuilt window whose daily rows are gone are now deleted alongside the upsert, so the two tiers cannot disagree. An empty daily tier still short-circuits, so a wiped tier cannot cascade into deleting monthly history. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 Add tests covering the source window, rollup SQL and reconciliation schedule Closes the acceptance criteria that had no automated check: - the monthly rollup issues no source-table SQL, asserted by capturing the queries it actually sends - the window ladder at 2 / 7 / 62 days, including a row that finishes after the narrow window has moved past its created_at and so never re-enters it - the reconciliation schedule row, its idempotency and its reverse The schedule tests call the migration's function directly. The suite runs with --no-migrations, so data migrations never execute and asserting on the beat row would fail regardless of the migration being correct. Also moves the dotenv load in settings/base.py above the Celery block. CELERY_BROKER_BASE_URL, _USER and _PASS were read above it, so they could not be supplied by an env file at all and had to be ambient. Ambient values still take precedence, so deployed behaviour is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 Trim comments in tasks.py and revert the unrelated settings change Cut the verbose comments and docstrings down to the purpose and the non-obvious bits. Code is unchanged. Restore backend/settings/base.py to main — moving the dotenv load ahead of get_required_setting was a local test convenience, not part of this change. The test rig exports the broker vars itself, so CI never needed it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 Renumber the reconciliation migration to 0005 UN-3445 landed 0004_pg_periodic_tasks on main after this branch was cut, leaving dashboard_metrics with two 0004s depending on 0003 and nothing depending on either. Django saw two leaf nodes and refused to build the graph, so `migrate` failed before applying anything — every app, not just this one. Depend on 0004_pg_periodic_tasks and renumber to match, so the short prefix form stays usable for a rollback. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 [FIX] Address review: PG transport, scoped orphan sweep, retry posture The reconciliation row could not run on the PG transport — two functions share the task name dashboard_metrics.aggregate_from_sources and the worker one took no arguments, so the mirrored row dispatched source_window_days into a zero-arg function and the message was dropped. The worker proxy and the internal endpoint now plumb it, and 0005 declares the PG twin rather than leaving the mirror to invent one. The orphan sweep is scoped to the (organization, month) partitions the rollup actually produced: an incomplete daily tier passed the empty-tier guard and deleted monthly rows it could not vouch for. Its deletion count now reaches the task result and a WARNING. DatabaseError and OperationalError propagate from the monthly rollup so the configured autoretry fires, instead of being logged once behind success: True. The prefilter is never narrower than the query window, so a widened source_window_days cannot skip the orgs it exists to repair. bulk_create takes an explicit batch_size. The Beat/PG drift guard named 0002 and 0004, so it kept comparing three schedules against three while this PR added a fourth. It now discovers every migration in the app, replays their RunPython forwards in order, derives the Beat cadence from the schedule row, binds every declared kwarg to its task signature, and asserts every post-install Beat write bumps PeriodicTasks.last_update. The reconcile row moves to 04:40 UTC: */15 fires at minute 0, and the per-tier lock keys do not block each other, so 04:00 started two full aggregations at once. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 [FIX] Address Athul's review: upsert-only monthly, per-schedule lock, honest success The orphan delete is removed. The design agreed on this ticket (comments 44768/45016) is a pure INSERT ... ON CONFLICT DO UPDATE; the delete was scope beyond it, and it converts a recoverable undercount into unrecoverable loss — the daily rows that would rebuild a deleted monthly row are exactly the ones that were missing. A stale total is recoverable with backfill_metrics. The reconciliation pass no longer shares a lock key with the 15-minute schedule. The 15-minute row is an IntervalSchedule and drifts against a fixed crontab, so on a shared key the once-daily repair loses the race roughly one day in seven, returns skipped=True and is never retried. A run in which every metric for every org failed no longer reports success: True. The result's success now reflects the error count, the completion log rises to WARNING, and the worker-side guard reads skipped_reason and errors as well as skipped — it saw none of these three did-nothing shapes before. A failed monthly rollup is distinguishable from an empty one: upserted=0 collided with the legitimate no-op and the no_active_orgs return. source_window_days is validated and bounded. It arrives as JSON from a Beat row that is editable in the admin: negative puts the window in the future, 0 never refreshes yesterday, 365 restores the multi-month scan this ticket exists to remove. Tests: a golden test seeds source rows, lets the real aggregation populate daily, and compares the rolled-up monthly against the pre-change derivation computed independently from get_documents_processed — AC-4 was claimed Met and had no equivalence assertion. Fixture offsets derive from the month boundary rather than fixed day counts, which land in the wrong month for the last days of any month. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ecution on (status, created_at) (#2264) * UN-3972 [PERF] Index workflow_file_execution on (status, created_at) The dashboard metrics cron's documents_processed and failed_pages queries filter this table on status + a created_at window, but all four existing indexes lead with workflow_execution_id. With no entry point here the planner drives top-down from the org and sequentially scans all 1.28M rows of workflow_execution — 83% of the cron's DB time on production. Built CONCURRENTLY with atomic = False; a plain AddIndex would hold a SHARE lock over a 3.4GB table taking live inserts. Guarded against a leftover INVALID index from an interrupted build, which IF NOT EXISTS would otherwise keep silently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3972 [PERF] Trim the migration docstring to the project ceiling The docstring restated the prod plan, deployment runbook and recovery steps. That detail belongs in the PR, not in a file every future agent scans. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3972 [PERF] Guard the index migration's non-atomic CONCURRENTLY shape with tests The suite runs with --no-migrations, so 0007 is never executed in CI. Regenerating it with makemigrations, or dropping atomic = False / CONCURRENTLY while tidying, would land a plain AddIndex — a SHARE lock held for the whole build on a 3.4 GB table that takes live inserts — with every test still green. Five DB-free assertions on the migration module and the model's Meta.indexes: non-atomic, concurrent in both directions, the INVALID-index guard present, AddIndex confined to state_operations, and model/migration agreement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3972 [FIX] Assert the index definition, not just its validity, and pin reversibility The CREATE INDEX CONCURRENTLY IF NOT EXISTS matches on name alone, so a hand-built index with different columns was kept while Django recorded (status, created_at) into model state — a permanent, invisible divergence that makemigrations --check cannot see. The guard now compares pg_get_indexdef against the expected btree definition and qualifies the lookup by current_schema(), since app tables live in the unstract schema. Also pin that every database_operation is reversible: dropping the guard's reverse_sql=noop killed the whole rollback path with all five tests still green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3972 [FIX] Address Athul's review: guard semantics, whole-migration assertions Three mutations that were green are now caught: appending a bare AddIndex after the SeparateDatabaseAndState (a real lock-taking build on a 3.4 GB table, invisible because every assertion read operations[0]); flipping the guard's NOT indisvalid polarity, which either raises on every healthy deploy or never fires at all; and a typo in reverse_sql, which makes rollback a silent no-op through IF EXISTS while Django unapplies the migration. The CREATE assertion matches the column order by regex instead of an exact byte sequence — removing one space used to fail it, a false-failure mode whose only outcome is someone loosening the assertion. Docstrings: the plan citation now points at UN-4045, which supersedes the earlier workflow_file_execution reading; "every existing index leads with workflow_execution_id" was false (the PK leads with id); the exact CREATE statement an operator should run out of band is spelled out, with a warning off the struck two-index variant; and the models.py comment no longer implies the index fixes both cron queries when it fixes one until UN-3973 narrows the window. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…edule by tier and indexing workflow_execution on created_at (#2265) * UN-3973 Derive monthly metrics from the daily tier, narrow source window to 2 days The dashboard aggregation widened its DAY-granularity query to the first of the previous month so monthly buckets could be summed in Python from the same rows. Every run re-read 32-62 days of source data per metric, per org, 96 times a day. Monthly is now rolled up from event_metrics_daily in one statement for all orgs, so the source queries only need the daily window. That window drops to 2 days, sized against the measured worst created_at -> terminal-status lag of ~2h. A once-daily 7-day pass reruns the same task at a wider bound to repair gaps left by cron downtime. The active-org prefilter is decoupled from the daily window and pinned at 7 days: metrics filtered on another column (hitl_completions on approved_at) can land for an org whose executions are older than the source window. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 Address Sonar and Greptile review findings Sonar: - S117: rename apps.get_model() locals in 0004 to snake_case - S3776: cut _run_aggregation cognitive complexity from 22 by hoisting the static metric config tables to module level and extracting the per-org body, the active-org prefilter and the result shape into helpers Greptile: - Monthly rows in the rebuilt window whose daily rows are gone are now deleted alongside the upsert, so the two tiers cannot disagree. An empty daily tier still short-circuits, so a wiped tier cannot cascade into deleting monthly history. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 Add tests covering the source window, rollup SQL and reconciliation schedule Closes the acceptance criteria that had no automated check: - the monthly rollup issues no source-table SQL, asserted by capturing the queries it actually sends - the window ladder at 2 / 7 / 62 days, including a row that finishes after the narrow window has moved past its created_at and so never re-enters it - the reconciliation schedule row, its idempotency and its reverse The schedule tests call the migration's function directly. The suite runs with --no-migrations, so data migrations never execute and asserting on the beat row would fail regardless of the migration being correct. Also moves the dotenv load in settings/base.py above the Celery block. CELERY_BROKER_BASE_URL, _USER and _PASS were read above it, so they could not be supplied by an env file at all and had to be ambient. Ambient values still take precedence, so deployed behaviour is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 Trim comments in tasks.py and revert the unrelated settings change Cut the verbose comments and docstrings down to the purpose and the non-obvious bits. Code is unchanged. Restore backend/settings/base.py to main — moving the dotenv load ahead of get_required_setting was a local test convenience, not part of this change. The test rig exports the broker vars itself, so CI never needed it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3974 [PERF] Split the dashboard metrics schedule by tier and index workflow_execution on created_at Schedule split. One schedule ran every 15 minutes and wrote all three metric tiers. Dashboard daily and monthly figures do not need 15-minute freshness, so they move to hourly — 96 runs a day becomes 24 for the expensive DAY-granularity half of the work, while the hourly tier keeps its cadence. Both schedule rows point at the same task and differ only in a `tier` kwarg; a second task name would need its own worker registration and internal endpoint for the PG path. The lock is now keyed per tier, so the two runs that collide at the top of every hour do not starve each other. Omitting `tier` still writes all three tiers, so a manual trigger never silently writes nothing. Prefilter index. The active-org prefilter measures 1,849ms per call on production — the slowest single query on the instance. Nothing on workflow_execution leads with created_at: the two composite indexes are date-ordered only within one workflow or pipeline, and the partial index is empty in steady state. The split raises this query's call count, and UN-4045 will leave three more metric queries on the same bare date-range shape, so the index lands with the split rather than after it. Built CONCURRENTLY with atomic = False and guarded against a leftover INVALID index, matching migration 0026. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3974 [PERF] Keep scheduler ownership out of the split migration and cut _run_aggregation's complexity Migration 0005 used update_or_create for the PG row of the schedule it was only re-keying, which reset pg_owned to False. converge_pg_scheduler disables a row's Beat twin when the PG scheduler adopts it, so on an adopted deployment the migration would have left the aggregation with no firer at all — Beat disabled, PG no longer owning it. It now updates only task_kwargs on that row, leaving enabled and pg_owned to the scheduler that owns them. Rollback is symmetric. Threading the tier through _run_aggregation took its cognitive complexity from 25 to 27 against a limit of 15. Extracted _collect_org_metrics and _aggregate_org, and hoisted the two static metric tables to module level so they are not rebuilt per call. Names match the same extraction on #2255 so the two reconcile cleanly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3974 [PERF] Trim comments and docstrings to the project ceiling The two index migrations carried 50-60 line docstrings restating the prod plan, deployment runbook and recovery steps. That detail belongs in the PR, not in files every future agent scans. Cut to purpose and key behaviour. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3974 [PERF] Cover all three acceptance criteria with tests The suite runs with --no-migrations, so neither 0005 nor 0029 ever executes in CI, and nothing pinned the schedule split's behaviour at all. 46 tests, at least one per acceptance criterion. AC-1 — cadence, and the tier reaching the task. 0005 creates one row and rewrites one, both scheduler tables agreeing, and the rewrite touches neither pg_owned nor enabled: on an adopted deployment converge_pg_scheduler has already disabled the Beat twin, so handing ownership back would leave the aggregation with no firer. Separately the internal endpoint and the worker proxy are pinned to carry `tier` — that leg fails silently, since _call_internal builds a body only when a tier is given and the existing worker test called the task without one. AC-2 — the split changes no figure. Runs the real _run_aggregation three times and diffs the metrics tables: `hourly` reproduces the pre-split hourly figures exactly, and hourly + daily_monthly reproduce every row `all` writes. Two guards keep it from going vacuous, the second because mutation testing caught the first version passing while _aggregate_single_metric was broken — the fixture produced only LLM metrics, leaving half the split unverified. AC-3 — the index. Migration shape (non-atomic, CONCURRENTLY both directions, the INVALID guard, AddIndex confined to state_operations), plus an integration test that EXPLAINs the query the aggregation actually issues, captured rather than rewritten: a hand-copied queryset would keep passing after the prefilter changed, which is the one thing it is for. Rows are inserted in ascending created_at order so the heap matches production's append order. The Query Insights half of AC-3 is a production reading and is deliberately not faked here. Every test verified to fail when the thing it guards breaks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 Renumber the reconciliation migration to 0005 UN-3445 landed 0004_pg_periodic_tasks on main after this branch was cut, leaving dashboard_metrics with two 0004s depending on 0003 and nothing depending on either. Django saw two leaf nodes and refused to build the graph, so `migrate` failed before applying anything — every app, not just this one. Depend on 0004_pg_periodic_tasks and renumber to match, so the short prefix form stays usable for a rollback. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3974 Renumber the schedule-split migration to 0006 behind UN-3973's 0005 UN-3445's 0004_pg_periodic_tasks is the parent of both this migration and UN-3973's reconciliation migration, so landing both would leave dashboard_metrics with two leaf nodes and no applicable graph. Depend on 0005_add_reconciliation_task instead, which puts the intended merge order (UN-3973 then UN-3974) in the graph rather than in the merge queue. This branch cannot migrate on its own until UN-3973 lands; its tests are unaffected, since the suite runs with --no-migrations. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 [FIX] Address review: PG transport, scoped orphan sweep, retry posture The reconciliation row could not run on the PG transport — two functions share the task name dashboard_metrics.aggregate_from_sources and the worker one took no arguments, so the mirrored row dispatched source_window_days into a zero-arg function and the message was dropped. The worker proxy and the internal endpoint now plumb it, and 0005 declares the PG twin rather than leaving the mirror to invent one. The orphan sweep is scoped to the (organization, month) partitions the rollup actually produced: an incomplete daily tier passed the empty-tier guard and deleted monthly rows it could not vouch for. Its deletion count now reaches the task result and a WARNING. DatabaseError and OperationalError propagate from the monthly rollup so the configured autoretry fires, instead of being logged once behind success: True. The prefilter is never narrower than the query window, so a widened source_window_days cannot skip the orgs it exists to repair. bulk_create takes an explicit batch_size. The Beat/PG drift guard named 0002 and 0004, so it kept comparing three schedules against three while this PR added a fourth. It now discovers every migration in the app, replays their RunPython forwards in order, derives the Beat cadence from the schedule row, binds every declared kwarg to its task signature, and asserts every post-install Beat write bumps PeriodicTasks.last_update. The reconcile row moves to 04:40 UTC: */15 fires at minute 0, and the per-tier lock keys do not block each other, so 04:00 started two full aggregations at once. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3974 [FIX] Address review: Beat reload, inherited ownership, boundary validation Rewriting live Beat rows through historical models fires no post_save, so DatabaseScheduler never reloaded: the existing row kept firing with no tier and the new row never fired at all. 0006 now bumps PeriodicTasks.last_update in both directions, as scheduler/ownership.py and mirror_pg_periodic_tasks.py already do. The new row inherits pg_owned and both enabled flags from the row it is split from instead of hardcoding Beat. In a PG-adopted environment the daily and monthly tiers had no firer at all while the hourly run still returned success. It also moves to minute 20. Minute 0 collides with */15 — and so does the suggested minute 30, since */15 fires at :00 :15 :30 :45 — and the per-tier locks are built so the two runs cannot block each other. An unrecognised tier is now rejected in post(), and the blanket except ValueError in _run is gone, so a ValueError from inside the aggregation reaches the logged 500 path rather than reading as a bad request body. Tests: the JSON round-trip assertion was a stdlib tautology that never read what the migration writes — replaced with an assertion on the updated row, the one firing the hourly tier in production. The ALL default is pinned off inspect.signature. The RunSQL table and column are derived from the model rather than grepped. The planner-choice assertion is deleted: a cost model on 12,000 rows is not production evidence. Also drops a full Organization count that ran on every tier for one log field, and gives two test modules the Django bootstrap their siblings carry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3973 [FIX] Address Athul's review: upsert-only monthly, per-schedule lock, honest success The orphan delete is removed. The design agreed on this ticket (comments 44768/45016) is a pure INSERT ... ON CONFLICT DO UPDATE; the delete was scope beyond it, and it converts a recoverable undercount into unrecoverable loss — the daily rows that would rebuild a deleted monthly row are exactly the ones that were missing. A stale total is recoverable with backfill_metrics. The reconciliation pass no longer shares a lock key with the 15-minute schedule. The 15-minute row is an IntervalSchedule and drifts against a fixed crontab, so on a shared key the once-daily repair loses the race roughly one day in seven, returns skipped=True and is never retried. A run in which every metric for every org failed no longer reports success: True. The result's success now reflects the error count, the completion log rises to WARNING, and the worker-side guard reads skipped_reason and errors as well as skipped — it saw none of these three did-nothing shapes before. A failed monthly rollup is distinguishable from an empty one: upserted=0 collided with the legitimate no-op and the no_active_orgs return. source_window_days is validated and bounded. It arrives as JSON from a Beat row that is editable in the admin: negative puts the window in the future, 0 never refreshes yesterday, 365 restores the multi-month scan this ticket exists to remove. Tests: a golden test seeds source rows, lets the real aggregation populate daily, and compares the rolled-up monthly against the pre-change derivation computed independently from get_documents_processed — AC-4 was claimed Met and had no equivalence assertion. Fixture offsets derive from the month boundary rather than fixed day counts, which land in the wrong month for the last days of any month. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc * UN-3974 [FIX] Address Athul's review: lock covers what is written, both kwargs, graph guard The lock is keyed by granularity written, not by enum member. ALL took a third key that excluded nothing, so an ALL run and the scheduled hourly run wrote EventMetricsHourly concurrently — reachable from the documented manual trigger and from the endpoint's own "omit tier" contract. ALL now takes both keys and releases whatever it took if it cannot take them all. Keys are namespaced by source window so the once-daily reconciliation pass, which is never retried, is not starved by the 15-minute schedule. source_window_days is accepted on all three legs. 0006 hard-depends on 0005, so the reconciliation row is a certainty rather than a hypothetical, and this branch's signatures rejected the kwarg it dispatches. The tier predicates come from one membership table, so a member added without an entry raises instead of acquiring the lock, iterating every org, writing nothing and returning success. The migration's bulk updates check their row counts. A filtered update matching nothing reported success while leaving the old row on kwargs="{}" — every tier every 15 minutes — alongside the new hourly row: strictly more load than before, silently. tier is validated at the request boundary with a warning log, and an explicit null is treated as omitted. New test_migration_graph.py builds the migration graph, which is what catches 0006's dependency on a node that is not on this branch; --no-migrations means nothing else does. Lock behaviour is now exercised rather than its key string asserted, merge_schedules has coverage at all, the equivalence file carries one absolute expectation and a frozen clock, and the prefilter asserts the index is usable under enable_seqscan=off rather than that the planner chose it on 12,000 synthetic rows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGZBF68CShem3pbUJM2tBc --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
| Filename | Overview |
|---|---|
| backend/dashboard_metrics/tasks.py | Splits aggregation by tier, narrows source windows, introduces per-tier locking and active-organization filtering, and derives monthly metrics from daily rows. |
| backend/dashboard_metrics/migrations/0005_add_reconciliation_task.py | Adds the daily seven-day reconciliation schedule for both Celery Beat and PostgreSQL scheduling. |
| backend/dashboard_metrics/migrations/0006_split_aggregation_schedule.py | Splits the existing aggregation schedule into frequent hourly-tier and hourly daily/monthly-tier jobs while preserving scheduler ownership. |
| backend/dashboard_metrics/internal_views.py | Validates and forwards aggregation tier and source-window arguments through the internal task endpoint. |
| workers/scheduler/dashboard_metrics_tasks.py | Carries scheduler task arguments through the worker proxy to the backend aggregation endpoint. |
| backend/workflow_manager/file_execution/migrations/0007_wfe_status_created_idx.py | Adds the concurrent status-and-created-at index used by file-execution metric queries. |
| backend/workflow_manager/workflow_v2/migrations/0029_we_created_at_idx.py | Adds the concurrent created-at index intended to support active-organization filtering. |
Flowchart
%%{init: {'theme': 'neutral'}}%%
flowchart LR
S[Source metric tables] -->|Every 15 minutes| H[Hourly aggregation]
S -->|Hourly at :20| D[Daily aggregation]
S -->|Daily 7-day reconciliation| R[Hourly and daily repair]
D --> M[Monthly rollup from daily]
H --> API[Dashboard APIs]
D --> API
M --> API
Reviews (4): Last reviewed commit: "UN-3883 [FIX] Pin the second lock suite ..." | Re-trigger Greptile
| return set( | ||
| WorkflowExecution.objects.filter(created_at__gte=cutoff) | ||
| .values_list("workflow__organization_id", flat=True) | ||
| .distinct() |
There was a problem hiding this comment.
Prefilter drops valid metric activity
When an organization has recent PageUsage, Usage, or HITLQueue activity but no WorkflowExecution created within the lookback, _active_org_ids excludes it from every metric query, causing its pages-processed, LLM, or HITL dashboard rows to remain stale or absent even after reconciliation.
Knowledge Base Used: Restore dashboard metrics and navigation
Prompt To Fix With AI
This is a comment left during a code review.
Path: backend/dashboard_metrics/tasks.py
Line: 631-634
Comment:
**Prefilter drops valid metric activity**
When an organization has recent `PageUsage`, `Usage`, or `HITLQueue` activity but no `WorkflowExecution` created within the lookback, `_active_org_ids` excludes it from every metric query, causing its pages-processed, LLM, or HITL dashboard rows to remain stale or absent even after reconciliation.
**Knowledge Base Used:** [Restore dashboard metrics and navigation](https://app.greptile.com/zipstack/-/custom-context/knowledge-base/zipstack/unstract/-/reverts/incident-mitigation_1813-20260302-dashboard-metrics-sidebar-440ae49.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.There was a problem hiding this comment.
@greptile-apps The finding is real, and the concrete case is Prompt Studio: an LLM call there writes a Usage row without ever creating a WorkflowExecution, so an org doing only Prompt Studio work in the lookback is excluded from llm_calls, challenges, summarization_calls, llm_usage, pages_processed and both HITL metrics.
It is pre-existing rather than introduced here, and this PR does not widen or narrow it.
At the merge base the prefilter is the same single table:
# origin/main, dashboard_metrics/tasks.py
active_org_ids = set(
WorkflowExecution.objects.filter(created_at__gte=daily_start)
.values_list("workflow__organization_id", flat=True).distinct()
)daily_start there was 7 days. This PR narrows the metric window to 2 days, which would have dragged the prefilter down with it — so _active_org_ids takes min(window_start, end_date - DASHBOARD_ACTIVE_ORG_LOOKBACK_DAYS) to floor it. Measured on the current head:
source_window_days= 2 -> cutoff 7 days back (floored at LOOKBACK)
source_window_days= 7 -> cutoff 7 days back (widened by the caller)
source_window_days=60 -> cutoff 60 days back (widened by the caller)
So the org set is identical to before the change, and widens with the reconciliation pass rather than being pinned.
Not fixing it in this PR, for two reasons. Widening the prefilter means unioning Usage, PageUsage and HITLQueue — three more queries per run on the hot path, on a change whose entire purpose is removing DB load from this cron (UN-3883: ~55 min of DB time per 6 h). And UN-4045 deletes the prefilter outright by grouping every metric query by organization, at which point the org set comes from the metric queries themselves and the gap closes without a union.
Tracking it there rather than papering over it here.
There was a problem hiding this comment.
Agreed. This is a real limitation, but it is pre-existing and not introduced by this PR. The min(...) floor preserves the previous 7-day prefilter for the normal 2-day run and correctly widens it for reconciliation, so this change does not regress the affected organizations. Given the hot-path load objective and the planned UN-4045 fix that removes the prefilter, I would not block this PR on the finding; it should be tracked as follow-up work there.
…e metrics docs Remediation pass over #2276. Fixes only — no aggregation behaviour changes. High: - test_aggregation_tier.py needed a live Redis but sat in the unit tier, which provides none: 8 errors, `test (unit)` red. Settings never override CACHES, so the suite inherits production django_redis. The lock cases now pin locmem, which has identical add/get/delete semantics; they run in CI for the first time. - backfill_metrics computed start_date without truncating to midnight while the cron truncates, so the oldest day was written as a partial day's bucket. The monthly rollup now sums the persisted daily tier rather than recomputing from source, so that short value became permanent once past the reconcile window. - 0006's rollback runbook named `migrate dashboard_metrics 0005` with no ordering. 0005 and 0006 do not exist in the previous release, so that command errors after the image rolls back; and 0005's own row carries source_window_days, which the old signature also rejects. Corrected to 0004, reversed before the image. Medium: - source_window_days > 90 passed the view and raised inside the task, returning 500 for a bad request. _int_arg now takes the task's own bound. - test_a_blocked_run_releases_whatever_it_took never executed the rollback it claimed to prove: keys sort daily_monthly first, so ALL failed on its first key and the rollback loop ran zero times. Reordered; verified it now fails without the rollback. - 0029's index guard checked validity only, so a hand-built (created_at DESC) index passed while Django recorded fields=["created_at"]. Lifted 0007's definition and current_schema() checks across, with tests. - Releasing the lock keys was unisolated: one cache.delete raise replaced a completed run's return value and stranded the remaining keys. - The "every 15 minutes" cadence was restated in five places the split made wrong. Removed the cadence from prose rather than restating it; it lives in 0005/0006. - The rollup docstring and its test claimed a deleted day leaves the monthly total in place. It does not — the group survives with a smaller sum. - README: the 7-day window is a lag ceiling, not only a downtime one; the hourly tier never self-repairs beyond 24h; added the deploy backfill at --days 62 (60 misses a day when run on the 31st). - docker-compose and the cloud chart both asserted the periodics never overlap and the Redis lock self-guards. The split made both false; comments corrected. Also: ruff 0.3.4 (the pinned gate) reported 13 errors and 8 unformatted files, all in tests this PR adds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Aja2YP7hoVWqocsUd6jCkN
…ation reverse on the operation Two siblings of fixes in 3a49d7479 that the first pass missed. - TestTheLockIsPerSchedule (test_tasks.py) takes un-namespaced lock keys on the real django_redis cache and brackets each test with cache.clear(), which django-redis implements as FLUSHDB — it wipes every key in that database, and the Celery broker shares db 0 in the test env. Pinned to locmem like its sibling in test_aggregation_tier.py. Verified: the class fails 4/4 without the override on an unreachable Redis and passes 4/4 with it. - test_it_builds_and_drops_concurrently greps the migration source, and the guard added in 3a49d7479 put "DROP INDEX CONCURRENTLY IF EXISTS" into two RAISE EXCEPTION messages — so the assertion held whatever reverse_sql was. Asserted on the rendered operation, as the 0007 sibling already does. Verified: mutating reverse_sql to RunSQL.noop now fails the test; before this it left all nine green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Aja2YP7hoVWqocsUd6jCkN
|
Unstract test resultsPer-group results
Critical paths
|



What
Three changes to the dashboard metrics cron, already reviewed and merged individually as #2255, #2264 and #2265:
Why
The cron was using roughly 55 minutes of database time every 6 hours on production. It re-read between 32 and 62 days of raw data for all 38 organisations, 96 times a day, to produce figures that only change once a month.
Nothing a customer sees changes. Hourly dashboard numbers are still at most 15 minutes old; daily and monthly numbers are now at most an hour old instead of 15 minutes, which is the one deliberate trade.
How
_rollup_monthly_from_dailyreplaces the per-org monthly source queries with oneINSERT ... ON CONFLICT DO UPDATEoverevent_metrics_dailyfor every org. Upsert-only, per the design agreed on UN-3973: a monthly row the daily tier no longer produces is left in place, because a stale total is recoverable withbackfill_metricsand a deleted one is not.aggregate_metrics_from_sources(tier, source_window_days)is called by three schedule rows. Both kwargs travel on both transports — Beat calls the Django task directly, the PG scheduler goes worker proxy → internal endpoint → the same function.ALLtakes both, so it genuinely excludes a concurrent hourly run; distinct windows are distinct jobs, so the once-daily reconciliation pass can never be starved by the 15-minute schedule it is never retried after.CONCURRENTLYunderatomic = False, withAddIndexconfined tostate_operations. Each migration asserts the index is valid and has the expected definition before recording itself applied.Can this PR break any existing features?
The realistic risks, and what bounds each:
event_metrics_daily. Gaps shorter than 7 days repair themselves on the next reconciliation pass; anything older needsbackfill_metrics. Bounded to the current and previous month.0006is reversed — the schedule rows carry atierkwarg the previous release's signature rejects. Documented in the migration docstring;migrate dashboard_metrics 0005restores it.tierand raisesTypeErroruntil it rolls. Self-healing, and no aggregation is lost because the next tick succeeds.Database Migrations
Four, in three apps. No schema changes to any metrics table.
dashboard_metrics/0005_add_reconciliation_taskdashboard_metrics/0006_split_aggregation_schedulefile_execution/0007_wfe_status_created_idxworkflow_file_execution (status, created_at)workflow_v2/0029_we_created_at_idxworkflow_execution (created_at)Both index migrations no-op via
IF NOT EXISTSif the index was built out of band first, which is the preferred production path — the exact statement is in each migration's docstring.Env Config
None.
Deploy Steps
Run once, before the first aggregation after deploy:
Monthly is now derived from the daily tier, so that tier has to be complete across the rollup window first.
--skip-monthlyis deliberate: repair daily and let the rollup derive monthly.Relevant Docs
backend/dashboard_metrics/README.mdis updated — schedules, windows, staleness bounds, and the ownership overlap withbackfill_metrics.Related Issues or PRs
Merged into this branch: #2255 (UN-3973), #2264 (UN-3972), #2265 (UN-3974). Parent: UN-3883.
Dependencies Versions
No dependency changes.
Notes on Testing
event_metrics_daily, not source tablesindisvalid = tCONCURRENTLY, no write-blocking lockget_documents_processedfree of a seq scan onworkflow_file_executionget_failed_pagesfree of a seq scan onworkflow_executionget_recent_activityunder 1 s(created_at DESC)index descoped in comment 45015The two Confirm on prod rows are post-deploy readings against Query Insights, not outstanding work.
Screenshots
Not applicable — no UI change.
Checklist
I have read and understood the Contribution Guidelines.