Replies: 1 comment
|
Answered this together with #1312, since the two are the same arc and the questions there are the concrete end of it: #1312 Short version of the position, so this thread carries it too: keep the two origin fields, materialize the effective value into Design discussion continues on #1312. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Routing this here as a design RFC, following the same pattern as the recent ones. Full draft lives at
docs/rfc/rfc-quality-model.mdon my fork; this is the summary for discussion.Problem
quality_scoreis a single overloaded field. Six sources write to it (AI scorer, implicit scorer, async batch scorer, manual rating via MCP + HTTP twins, decay association-boost, Milvus conflict resolution) and ten consumers read it (decay multiplier, forgetting retention, four rerank weights, graph sort, analytics tiers, bootstrap ranking) — none can tell where the number came from.The concrete failure: when a human rates a memory, the rating is blended (
0.6*user + 0.4*existing) into the same field the machine computes. Human opinion and machine guess overwrite each other — the rating gets diluted or lost, or the machine's later score erases the human's judgment. One field, two authors scribbling over each other.Empirical state on my deployment (22,709 memories): 246 have
quality_score(1.1%), zero havequality_provider, 138 haveuser_rating. In practice effective quality today is human rating + a 0.5 default — the AI scorer has effectively never persisted here. So the overload is currently masked, but it becomes active the moment the scorer runs and then a rating lands on top.Proposal (decisions taken, open questions below)
computed_quality(machine, keepsquality_provider/quality_components) anduser_rating(human) as distinct metadata fields, each preserving its own history. Additive — no schema change.0.6/0.4blend is removed.NEVER do Xlesson may earn a 👎 yet must survive). Accumulated 👎 do lead to forgetting.metadata.get('quality_score')raw (including a ranking path and two that operate on plain dicts with noMemoryobject), so a property-based projection would miss them. Writing the composed value intoquality_scoreat rate/score time means every existing reader gets it with zero consumer changes. The origin fields stay separate for reversibility.quality_providerstays strictly "which engine computed"; rating provenance lives in its own field, never masquerading as a provider (which would silently degrade toimplicitin the sync codec's closed vocabulary).Open questions I'd like your read on
-1/0/+1. How should it map onto the0..1scale decay/forgetting/ranking consume?(rating+1)/2makes 👎 = 0.0 exactly, which under "human wins" drops it to the retention floor. A floor above 0 for 👎? A dedicated curve? This feeds the retention thresholds you own.rating_history(N distinct negatives? net-negative sum? window?) to fitforgetting.py?quality_score, given the raw-metadata readers?compress_metadata_for_synccurrently dropsuser_rating; to sync the human opinion across hosts it needs a positional CSV field gated by a codec version bump. OK to coordinate that?Prior art I'm building on, not reinventing: the self-service RFC's "confidence + provenance first-class, separate from the raw score" (§4), and the
original_quality_before_boostpreservation precedent inconsolidation/decay.py. Happy to split implementation into small PRs once the shape is agreed. This depends on landing before the AI scorer is switched on, so computed and rating don't fight.All reactions