Replies: 9 comments 1 reply
|
@LING-6150 Thanks — this is a solid RFC. Approving the overall direction; let's move to implementation. The three sections you flagged
Open questions
Next step: PRs Please go ahead with one PR per phase against
For the migration PR, include a real run against a database with existing materials (old sessions plus new uploads side by side) in the PR description, not just unit tests. Keep each PR reviewable on its own and link them back to this discussion. |
|
@wyuc Thanks! Tracking issue: THU-MAIC/OpenMAIC#1736. Noted on all points. Empty-folder deletion is now in the first batch (page-only, in Phase 2, with the UI in Phase 4), and Phase 1 is split into 1a (manifest, folders, asset_root_refs, lifecycle checks and claims, including removing the DROP COLUMN asset_id bootstrap statement), 1b (owner-level extraction and cache) and 1c (migration, with a real run against a database that has old sessions and new uploads side by side). Each PR will link back here. Starting with 1a. One question for 1a: I plan to keep 1a to schema, roots and lifecycle checks only, and start writing new uploads into the pool in 1c together with the pool-first read path and the backfill, so no producer lands before its reader. Does that split work for you? |
|
@LING-6150 Thanks for the follow-ups. Decisions on everything open, plus two scope cuts: 1. 1a stays schema-only. Agreed: 1a adds the schema, 2. No undo and no result cards in the first batch. Library changes are low-risk (the agent cannot delete, and every create/move/rename is easy to reverse by hand or by asking the agent), so undo is not worth its cost right now. That removes the operation records and revision counters from PR 2 and drops Phase 3 entirely. Agent changes are shown through the existing tool-call rows, with a clear one-line summary per tool (e.g. "Created folder 'Functions'", "Moved 2 materials to 'Functions'", "Started parsing 2 materials"). No custom card component and no folder links for now. Please update the RFC body and the phases in #1736: 1a / 1b / 1c → tools and 3. Land everything through an integration branch. Please create I checked your deployment concern against the collector: you are right. A pre-1a collector's unreferenced sweep stamps any committed entry with no
Also, in 1a please record a reference-rule version in the database and have the collector skip its entry level when the database's version is newer than the one it knows. It cannot protect this cutover, because old collectors do not know the check, but it means the next new reference kind will not need the same manual rule. 4. Naming. In the UI the feature is called Knowledge base (知识库 in Chinese). Code and APIs keep 5. Review. A review of #1745 is on its way. Please keep it open against the integration branch while that runs. |
|
@LING-6150 I've created |
|
@wyuc Thanks, all noted. Update:
|
|
Update: 1c is open as a draft: THU-MAIC/OpenMAIC#1763 · feat(materials): store uploads… (stacked on #1753). It is the first phase that writes material roots; the deployment boundaries and the backfill switch are in its description. One question : will integration/material-library be deployed anywhere before it merges into main? @wyuc |
|
@LING-6150 Thanks — great follow-up on the old-bootstrap boundary.
|
|
Update: Phase 1a (#1745), 1b (#1753), and 1c (#1763) are merged into integration/material-library. The three required follow-ups from the 1c review are also merged in #1791: losing backfill allocations are cleaned up, failed pool reads can fall back to retained old objects, and old-object reads verify the recorded digest. I'm preparing Phase 2 as one PR, following the current RFC. Owner extraction will be enabled only when the complete producer/consumer paths, including document-image ingestion and reference rewriting, are ready. One delivery-boundary question: §5 includes source deletion as a shared, page-only operation, but the delivery table explicitly names only empty-folder deletion/control. I propose delivering the shared source-deletion operation and HTTP route in Phase 2, with the confirmation UI in Phase 3. Phase 2 would therefore cover the real deletion transaction and late-worker concurrency tests. Does that match the intended boundary? If all source-deletion functionality should instead ship in Phase 3, Phase 2 will test deleted-row rejection only, and the real deletion transaction tests will ship with Phase 3. |
|
@LING-6150 Answered on #1792. In short: source deletion ships as its own small PR into the integration branch after Phase 2, as option A with a focused review and a real run, with the confirmation control in Phase 3. Phase 2 is split after commit 6 into two stacked PRs. All scope decisions are accepted. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Status: Accepted, in implementation · Area: materials, storage, agent runtime, workbench · Builds on: #1153, #1669 (Discussion #1657)
Updated 2026-10-01 with the decisions from this discussion: no undo or result cards in the first batch (§6), a revised phase plan delivered through
integration/material-library, the release conditions for roots, and the UI name (§10).Summary
This RFC gives uploaded files an owner-scoped library with folders, extraction reused across conversations, one set of operations shared by the page and the agent, and a one-line summary of each agent change in the conversation. Originals and extraction outputs are stored in the asset pool, with library entries as their reference roots. Composer uploads go into Unfiled and are attached to the conversation, and the composer's
@menu searches the whole library.In the UI the feature is called Knowledge base (知识库); code and APIs keep material and material library (§10).
The library is a page opened from the bottom of the left rail. It is not a tab next to Chats and Courses, nor a pane beside the chat. The agent can organize materials and start extraction;
store_materialis not given to the agent in the first batch, anddelete_materialis page-only.All statements about current behavior were checked against upstream
mainat1c70e86a. Everything else is a proposal.Problems and goals
Uploads are stored durably, but there is nowhere to manage them.
POST /api/materialsrecords anowner_materialrow and writes the file toMaterialByteStore. There are no folders and no rename or delete operation; thedeleted_atcolumn exists, but nothing writes it.GET /api/materialsrequires a session id and lists only that session's materials.Extraction cannot be reused across conversations.
bindOwnerMaterialsToSessioncopies the file's bytes into the session and creates a new session material id, and extraction runs on that session row. A second conversation starts from scratch.Library assets need to stay alive without a course. Today only
document_asset_refscommits a pending pool allocation and keeps it alive: a new allocation that no course names expires, and removing the last course reference starts collection grace. Library uploads need a reference root of their own.The goals are:
Relationship to #1153
#1153 defined the storage identity of materials. Its parts 0, 1, 2 and 4 were merged as #1154, #1164, #1168 and #1172. #1242 later states that the "asset registry keeps zero runtime write wiring" and that "legacy/inline conversion modules [were] deleted"; the manifest and the extraction cache were removed in the same change. #1524 brought workbench image and video generation back to the pool. Material uploads, extraction and
use_material_mediastill use byte stores.maintodayid@versionWhat is different from last time
In #1168, the manifest was a per-principal KV list keyed by content digest. It kept files from being collected by holding its own pool entries. #1242 removed it. This proposal is different in three ways:
First-batch scope and non-goals
In scope:
@;Material and course folders stay separate. Adding materials may use a dialog, and an original opens in a new tab or downloads.
Out of scope:
fetch_urlmaterials into the library;Design
1. Upload is ingest
Uploading from the composer or from the library page creates a source material in Unfiled (
folderId = null). A composer upload is also attached to the conversation. Choosing a material that is already in the library does not upload it again.These labels map onto the existing status values. Extraction queueing, leases and completion, which today live on session materials, move to the owner level. The details belong in the Phase 1 PR.
Upload and extraction stay separate states, so a stored file is never assumed to be searchable. A message cannot be sent while any of its files is still uploading, and the composer and the page use the same ingest and admission rules.
2. Manifest, folders, pool and roots
The owner manifest records identity and organization.
owner_materialgainsasset_id(the original in the pool), a nullablefolder_id, adisplay_nameseparate from the original filename, the current extraction attempt and state, and a pointer to the latest completed extraction. Material ids stay globally unique.Folders are stored as
{ owner_id, id, name, normalized_name, created_at, updated_at }, unique on owner and id and on owner and normalized name. They use the same name normalization as course folders. Unfiled is the absence of a folder, not a folder row. Images and audio tracks produced by extraction become owner material kinds linked byderivedFrom(§3).Every retained file goes into the pool. Originals, extraction artifacts, readable text or transcripts, and derived media are allocated under
assetPrincipalForOwner. Media embedded in a stored artifact is ingested separately and replaced by references to the allocated assets. Lineage (source, extractor version, and page or time when supplied) is recorded in the library, not in storage.Reference roots keep materials alive. The proposed table is
asset_root_refs(root_kind, root_id, asset_id), with a composite primary key, an index onasset_id, and a cascading foreign key toasset_entries. The library writes('material', materialId, assetId)rows. Storage does not interpretroot_kind, and because material ids are globally unique, root rows do not change when a claim moves the material.Claims move the whole library. The owner-material claim participant is extended to folders and extraction metadata; the existing assets participant already re-keys asset principals. Folder conflicts follow the course-folder claim rules. Schema, locking and claim details belong in the Phase 1 PR.
Invariants
fenceOwnerWrite. The transaction keeps the existing sorted entry-lock order and checks that every asset exists and belongs to the same owner.3. Extraction and stable ids
Extraction becomes reusable owner-level state.
extract_materialkeeps its "ensure extraction has started" semantics: idle and failed sources are queued; running and completed ones are left alone.wait_for_materialskeeps its bounded-wait semantics and gains an optional libraryscope(§5). Changing the extractor does not re-run completed sources; forced re-extraction needs a future, explicit policy.A source id reads its latest completed text. Today, reading a source returns
source_requires_derivative, and extraction creates separateextractionortranscriptmaterials with their own ids. Under this proposal, reading or searching by source id uses the latest successful extraction; if there is none yet, the call returns the extraction state and suggestsextract_material.Text and transcripts become artifacts of the source rather than separate visible materials; keeping internal derivative rows would also work (see the open question on storing extracted text). Media derivatives such as images and audio tracks keep their own ids, inherit the source's folder, and cannot be moved, renamed or deleted on their own;
list_materialsexposes their lineage and any page or time the extractor supplied.use_material_mediaaccepts media source and derivative ids. Reading image or audio materials still returns unsupported-kind guidance.The extraction cache is scoped per owner. Its key combines the owner, the content identity, the extractor
id@versionthat actually ran, and any result-affecting options. A cross-owner hit would reveal that someone else uploaded the same file, so there is no cross-owner lookup. The cache does not keep assets alive; a cache entry whose assets no longer exist or are no longer retained is a miss. Details belong in the Phase 1 PR.Invariants
4. Conversations and courses
Conversations attach materials without copying them. A new link table holds unique
(session_id, material_id)pairs. Composer upload, the paperclip picker and@all attach the same way: a selection is attached when the message is sent (removing it before sending attaches nothing), and attaching the same material twice has no extra effect. The@menu searches all folders, shows each material's name, folder and state, skips files still uploading, and marks attached materials as selected. A course pick keeps its current meaning as the course target for the turn.After a linked material is deleted, later reads through the link fail, though text already in the conversation stays; the delete confirmation says so (see the open question on conversations after deletion).
The agent prompt follows source ids.
sessionMaterialsPromptBlockis rewritten around one flow: extract, wait, then read with the same id; it also covers derivatives, library scope, the organizing tools, and that deletion is left to the teacher. Every consumer of material ids, includingimport_pptx, resolves both new links and old session rows.Using a material in a course allocates a separate entry.
use_material_mediareturns a newly allocatedsrcunder the principal of the owner of the writable stage. The document write then records it indocument_asset_refs. This implements #1153 §6. Resolver and promotion details belong in the Phase 2 PR.A course-use entry is always separate, even for the same owner, so deleting the source cannot break the course. Preparing media does not by itself create a course reference: an entry that is never written into a course expires. Physical deduplication still counts against logical quota.
5. Shared operations and tools
The page and the agent call one implementation. HTTP routes and runtime tools are thin adapters that resolve the owner. Deciding how to classify materials is left to the agent and skills. Listing the owner's library and reporting limits are new API responses, not the behavior of today's
GET /api/materials.list_materials({ scope?, folderId?, query? })list_material_folders({ query? })create_material_folder({ name })→{ folderId, created }move_materials({ materialIds, folderId });nullmeans Unfiledrename_material({ materialId, name })/rename_material_folder({ folderId, name })store_material({ assetId, name?, folderId? })extract_material({ materialId, scope? })delete_material({ materialId })read_material({ materialId, offset?, scope?, revision? })search_material({ query, materialId?, scope? })wait_for_materials({ materialIds?, timeoutSec?, scope? })use_material_media({ materialId, stageId, scope? })scopedefaults to'session';'library'reaches the owner's unattached materials. In library listings, omittingfolderIdlists everything andnulllists Unfiled only;queryfilters by name and metadata. Uploads create their entry directly, withoutstore_material. Paging and adapter details belong in the Phase 2 PR.Every change goes through the owner write fence, and a move applies to all of its sources or to none. A library-scope wait must name its material ids. Literal search keeps its current limits and reports truncation. Creating a folder with an existing name returns that folder with
created: false, and a move or rename that changes nothing reports that nothing changed.6. Tool-row summaries
Each organizing or extraction call shows a one-line summary in its existing tool-call row, for example "Created folder 'Functions'", "Moved 2 materials to 'Functions'" or "Started parsing 2 materials". An extraction summary says whether work started, was already running, or had already completed. There is no custom card component and no link to a folder for now.
There is no undo in the first batch. Library changes are low-risk: the agent cannot delete, and every create, move or rename is easy to reverse by hand or by asking the agent. The first batch therefore has no operation records, revision counters or undo endpoint; undo can be added later if it turns out to be needed.
7. Navigation and refresh
The library page opens in the main area. It shows folders, card and list views, status, opening or downloading the original, and "chat with this material". Agent operations never change the folder the teacher is looking at.
During a run, material changes are reported through the existing
library_changedevent with a materials discriminator, and the client refetches. Extraction completes in the background, so outside a run the page also needs to pick up those results, as well as changes made in other tabs. Two options were considered, and the first batch uses B:/api/agent/owner-eventsalready streams durable session summaries with LISTEN/NOTIFY wakeups and a 30-second polling fallback. It would add a persisted library revision as an invalidation signal, reconciled on connect, reconnect, resync, degraded catch-up and claims; the library revision does not replace the session event replay cursor.Refreshing never navigates. In option A, the library revision is committed together with each change, and NOTIFY is only a wakeup, never the source of truth.
8. Limits and quotas
Existing limits stay configurable and are shown before upload. The owner library listing returns the settings and current usage:
OPENMAIC_AGENT_MAX_UPLOAD_BYTESMATERIALS_MAX_DOCUMENT_BYTESMATERIALS_MAX_COUNT_PER_OWNERMATERIALS_MAX_TOTAL_BYTES_PER_OWNERASSET_QUOTA_BYTES0disables itBoth the source quotas and the pool quota are kept (see the open question on quotas).
Checks in the UI are early hints; the server remains the authority. Derivatives count toward the pool quota but not the source-count quota, so extraction can fail after a successful upload, and that failure is reported. Error responses keep their current codes (
413for size,429for source quota).9. Compatibility and migration
Existing data stays readable until its migration commits. Ready uploads are migrated by a bounded backfill or on first access: allocate in the pool, publish the pointer and root, then remove the old bytes. Existing session materials keep their ids, copies and extraction rows, and a session can contain both old and new materials.
Schema, bootstrap, race handling and cleanup details belong in the Phase 1 PR; consumer compatibility belongs in the Phase 2 PR.
Two points matter for review. The
DROP COLUMN asset_idstatement in today's bootstrap must be removed, and an old process must not be able to run it again. Old bytes are removed only after the new pointer is confirmed, and nothing is deleted when a commit outcome is uncertain.10. Naming
In the UI the feature is called Knowledge base (知识库 in Chinese); code and APIs keep material and material library. Since the first batch has no semantic Q&A, the empty state and onboarding say plainly what it does: upload teaching materials, have the assistant file and parse them, and pick them with
@while preparing a course.Delivery phases
Every phase is its own PR to the integration branch
integration/material-library, in this order, and each PR settles the details this RFC leaves open. The branch merges intomainonce, at the end.asset_root_refsand lifecycle checks against both reference tables; claims; the reference-rule version; removal of theDROP COLUMN asset_idbootstrap statement@: operations, tools and scopes; attachment resolvers; revision-aware reads; the new prompt; tool-row summaries; invalidation and limits; course copy-on-use; the composer picker and@; empty-folder deletionPhase 2 lets the agent use the library before the page exists; no intermediate version links to a page that is not available yet.
Release conditions. Roots ship in the same release as 1a, and two hazards come from processes of an earlier release; the integration PR and the release notes state them:
ASSET_COLLECTION_GRACE_MS. Deployments that store users' original files may want a longer grace period than the one-hour default.DROP COLUMN IF EXISTS asset_idwhenever it starts, which drops every pool pointer at once. Once any new instance has stored a material in the pool, no older instance may start against the database: stop the older instances before the new ones take writes.Decisions
Decided in this discussion:
integration/material-library, which merges intomainonce.References
Feedback from anyone is welcome, especially on the reference-root table (§2) and the phases.
All reactions