Replies: 1 comment
|
Provider-native PDF/video input would be useful, but I think it needs to be added as a route capability rather than by widening the current image MIME list. In rc.2 the composer, wire admission, attachment store, durable message, retrieval response, and provider projection are all image-specific. A safe design would keep one immutable, Session-authorized local attachment as the source of truth, then let an exact provider route choose one of: native passthrough, bounded extraction, a hybrid of both, human-only retention, or an actionable unsupported-format result. For native passthrough, I would scope the remote binding to at least: That prevents a provider file ID from becoming a global durable identity or crossing tenant/route boundaries. The lifecycle also needs explicit PDF and video need different truth boundaries. For PDFs, record whether the route saw embedded text, OCR, or rendered pages and whether page citations cover text or images. For video, record frames/audio/subtitles and the provider sampling policy; a successful request must not imply every frame was observed. Provider URL ingestion, retention, deletion, residency, and SSRF controls also belong in the contract. I traced the current rc.2 boundaries and wrote the full capability matrix, binding state machine, failure router, and 27 acceptance tests here: The guide also links the current OpenAI file-input and Gemini PDF/video contracts as examples, while keeping them separate from what DeepSeek Harness currently ships. |
Uh oh!
There was an error while loading. Please reload this page.
some model api like chatgpt may allow pdf or video as input directly
please consider support pdf/video passthrough
All reactions