Repository navigation
0.1.9 - Tested prompt boundaries and adversarial fixtures
Hardening
- Application-owned trust boundary appended to all nine runtime stages, including customized prompts. Generated drafts get the boundary from the app, not model output.
- Structured JSON for generator/code reference data; exact stage output checks, invalid-summary rejection and enabled checker failure shown as non-approval.
- Synchronous question API now applies enabled question/relevance checks like the streamed flow.
- Static code review tightens exact imports/signatures, reflection, secret printing, environment mutation and import-time execution patterns. No generated code auto-executes.
Test matrix
- 442 source tests pass locally, rebuilt wheel/sdist metadata/install checks pass.
- 13 synthetic attack families across nine prompt constructions (117), 26 code input/policy constructions, runtime/history/failure fixtures, nine AST escape examples, eight malformed-label/citation cases and four benign controls.
- Genuine ZIP drop/review/consent/start/cited answer/stop/remove flow and all nine generator UI stages checked on installed wheel, desktop/phone, reduced motion, no auto-save.
- ChatGPT provided a separate simulated attack second opinion. It did not run the actual backend. Optional second followup returned a site error after one retry. No live-model attack-success rate is claimed.
Residual risks
No finite test matrix covers every attack. Prompt defenses and JSON serialization are not immunity. Models can still obey poisoned content or launder false claims; citation syntax does not prove support. Optional support checks are model-dependent and off by default. Streaming may show text before final checks, so there is no confidentiality guarantee. Static Python checks do not isolate capabilities or enforce all network endpoints. Bounded mode is NOT a security sandbox; host files/network remain accessible. Human review stays required.