QEV SDK 0.2.1
An integrated Python SDK and English local playground for the released QEV 0.1.1 model.
QEV is inspired by LAYA's typed decisions and built with Qwen3.5-2B's text and vision
backbone, language adapters and trained decision heads. The vision encoder is frozen.
SDK and playground
- Windows launchers now open the same English interface as
qev playground. - The 2048 game has been retired from the playground; research reports remain available.
- The model is labeled Multilingual. Published evaluations focus on English, and equivalent
accuracy across languages has not been established. - The pip wheel includes the playground, presets, six sample photographs and license notices.
qev playgroundstarts the interface;qev.load()provides Python inference.- First use downloads missing, pinned model files; subsequent launches reuse the cache.
qev predict,qev serve,qev downloadandpython -m qevshare the same runtime.- uv can run the package as an isolated tool or from the locked source checkout.
- Star and support links are optional. No install hook authenticates to GitHub or stars a repo.
- An English QEV agent skill adds typed text/image judgments, dry runs, validated responses
and review handling. It is a separate installable skill folder and ZIP using this SDK.
This release changes packaging and launch behavior. Model weights, inference modules,
calibration and the evaluation results below remain those of QEV 0.1.1.
Same-input comparison
All three frozen models were evaluated on the same 2,800 English questions on one
RTX 4060 Ti 8 GB, without new training, prompt selection or calibration fitting.
| Test | LAYA English | LAYA Typed Decisions | QEV 0.1.1 |
|---|---|---|---|
| Official typed-decisions, 2,000 questions | 36.05% | 76.95% | 77.00% |
| News topic, 400 examples | 95.00% | 95.25% | 82.50% |
| Emotion, 400 examples | 58.75% | 60.00% | 50.25% |
Typed-decisions is an adapted regression benchmark. The one-question lead over the
specialist does not establish an advantage. News and emotion are task-held-out relative
to the audited QEV adaptation sources; backbone pretraining overlap is unknown.
LAYA was smaller and faster, with better probability quality on typed-decisions.
Full metrics, confidence intervals, training exposure and source hashes are published.
The model introduction includes an actual photograph with all three typed answers,
downloadable inputs and recorded outputs, and coverage-versus-accuracy tables. The example
is a previously inspected verification fixture, not a new benchmark claim.
Downloads and installation
- Model: https://huggingface.co/ken-jo/qev
- Data: https://huggingface.co/datasets/ken-jo/qev-data
qev-0.2.1.zip: checkpoint, custom runtime and evaluation evidence.qev-data-0.1.1.zip: 15 historical corpus snapshots with source-specific licenses.qev-0.2.1-py3-none-any.whlandqev-0.2.1.tar.gz: Python SDK distributions.qev-skill-0.2.1.zip: agent skill, helper, reference material and licensed example photo.
The SDK contains code and playground assets. Initial loading, or qev download, retrieves
the adaptation and pinned Qwen base.
The adaptation file is 32.01 MB; the full inference model is 2.213B parameters and the
upstream weight download is approximately 4.55 GB. Python 3.12 is required.
Available on PyPI: pip install qev==0.2.1.
Run qev playground, or use uvx --python 3.12 --from qev==0.2.1 qev playground.
Verification and limitations
For SDK 0.2.0, 184 SDK/runtime tests and 14 skill tests passed. All 38 frozen inference modules, learned weights and calibration
remain unchanged. Packaged text and three-type photograph inference preserve fixture
answer probabilities exactly. A fresh Python 3.12 dependency environment, pip installation
of the wheel, uvx text inference and uv run playground execution passed. Seven UI
presets, two image resolutions and HTTP downloads of all six sample images passed.
The upstream backbone reused an existing cache; it was not downloaded again. Browser
visual inspection was unavailable. See release/sdk-verification.json for the exact scope.
During the SDK 0.2.0 release, the skill was also run against the local model server for a text request and a photograph
with choice, score and noul questions; results matched the recorded fixture. Linux CI
build/install checks and authenticated PyPI Trusted Publishing succeeded. Published PyPI
files were downloaded anonymously and checked against the CI artifact hashes.
For SDK 0.2.1, the wheel and source distribution were inspected, and packaged GPU text
and three-type image inference matched the previous fixture responses exactly. The live
English GPU interface exposes Multilingual labeling and returns 404 for the retired game
page and game API. The earlier unit-test suite was not rerun for this packaging patch.
Current publication evidence is recorded in release/sdk-verification.json,
release/pypi-publication.json and release/huggingface-publication.json.
Separate image/workflow results: CIFAR-10 guard 95.83%, fresh procedural workflows
69.38%, local-photo HTTP p95 114.94 ms on RTX 4060 Ti 8 GB. The latter excludes loading,
WAN and concurrency. Visual 2048 remained unsolved: 0 wins in 72 exploratory games.
Uncertainty and visual reasoning limits are documented alongside successful results.
Public documentation is English. Personal Korean notes remain archived separately.
Corpus archives retain their original bytes, record IDs and splits; historical stages
overlap. No public demo Space is created by this release.