Repository navigation
Releases: webdevtodayjason/archiver
Release list
LAST LIGHT 0.1.2
Publishes LAST LIGHT under its own release line after the failed farm submission selected Tiiny Brain 0.1.3. The shipped LAST LIGHT tree passes the farm scanner and its offline, read-only Python 3.11 selfcheck.
LAST LIGHT shelf 0.1.0
The prebuilt shelf that LAST LIGHT downloads on first run. This is data, not
code: there is no program in it. The app fetches it, checks it by length and
then by hash, and unpacks it into its own data directory.
size 73196259 bytes
sha256 85dacb04a56693295ffe23e2251f746a7250a366a475469eb1ce4fda43181e5c
What is on it, 5,955 documents:
- 5,934 Vikidia articles (https://en.vikidia.org), under CC BY-SA 3.0 or the
GFDL, the same terms as Wikipedia. - 21 survival and reference PDFs, 4,189 pages. Most are works of the US federal
government and so are not under copyright in the United States: FEMA, USDA,
EPA, USGS, HHS, the Army's Borden Institute, Oak Ridge. Two are public domain
by age (Harding 1907, Gibson 1881, Hasluck 1917). One carries a written grant
on its own page (the SODIS Safe Water School manual). - Three of those 21 are Hesperian titles under CC BY-NC-SA: Where There Is No
Doctor, Where There Is No Dentist, A Book for Midwives. Non-commercial.
Two plain text files travel with it. ATTRIBUTION.txt carries the Vikidia credit
and licence, which a share-alike licence requires and which the app enforces:
it refuses to unpack a bundle that does not have it. SOURCES.txt lists every
document and why it is free to pass on.
Anything whose terms could not be established in writing was left off rather
than guessed at. The corpus this was cut from holds 109 PDFs; 88 of them are not
here.
LAST LIGHT 0.1.1
The first run can now go and get its books.
0.1.0 shipped with nowhere to fetch a shelf from, so a fresh install served,
said the shelf was not here and refused to download one. The shelf is published
now, as its own release on this repo, and this build points at it.
shelf https://github.com/webdevtodayjason/archiver/releases/tag/shelf-v0.1.0
size 73196259 bytes
sha256 85dacb04a56693295ffe23e2251f746a7250a366a475469eb1ce4fda43181e5c
5,955 documents: 5,934 Vikidia articles under CC BY-SA 3.0, and 21 survival and
reference PDFs that are free to pass on. The download resumes if it breaks, is
checked by length and then by hash before anything is unpacked, and is unpacked
through a filter that refuses absolute paths, parent traversal and links. A
bundle with no ATTRIBUTION.txt in it is refused outright.
All three shelf constants read the environment first (LAST_LIGHT_SHELF_URL,
LAST_LIGHT_SHELF_SHA256, LAST_LIGHT_SHELF_BYTES), which is how a mirror gets
pointed at. Nothing else changed, and the app still works with no device, no
network and no corpus.
LAST LIGHT 0.1.0
The off-grid library, packaged for the Tiiny App Farm. Ask a question in plain
words, get an answer built only out of the books on the device, with the book
and the page printed under it. When the shelf does not cover the question it
says so instead of inventing an answer.
This is the books half of the archiver repo. The notes half ships separately as
tiiny-brain and is not affected by this tag.
What is in the farm build:
- ask and cite, browse the shelves, look a subject up by name
- no ingestion. Turning a PDF into text needs poppler, reaching poppler means
starting another program, and the farm's archive scanner refuses a tree that
can do that. The desk tool in the same repo still ingests. - a prebuilt shelf instead: downloaded once, resumable, checked by length and
then by hash before anything is unpacked, and unpacked through a filter that
refuses absolute paths, parent traversal and links. - with no shelf yet it serves and says so, and offers to fetch one.
- Python standard library only.
The shelf URL is still a placeholder. Nothing has been published to fetch, so
--download-shelf refuses and says why rather than reaching for something that
is not there.
python3 last-light --serve [--port N]
python3 last-light --selfcheck
python3 last-light --shelf
0.1.3: load notes from the page
Loading notes should not have been a terminal command, and now it is not.
The empty board carries the door: "No notes yet. Point Tiiny Brain at a folder
of markdown", with a button. The same button sits in the top bar and the same
panel is under Settings. Give it a folder, or browse to one from this machine,
and one press runs the walk, the filing, the chunking, the embedding on your
Tiiny and the name index, telling you which of the five it is on. The board
redraws itself when it finishes. Point it at the same folder later and it only
does what changed.
The five commands still work exactly as they did. They are just no longer the
only way in.
Every step that cannot run now says which step and keeps what came before it: a
Tiiny nobody has set up, a Tiiny that did not answer, a Tiiny answering and
serving no embedding model, a folder with no markdown. The names are built
either way, so a sleeping Tiiny costs you the ability to ask and nothing else,
and try again picks up where it stopped. An empty board is never left without
words on it.
Three faults turned up on the way and are fixed here: a load started before the
device was configured used to leave the page saying "loading" for ever; the
board could stay empty after a load while the numbers under it counted the new
notes; and the two routes the page polls no longer wait on the database during
the slowest step. Two different vaults with an Inbox at the top now stay two
notes.
PDFs in the folder are counted and reported as not loaded in this build.
Python standard library only. No subprocess anywhere in the shipped tree.
Tiiny Brain 0.1.2
A fresh install could not answer its own first two questions. /api/vitals and /api/settings both count rows in entity, and nothing on the cockpit's path created that table, so on a database nobody had built an entity index in the query raised, the handler dropped the socket, and the browser showed ERR_EMPTY_RESPONSE with an empty page behind it.
- Every table the app owns is now created on every connection, by one function the three modules share. A fresh database and one left half built by 0.1.1 both come out complete.
- A route that raises answers 500 with the reason instead of closing the connection in silence. The traceback still goes to farm.log.
- The corpus moves to the data directory the farm keeps beside the version folders, so an index survives an update.
ARCHIVER_HOMEstill wins and a plain checkout is unchanged. Your notes were never involved: they are markdown files in your own folder. /favicon.icoanswers 204 and the page carries its own mark.- First CI: the tests and the selfcheck run on Linux, macOS and Windows.
Tiiny Brain 0.1.1
The cockpit reads TIINYAPP_PORT when no port is given, so farm start tiiny-brain --port N moves it.
Tiiny Brain 0.1.0
Your own notes, held by the machine on your desk.
Point it at a folder of markdown. It reads every note, works out which names keep turning
up, and draws the ones that keep turning up together. Click a name and its neighbourhood
pulls forward. Ask a question and the answer cites the notes it came from, or says plainly
that nothing it retrieved covers it.
Two ways of answering sit behind that box, and the split is the point. "What did we decide
about X" is a retrieval question, so it goes to nearest-neighbour search. "Who is Richard"
is not. Search hands the model whichever notes name him near words that match the question.
What answers it is every passage that names him at all. So a who-or-what-is question about
a name the archive knows reads the mentions instead, and the answer says which path it took.
Write a note from inside it and the note lands in your own vault as markdown, gets chunked
and embedded, and is answerable a few seconds later. Your notes stay plain files in your own
folder the whole time.
The light in the top bar tells you whether it is working: whether the device answered,
whether embeddings ran, whether the answering model is reachable, and whether every chunk
is embedded. It names the one that is broken instead of making you read a log.
This archive is the notes half. The same engine turns a shelf of scanned PDFs into citable
text, which needs poppler and an OCR engine and is not in here, because an app that only
reads your notes has no business running other programs. Both halves are in the repo.
Python 3.9 or newer, standard library only. A Tiiny for the embedding and the answering.