Repository navigation
v0.4.0
Added
- The system prompt now varies with the model as well as with the window. A
frontier model infers the shape of a job from the job; a local one of
twenty or thirty billion parameters does better having been shown, so the
worked example and the terse-answer examples go to the models that need
them. Unknown counts as needing them, because that is what a model nobody
recognises usually is. The rules themselves go to everyone. ask_user: the model can stop and put a question to the user with two to
nine options, and the answer comes back as the tool result. It is for a
fork in the task that only the user can settle, asked before the work
rather than after it. Pick with1-9;eschands the decision back,
and the model is told to take the option that changes the least and say
which. A run with-phas nobody at the keyboard, so the question is
printed and answered "nobody is here" instead of hanging.- A permission prompt takes four answers, not three:
yonce,aalways,
sskip,nno. Skip means leave this step out and carry on with the
rest; no means stop, because the user has something to say. Both are told
to the model as "it did NOT happen", and neither may be reported as done. - Four modes for how much thoth asks before it acts, cycled with
shift+tab, named with/mode, and set for one run with--mode:
manual(the default, ask every time),accept edits(file changes go
through, the shell and the network still ask),auto(nothing asks) and
plan(nothing is changed at all). Accept-edits stops at the shell on
purpose: a file change is a diff with an undo behind it, a command is
neither. The mode is never saved, so one turned on for a sandbox is not
still on tomorrow against a real repository. - Plan mode answers with the plan, then asks what to do with it: carry it
out with edits accepted, carry it out asking each time, or keep planning.
The first two switch the mode and send the plan back to be carried out,
so there is nothing to retype. The tools that write are refused while it
is on, and the refusal tells the model to describe the change instead. shift+enterstarts a new line in the input instead of sending it, and
the input grows to as many rows as the message has lines, up to ten.
alt+enterandctrl+jdo the same for terminals that never deliver the
first one, and a line ending in\breaks where even those do not get
through.upanddownwalk the lines of a multi-line message, and are
the history again when there is only one line. On unix thoth now asks the
terminal for key disambiguation, without which shift+enter arrives as a
plain enter; the windows console tells them apart on its own.write_filetakesappend, for adding a section to the end of a file.
It needs no prior read, because nothing already in the file is touched.
It is how a file too long for one call gets written: the first section
normally, every one after it appended. Doing that withedit_filemeant
inventing a uniqueold_stringout of the last lines of the file, and the
last lines of a Rust file are}and}.
Security
- A tool result that reads like an instruction aimed at the model gets a line
saying so, whatever it came from: a file, a page, a search result. The
phrases an injection needs are the tell, since it has to cancel what came
before and usually asks to be kept quiet. Refusing was already reliable;
saying so was not, and a hidden order the user never hears about is the
half of the attack that still works. - Anything fetched from the network arrives wrapped in a line saying what it
is: content someone else wrote, to be read and never obeyed, and to be
reported if it asks for a command, a file change or for the instructions to
be set aside. The rule was already in the system prompt, two thousand
tokens before the page shows up; the warning that holds is the one next to
the payload. Search results are wrapped too, since a title and a snippet
are the cheapest place on the internet to put a sentence in front of
somebody else's agent.
Fixed
- A denied action is not reported as done. Answering
nto a write outside
the working directory stopped the write, and the model still signed off
with "written the file, it has hello in it". The refusal now says the
action did NOT happen and not to report it as done, and the end of the
request counts what was denied, so a false sign-off is contradicted where
the user is reading. - A background process thoth started and the model never stopped is named at
the end of the request. "Closing the server now" followed by the turn
ending leaves it holding the port for the rest of the day, and whether it
is still up is a question the operating system answers rather than
something to take the model's word on. - A tool call missing a field is told which tool and what a good call
carries. serde said "missing fieldold_string" and stopped there, which
names neither, and the model has to guess which of its calls went wrong.
The schema already lists the required fields, so the error can too. - Several files asked for in one turn no longer blow the window apart. Each
tool result was capped on its own, so six read_file calls in one reply
arrived as six full results and the request after them went over the
window. A server does not complain about that: it drops the front of the
prompt, which is the system prompt and everything agreed so far, and the
model answers the last file with "what would you like me to do with this?".
One turn's results now share one turn's room, and a result that runs out
of it says where it was cut and to ask again. The same job in an 8k window
went from losing the task entirely to finishing it across two compactions. - Running the tests again after a fix is no longer blocked as a repeat. The
loop breaker counts a command by its text, andbun testafter an edit is
the same text and a different answer; a model that fixed the bug and went
to confirm it got "this exact shell call was already run 3 times, the
result will not change". A change to any file now clears what was counted
about commands, which is the one habit worth encouraging. - A path that is not there says where the working directory is. "The system
cannot find the path specified" cost three more calls and apwdevery
time a model guessed an absolute path wrong. - Adding a dependency counts as changing code for the note above: a version
number written into a manifest and never resolved is the same broken
handoff as code that was never built. - A request that changed files and ran nothing says so. The prompt tells the
model to build what it changed, and a model that skipped it still signs off
with "build passed"; whether a command ran is not the model's word about
itself, thoth watched every tool call the request made. It now prints a
note, so "it builds" over an untouched compiler is contradicted on the
spot. - The preview of an
edit_filewith an emptyold_stringgives the reason
the tool will give. It used to say "old_string not found in file", sending
the reader to look for a typo in a string that is not there at all. - The write cap does not shrink a hosted model to 170 lines a file. It used
to be derived from the tool-result cap, which guesses small when a profile
declares nocontext_window, on purpose: one grep must not eat the context
of a server nobody described. A write is capped for the opposite reason,
so an undeclared window now means 40k characters there, not 6k. - A compaction in the middle of a request no longer leaves files unreadable.
The loop breaker counts identical tool calls, and it kept counting across
the compaction that had just thrown their results away, so a re-read of the
file the model was working on came back "STOP: this exact read_file call was
already run 4 times" with no way to get the content. It burned the rest of
the turn going in circles. - Answers come back in the language the request was written in. Local models
read a system prompt full of English and answer in English (or in Chinese)
whatever the style rule says, so a request in a non-Latin script now names
the language outright. The name is remembered for the session, because the
reminder rides on the user message and a compaction deletes every one of
them: after the first compaction the whole conversation went back to
English. The summary is asked for in the user's language too. - "old_string appears 2 times" now says which lines they are. Telling a
model its anchor is not unique and to "add surrounding context" leaves it
guessing which of the matches it was aiming at; the line numbers let it
pick the context in one turn instead of two. - An
edit_filewith an emptyold_stringis answered with the tool that
does what it was trying to do. An empty anchor matches between every pair
of characters, so a model trying to add a[[bin]]section to a 7-line
Cargo.toml got "old_string appears 61 times, add surrounding context to
make it unique": true, and no help at all. It now says to usewrite_file
withappend. - One edit in a
multi_editthat asks for no change (old_string equal to
new_string) is skipped and named in the result, instead of failing the
whole call. Throwing away three real edits because the fourth had nothing
in it cost a turn every time a model got one of four wrong. A call where
every edit is like that is still an error. - Reading a file again after changing it is no longer treated as going in
circles. The loop breaker counts identical calls, so a model that read a
file, edited it wrong, and read it back to see the damage was told "this
exact read_file call was already run 3 times, the result will not change"
while the file on disk said otherwise. A change to a file now clears what
was counted about reads of it. Only the count: what makes the re-read
supersede the copy taken before the change stays, because after a change
that copy is not merely old, it is wrong, and leaving it would put both
versions of the file in front of the model at once. - Hitting the context limit mid-generation no longer costs two compactions.
The recovery path compacted and jumped straight back to the model, past the
point where the auto-compact request from that same turn is answered, so
the next turn compacted again with four messages in the conversation and
threw away the file it had just read. - Long files are built up in sections instead of being lost. A model that
tries to emit a whole 1200-line file in one write_file has to fit every
character of it, and its reasoning, inside one generation; past the window
it is cut off mid-file and all of it is gone. write_file now refuses a call
larger than half of what a tool result may take and says how to write the
file in parts, and the prompt asks for that shape up front. - The model stopped deleting files to rewrite them. write_file already
counts a file thoth wrote itself as read, so overwriting it was allowed all
along; the prompt said only "write_file is for brand-new files", and a
model that wanted to start over on a 1300-line file it had just written
reached for delete_file instead.
Changed
- Release notes are the changelog entry for the tag, put there by one
publish job instead of by all five build jobs at once. Each of them was
generating notes of its own, which is where the repeated
"Full Changelog" lines in v0.3.0 came from.
Full changelog: v0.3.0...v0.4.0