Repository navigation
Releases: thoth-coder/thoth
Releases · thoth-coder/thoth
Release list
v0.4.5
Added
- Lua. One row in the stack table:
luac -pper file plus luacheck, marked as
a stack that running does not check, because a chunk is only compiled when
it is first required and a typo in a branch nothing took is still there. - thoth names the reference to search, per stack:
site:docs.rs,
site:pkg.go.dev,site:docs.python.organd so on, out of the same table
as everything else. Search is eight results scraped off one page, and a
model guessing a phrase gets tutorials and stale blog posts; a version
number recalled from training is the thing most likely to be wrong.
Fixed
- The interface behaved far worse than
-pon the same task, and this was
why. With an editor attached thoth adds aproblemstool, and its
description told the model to reach for it "before falling back to a full
build". Editor diagnostics are not a build: they were invisible to the guard
that refuses a second whole-file rewrite, so a model doing exactly what the
tool told it to do got refused for the rest of the request with nothing it
could do to clear the refusal. A-prun in a bare directory has no editor
and never saw it. The build is the check now andproblemsis the second
opinion, in that order, in both the tool description and the prompt; the
language server lags behind edits and can need restarting before a fixed
error clears, so a diagnostic that survives a passing build is stale and
says so, rather than sending the model back to change working code. - Two of thoth's own rules were sentences addressed to the model with nothing
behind them, which is the wrong shape for a rule about what leaves the
machine. A search query that carries line breaks, a whole file's worth of
text, or something shaped like a credential is now refused in code, since a
query is the one thing thoth transmits. Reading.envand its family, and
.pem,.key,.npmrc,id_rsaand the rest, is a permission prompt
rather than a request: the path is shown and the answer is the user's. Auto
mode still lets both through, because auto mode is the user saying nothing
should ask.
Full changelog: v0.4.4...v0.4.5
v0.4.4
Added
- A reply with nothing in it is asked again rather than taken for an answer.
A model that spends its turn thinking and emits no text and no tool call
was ending the request wherever it had got to: in one run that left a
cargo initskeleton, a hello-world main, and a report that the work was
done. Twice, then it is handed back to the user as before. - A request that changed code and has not run anything against it since gets
one more turn,
with thoth asking for the check by name before the request ends. The note
that said so was already there, but it told the user, after the model had
stopped and could no longer act on it. In a run of six Rust projects the
one case where nothing was ever run was also the one that compiled cleanly
and did nothing at run time: five tasks spawned into a channel that
collected zero results. Asked to check its own work, the model found it and
fixed it. Once per request, and only where the project has a check to name.
The question is what has happened since the last edit, not whether any
command ran at all: a model that builds, then edits twice more and stops
has run a command and checked nothing, and two runs of the suite caught
exactly that, each leaving a project that did not compile.
Fixed
-pmode strips escape sequences from what it prints, which it never did.
The interface has taken them out of every line it draws since 0.4.2, and
this half was missed: a file thoth was asked to read could still repaint or
wipe the terminal of anyone runningthoth -p, and a title-setting OSC or
a right-to-left override went straight through. Found by a case that seeds
a source file carrying both escapes and instructions addressed to the
model; the model ignored the instructions and deleted them from the file,
but the escapes reached the terminal.- A file cannot be written whole twice in one request with nothing run
against it in between. The model would write a file, be told nothing by
anyone, decide on its own that it was wrong, and write the whole thing
again: four passes over the same 330 lines in one request, each costing a
full generation and each made out of the same guesses as the last. The
refusal names this project's check, out of the same table everything else
reads, and points atedit_filefor changing the part that is actually
broken. A check running clears it, and so does writing a different file. 2>&1is taken out of a command on Windows. PowerShell does not merge a
native program's stderr into its stdout the wayshdoes: it wraps every
line in an ErrorRecord, socargo check 2>&1came back as a
NativeCommandError with a caret diagram around the word "Compiling" and
the real error buried underneath, and a command that succeeded was marked
as failed. Both streams were already captured and labelled separately.- A repeated tool call is only stopped when its answer repeats too. thoth
hands out a background log path and asks the model to watch that file,
then refused the third read with "the result will not change" while the
build was still writing to it. The one call thoth asks to be made twice
was the one it punished. - A background command says how to wait for it, in the shell it will be run
in, instead of "wait a moment first" with nothing that could do the
waiting. Reading a log in a tight loop only ever shows the same bytes. - A command that fails because the shell rejected the syntax says which
shell it is.&&between two commands is a parse error on Windows and
cost a whole turn every time a model wrote one out ofshhabit. - A model call that stops answering no longer hangs thoth for ever. Only the
connection had a timeout, so a server that accepted it and then went quiet
left the spinner turning and the elapsed timer counting with nothing behind
either, which looks exactly like a slow model and never ends. Silence
before the first byte is allowed ten minutes, because on a large local
model the whole prompt is evaluated before a single token comes back;
silence in the middle of a reply is allowed two, because by then tokens
arrive steadily and a long gap is a stream that has stopped. - A project created during the request is now recognised once its manifest
exists. The stack was read once at the start, so "build me a project"
began in a directory with nothing to detect, and nothing the model ran
afterwards counted as a check.
Full changelog: v0.4.3...v0.4.4
v0.4.3
Changed
- The binary is a third smaller: 8.4 MB to 5.7 MB on Windows. Link-time
optimization across the whole program, one codegen unit, symbols stripped,
and no unwinding tables, since thoth has nothing that has to run while a
panic travels and the terminal is put back either way. - The demo screens behind
--vieware no longer compiled into a release
build at all. They were dead code in one, but only because the optimizer
noticed.
Full changelog: v0.4.2...v0.4.3
v0.4.2
Added
- Copying text out.
ctrl+y(or/copy) puts the last reply on the
clipboard and/copy allthe whole conversation; thoth asks the terminal
to do it with OSC 52, so it works over ssh and inside tmux and pulls in no
clipboard library.ctrl+thands the mouse back to the terminal so
click-and-drag selects text the way it does in every other program, and
ctrl+tagain gives thoth the scroll wheel. - Slash commands complete as you type them. A
/at the start of the
message opens the list, with what each command does beside it, narrowing
with every letter;taborentertakes the highlighted one. It is the
same picker@pathalready used. thoth --view <name>in a debug build draws any screen of the interface
with made-up contents and prints it as text: the start screen, a session
mid-task, a permission prompt, the chooser, plan mode, both pickers, the
four permission modes, the profile screen, and the whole thing at 46
columns, a turn mid-stream, a scrolled-back transcript, hostile content,
markdown, and a chooser with more options than rows to draw them in.
thoth --viewlists them and--view-size WxHsets the
terminal. Screens that only appear after a permission prompt or an
ask_usercall were minutes of model time away from being looked at,
which is where their bugs were living.cargo test every_view_draws
renders all of them at two sizes.
Fixed
- A shell command could draw on top of the interface on Windows. Redirecting
stdout and stderr is not enough there: a console program inherits its
parent's console and can write to it through the console api, which is what
Write-Progress,Invoke-WebRequest's progress bar and a coloured
Write-Hostdo, and a child that changes the console mode takes the
terminal's escape handling with it. Commands now run with no console of
their own. - PowerShell wrote its output in the console codepage, so anything past ascii
came back through the pipe as mojibake: Thai, Japanese and accented text in
a command's output was unreadable, to the user and to the model both. The
pipe is UTF-8 now. Commands also run-NonInteractive, so one that would
have asked a question fails instead of waiting for an answer nobody can
give. - Escape sequences from tool output reached the terminal.
cargo test,
gitandnpmemit colour the moment they think a terminal is listening
and the shell tool captures it verbatim, so a\x1b[31min a result
repainted the rest of the screen and\x1b[2Jwiped it. Every line thoth
draws now has its escapes, control characters and bidi overrides taken out
first, wherever it came from: tool output, a file, a web page, the model's
own reply, the model name the server gave, or a paste. - A question with more options than the terminal had rows drew as many as
fitted and no more: the rest could be walked onto with the arrows but
never seen, and the question itself could be pushed off the screen. The
list scrolls with the highlight now, says how many are below it, and never
takes the last row of the transcript. - A tool call with a permission prompt open said "running" underneath it.
It has not run, and may never run: it now says it is waiting for you. - "plan ready" and its three keys were written into the transcript and drawn
again on the state line one row below. The transcript keeps the moment,
the state line keeps the keys. - Tabs in tool results and error messages were drawn as a single character,
so a stack trace lost its indentation. - On a narrow terminal the status line dropped "scrolled up" and "selecting"
to make room for the token counts, so the two states that explain why the
screen has stopped behaving were exactly the ones not shown. They win over
the numbers now.
Changed
- The input sits in a box of its own instead of under a rule. It is the one
part of the screen written into rather than read, and the border says so;
it also carries whether thoth is listening, in the colour, which saves
saying it in words somewhere else. - A tool's result hangs off the call on a corner (
⎿) with the rest of it
lined up past the corner. A result and the next line of prose were both
just indented text before. - A shell command is drawn once, not twice: the preview under the call
repeated what the header line already said. It comes back the moment the
header line is too narrow to hold the whole command, and a command that
does not fit now wraps at its spaces instead of being cut, because a
permission prompt is answered on what is drawn. - The model's question is drawn above the input rather than below it: it is
asked before anything is typed, while@pathcompletes what is being
typed, and the two belong on opposite sides of the box. - Markdown markers are obeyed instead of printed:
##in front of a heading
is dropped rather than shown next to a heading that is already bold,-
becomes a bullet with the wrapped rest of the item lined up under it,>
becomes a bar down the side, a ``` fence becomes a rule carrying the
language, and---becomes a rule. - Each tool call is marked with a dot coloured by how it went: still going,
done, failed. A screenful of them reads down the left edge. - The start screen names the three things nobody discovers by typing a
sentence at thoth:/,@and shift+tab. - The chooser for a question from the model is a list you move through with
up/downand take withenter, drawn under the input with the
highlighted row marked.1-9still picks one straight off. - Its last row is not one of the model's options: it opens the input box so
you can write the answer in your own words when none of them is what you
wanted. The model is told plainly that you took none of its options, so it
does what you asked rather than the offer yours sounded closest to.esc
goes back to the list, and a message half-typed when the question arrived
is put back afterwards.
Full changelog: v0.4.1...v0.4.2
v0.4.1
Added
- thoth knows what checks a project, per stack, and says so. TypeScript,
JavaScript, Rust, Go, PHP, Python, Ruby, C#, C/C++ and Java/Kotlin each
have a row in a new table naming the command that checks the whole project
without running it, so the prompt can name it instead of leaving the model
to guess from whichever ecosystem it saw most of. Supporting another
language is adding a row. - The same table says whether running the code already was that check. For a
compiler it is; for TypeScript, JavaScript, PHP, Python and Ruby it is not,
and the prompt says so where it applies: a server that starts and answers
every request the test made can still be broken in every line that did not
execute. That case satisfied every existing rule about checking your work,
which is how a Bun API shipped with six type errors and a clean report. - A note after a request that changed code of one of those stacks and never
checked it. The existing "files changed and nothing was run" note cannot
catch this one, because something was run.
Changed
- Errors are fixed from the checker's output now, not from reading the
source: never fix an error you have not read, run the command, fix exactly
what it printed, run it again. - "Restructure this", "split it up", "do it properly" with no target layout
named is called out as a fork forask_user, with the worked example to
match. Several structures are right, the project cannot say which one was
meant, and the old rule stated the case too abstractly for a small model
to recognise it.
Full changelog: v0.4.0...v0.4.1
v0.4.0
Added
- The system prompt now varies with the model as well as with the window. A
frontier model infers the shape of a job from the job; a local one of
twenty or thirty billion parameters does better having been shown, so the
worked example and the terse-answer examples go to the models that need
them. Unknown counts as needing them, because that is what a model nobody
recognises usually is. The rules themselves go to everyone. ask_user: the model can stop and put a question to the user with two to
nine options, and the answer comes back as the tool result. It is for a
fork in the task that only the user can settle, asked before the work
rather than after it. Pick with1-9;eschands the decision back,
and the model is told to take the option that changes the least and say
which. A run with-phas nobody at the keyboard, so the question is
printed and answered "nobody is here" instead of hanging.- A permission prompt takes four answers, not three:
yonce,aalways,
sskip,nno. Skip means leave this step out and carry on with the
rest; no means stop, because the user has something to say. Both are told
to the model as "it did NOT happen", and neither may be reported as done. - Four modes for how much thoth asks before it acts, cycled with
shift+tab, named with/mode, and set for one run with--mode:
manual(the default, ask every time),accept edits(file changes go
through, the shell and the network still ask),auto(nothing asks) and
plan(nothing is changed at all). Accept-edits stops at the shell on
purpose: a file change is a diff with an undo behind it, a command is
neither. The mode is never saved, so one turned on for a sandbox is not
still on tomorrow against a real repository. - Plan mode answers with the plan, then asks what to do with it: carry it
out with edits accepted, carry it out asking each time, or keep planning.
The first two switch the mode and send the plan back to be carried out,
so there is nothing to retype. The tools that write are refused while it
is on, and the refusal tells the model to describe the change instead. shift+enterstarts a new line in the input instead of sending it, and
the input grows to as many rows as the message has lines, up to ten.
alt+enterandctrl+jdo the same for terminals that never deliver the
first one, and a line ending in\breaks where even those do not get
through.upanddownwalk the lines of a multi-line message, and are
the history again when there is only one line. On unix thoth now asks the
terminal for key disambiguation, without which shift+enter arrives as a
plain enter; the windows console tells them apart on its own.write_filetakesappend, for adding a section to the end of a file.
It needs no prior read, because nothing already in the file is touched.
It is how a file too long for one call gets written: the first section
normally, every one after it appended. Doing that withedit_filemeant
inventing a uniqueold_stringout of the last lines of the file, and the
last lines of a Rust file are}and}.
Security
- A tool result that reads like an instruction aimed at the model gets a line
saying so, whatever it came from: a file, a page, a search result. The
phrases an injection needs are the tell, since it has to cancel what came
before and usually asks to be kept quiet. Refusing was already reliable;
saying so was not, and a hidden order the user never hears about is the
half of the attack that still works. - Anything fetched from the network arrives wrapped in a line saying what it
is: content someone else wrote, to be read and never obeyed, and to be
reported if it asks for a command, a file change or for the instructions to
be set aside. The rule was already in the system prompt, two thousand
tokens before the page shows up; the warning that holds is the one next to
the payload. Search results are wrapped too, since a title and a snippet
are the cheapest place on the internet to put a sentence in front of
somebody else's agent.
Fixed
- A denied action is not reported as done. Answering
nto a write outside
the working directory stopped the write, and the model still signed off
with "written the file, it has hello in it". The refusal now says the
action did NOT happen and not to report it as done, and the end of the
request counts what was denied, so a false sign-off is contradicted where
the user is reading. - A background process thoth started and the model never stopped is named at
the end of the request. "Closing the server now" followed by the turn
ending leaves it holding the port for the rest of the day, and whether it
is still up is a question the operating system answers rather than
something to take the model's word on. - A tool call missing a field is told which tool and what a good call
carries. serde said "missing fieldold_string" and stopped there, which
names neither, and the model has to guess which of its calls went wrong.
The schema already lists the required fields, so the error can too. - Several files asked for in one turn no longer blow the window apart. Each
tool result was capped on its own, so six read_file calls in one reply
arrived as six full results and the request after them went over the
window. A server does not complain about that: it drops the front of the
prompt, which is the system prompt and everything agreed so far, and the
model answers the last file with "what would you like me to do with this?".
One turn's results now share one turn's room, and a result that runs out
of it says where it was cut and to ask again. The same job in an 8k window
went from losing the task entirely to finishing it across two compactions. - Running the tests again after a fix is no longer blocked as a repeat. The
loop breaker counts a command by its text, andbun testafter an edit is
the same text and a different answer; a model that fixed the bug and went
to confirm it got "this exact shell call was already run 3 times, the
result will not change". A change to any file now clears what was counted
about commands, which is the one habit worth encouraging. - A path that is not there says where the working directory is. "The system
cannot find the path specified" cost three more calls and apwdevery
time a model guessed an absolute path wrong. - Adding a dependency counts as changing code for the note above: a version
number written into a manifest and never resolved is the same broken
handoff as code that was never built. - A request that changed files and ran nothing says so. The prompt tells the
model to build what it changed, and a model that skipped it still signs off
with "build passed"; whether a command ran is not the model's word about
itself, thoth watched every tool call the request made. It now prints a
note, so "it builds" over an untouched compiler is contradicted on the
spot. - The preview of an
edit_filewith an emptyold_stringgives the reason
the tool will give. It used to say "old_string not found in file", sending
the reader to look for a typo in a string that is not there at all. - The write cap does not shrink a hosted model to 170 lines a file. It used
to be derived from the tool-result cap, which guesses small when a profile
declares nocontext_window, on purpose: one grep must not eat the context
of a server nobody described. A write is capped for the opposite reason,
so an undeclared window now means 40k characters there, not 6k. - A compaction in the middle of a request no longer leaves files unreadable.
The loop breaker counts identical tool calls, and it kept counting across
the compaction that had just thrown their results away, so a re-read of the
file the model was working on came back "STOP: this exact read_file call was
already run 4 times" with no way to get the content. It burned the rest of
the turn going in circles. - Answers come back in the language the request was written in. Local models
read a system prompt full of English and answer in English (or in Chinese)
whatever the style rule says, so a request in a non-Latin script now names
the language outright. The name is remembered for the session, because the
reminder rides on the user message and a compaction deletes every one of
them: after the first compaction the whole conversation went back to
English. The summary is asked for in the user's language too. - "old_string appears 2 times" now says which lines they are. Telling a
model its anchor is not unique and to "add surrounding context" leaves it
guessing which of the matches it was aiming at; the line numbers let it
pick the context in one turn instead of two. - An
edit_filewith an emptyold_stringis answered with the tool that
does what it was trying to do. An empty anchor matches between every pair
of characters, so a model trying to add a[[bin]]section to a 7-line
Cargo.toml got "old_string appears 61 times, add surrounding context to
make it unique": true, and no help at all. It now says to usewrite_file
withappend. - One edit in a
multi_editthat asks for no change (old_string equal to
new_string) is skipped and named in the result, instead of failing the
whole call. Throwing away three real edits because the fourth had nothing
in it cost a turn every time a model got one of four wrong. A call where
every edit is like that is still an error. - Reading a file again after changing it is no longer treated as going in
circles. The loop breaker counts identical calls, so a model that read a
file, edited it wrong, and read it back to see the damage was told "this
exact read_file call was already run 3 times, the result will not change"
while the file on disk said otherwise. A change to a file now clears what
was counted about read...
v0.3.0
Added
- Hosted apis are first-class. Anthropic gets a native transport
(/v1/messages) with prompt caching on the system prompt, the tool
schemas and the end of the history. OpenAI, Google's OpenAI-compatible
endpoint, OpenRouter and the rest work through the existing OpenAI path.
A profile picks the protocol withapi = "auto" | "openai" | "ollama" | "anthropic". - Per-profile
headers, for endpoints that do not take a bearer token
(Azure'sapi-key, gateways with their own). - Running cost. Set
price_in/price_out/price_cachedon a profile
and the status bar,/statusand-pshow what the session has spent. /modelsnumbers the list, and/model 3picks from it.multi_edit: several replacements in one file in a single call, applied in
order, all of them or none. One round trip instead of one per change, one
diff to approve, and no way to leave a file half edited./undo. Every file a request changes is snapshotted before the change,
and one request is one checkpoint, so/undoputs all of it back at once
(/undo listshows what is there). A file someone edited after thoth
wrote it is reported and left alone rather than overwritten, and the last
20 checkpoints live under~/.thoth/projects/<key>/undo/so a crash does
not take them with it.move_fileanddelete_file. Renaming and deleting used to mean reaching
for the shell, which skips the read registry, shows no diff, and turns one
"always allow" into a standing permission to runrm. Deleting needs the
whole file read first (the same condition as overwriting it), only works
inside the working directory, and never touches a directory.todo: the plan for the task, written before starting and rewritten as it
goes. It keeps a model from losing track of step three of five, and lets
you see the plan is wrong before the work is done. Only the latest version
of the list stays in the context.- The config screen asks which endpoint a new profile talks to (anthropic,
openai, gemini, openrouter, ollama) and fills in the url, so only the
model and the key are left to type. - Named config profiles.
thoth config(orthoth cfg) opens a screen to
edit them,/configdoes the same inside a session and applies the saved
profile to the running conversation.thoth -P NAMEruns one profile
once,thoth config use NAMEmakes it the default,thoth config list
shows what exists. Config files from 0.2 and earlier keep working; they
are read as a profile nameddefault.
Changed
- Tool calls the model writes as text (
<tool_call>{...}</tool_call>, or a
bare json object) are read and run instead of printed. A json object only
counts when it names a tool that exists, and anything that does not parse
comes back to the transcript untouched. - What a request costs is now budgeted against the context window instead of
fixed at numbers tuned for 16k: how much of a tool result is kept, how much
of a fileread_filereturns, and how much of the instruction file goes
into the prompt. A small window keeps its room, a large one gets to use it. grep,web_fetch,list_dirandglobcut their own output to the
budget too, the same wayread_filedoes, so the line that says the
result was incomplete is not itself the thing that gets cut off. They had
fixed limits of their own (11k characters, 15k, 500 entries) that ignored
how much room there actually was.read_filedoes its own cutting, so its "showing lines 1-240 of 900, read
on with offset=241" survives. It used to be replaced by a blind truncation
that left the model with no idea there was more, or how to ask for it.- Repeating a read-only call with the exact same arguments replaces the
older result in the context with a one-line note. Two copies of the same
file are one copy of dead weight; a read of a different range, or a search
with a different pattern, is left alone because it says something else. - The prompt ends with one worked example of a small task from grep to final
answer, and says how many tool calls are left for the request. Small models
follow an example better than another rule. - The style rules now say what not to write: no listing the edits again, no
repeating the plan just carried out, no explaining code that was not asked
about, and an answer as long as the question deserves. - Rules about git and about the editor's
problemstool are only sent when
there is a repo and an editor to use them on, and the tool itself is only
offered then: 800 characters of every request that was buying nothing. - Auto-compact works on every api, not just Ollama: it measures against the
profile'scontext_window. The field used to be callednum_ctx, which
is still read. - thoth no longer picks a model on its own when the endpoint is hosted. A
local server usually has one or two and guessing is a kindness; a hosted
one has hundreds and guessing spends the user's money, so it lists some
and asks. - The Ollama probe only fires at a local address. A hosted endpoint never
sees a request to a path thoth guessed at. stream_optionsis an OpenAI extension that not every compatible server
takes. When one rejects it, thoth drops the field and retries instead of
failing the turn: losing the token counts beats losing the answer.- The config file moved to
~/.thoth/config.tomlon every OS, next to the
state thoth already kept there, instead of the platform config directory
(%APPDATA%\thothon Windows,~/.config/thothelsewhere). The old path
is still read when the new one is missing, so upgrading changes nothing
until the next save. @pathnow opens a picker under the input instead of completing on tab
only: up/down move, tab or enter takes the highlighted entry, esc closes
it, and picking a directory lists what is inside it.- Startup screen: logo, version, working directory and, on the Ollama
native api, the context window that used to be a transcript line.
/clearshows it again.
Fixed
- The directory an "always allow" covered for a file outside the working
directory was recorded as the model spelled it, so reading../notes.txt
saved a grant ending in/..: no later path matched it, and/allow
showed the user something they could not place. Found by running it. globstopped at 500 matches and said nothing about it, so a truncated
list read like the whole answer. It says what it stopped at now.- The editor context (active file, selected text, the Problems panel) went
into every request unbudgeted: a large selection or one page-long type
error could take a good part of a small window. Both are capped, and say
when they were cut. - The permission preview for overwriting a file with CRLF line endings
showed every line of it as changed. It now previews what will actually be
written. - Every multi-line
edit_filefailed on a file with CRLF line endings,
which is most files on a Windows checkout.read_fileshows lines with
the carriage return stripped, so the text the model copies back could
never match the file byte for byte. Edits are now matched in the file's
own line endings, and a full overwrite keeps them instead of turning
every line into a change. --continueon a session that was killed between a tool call and its
result sent a transcript every api rejects, and nothing but/cleargot
past it. Half a turn is dropped when the transcript is loaded.- The
todotool refused a status it understood perfectly well: a model
writing "in_progress" or "Completed" instead of "doing" and "done" lost a
turn to a validation error. Those spellings are read as what they mean; a
word that means nothing here is still refused. - Security: one "always allow" for a file outside the working directory
covered every file everywhere for the life of the project. Reads and
writes outside it are now scoped to the directory the file is in; inside
the working directory one answer still covers the project. multi_edit,move_fileanddelete_fileshowed nothing once they were
always-allowed, unlikewrite_file,edit_fileandshell. Every one of
them now prints what it does whether or not it had to ask./undosaid nothing about a file it had not snapshotted because the
request was too large, so a request could come back looking complete while
one file was still changed. That file is now named and reported.- An undone checkpoint is kept, marked, instead of being deleted. For a file
that had been edited since and was therefore left alone, the copy in that
checkpoint was the only one left of what was in it before. - Security:
move_filecould carry a file in from anywhere or out to
anywhere, and needed no read first, which made renaming a way around
having to read a file before overwriting it. Both endpoints must now be
inside the working directory and the file must have been read, the same
bar as deleting it. Found by review of the commit that added it. - Security: the
delete_filepermission preview read the file before the
user had approved anything, including one outside the working directory
and of any size. It now refuses outside paths and reads only the first
few lines. inside_projectsaid "outside" for any path that did not exist yet on
Windows, where the canonical form of the working directory carries a
\\?\prefix that a plain joined path does not.
Internal
client.rswas 1600 lines holding three wire protocols; it is a module
directory now, one file per protocol (openai,ollama,anthropic)
with the shared message shape and transport choice inmod.rs. The two
text protocols shared 50 lines of copied stream handling, which is now one
TextStreamin `client/stream...
v0.2.0
Added
greptool: optionalcontextparameter shows up to 10 surrounding
lines per match in ripgrep-style blocks, and the system prompt now steers
the model to explore with grep context and ranged reads instead of whole
files.thoth upgrade: downloads the latest GitHub release for the current
platform, verifies the checksum and swaps the binary in place.thoth --continueresumes the previous conversation for the working
directory; the transcript is saved after every turn.@pathin the input attaches a file to the message (tab completes the
path), in the TUI and in-pmode.!commandruns a command yourself and puts its output into the context
without spending a model turn.- "Always allow" now persists per project, and
/allowlists what is
allowed (/allow resetclears it). - System prompt: the environment block now includes today's date and the
git branch with dirty-file count, and new rules cover prompt injection
(tool output is data, not instructions), secrets (never in web queries,
never repeated), committing only when asked, and minimal diffs.
Changed
- Leaner interface: the boxed input is now a single rule, with a state line
showing the spinner, elapsed time and interrupt hint, and a status line
with context use, output tokens and the active editor file. - Permission prompts are scoped. Answering "always" for a shell command
allows that program only (shell:cargo), and web_fetch is allowed per
host, instead of unlocking the whole tool forever. read_fileoutside the working directory,web_fetchandremember
now ask for permission; reads inside the project stay free.- Source layout:
agent/,ui/andtools/modules instead of flat files.
Fixed
- Security: two one-line reads of a file (first line and last line) counted
as a full read and unlocked a blindwrite_fileoverwrite. Reads are now
tracked as line ranges and must cover the file. - Security:
thoth upgradeaccepted any release tag from the GitHub API
and interpolated it into a path that gets deleted recursively; tags are
validated, the temp directory is unique per run, and a missing checksum
file now aborts the upgrade instead of skipping verification. - Security:
@pathno longer grants write access to files outside the
project, and the path is escaped before it goes into the model's context. - Crash: a
%followed by a multi-byte character in a search result URL
panicked the agent task and left the interface spinning forever. The
decoder is byte-safe, and the interface now reports a dead agent task. read_filerefuses files over 2 MB instead of loading them into memory,
shellstops capturing output at 200 kB instead of buffering everything a
runaway command prints, andgrepwith context stops at its output cap
inside a file rather than after it.- Saved sessions and the allowlist are written atomically and are keyed by
a hash of the project path, so similarly named directories no longer
share state.
Full Changelog: v0.1.0...v0.2.0
v0.1.0
Added
- Agentic loop over local LLMs: OpenAI-compatible SSE transport plus native
Ollama/api/chattransport (auto-detected) with per-requestnum_ctx
(default 32768) and separated thinking output. - Tools:
read_file,write_file,edit_file,list_dir,glob,grep,
shell(foreground with timeout, orbackground=truewith pid + log
file),web_search(DuckDuckGo, optional Google CSE with fallback),
web_fetch,problems(live IDE diagnostics),remember. - ratatui TUI: streaming transcript with incremental wrap cache, markdown
rendering, collapsible reasoning, unified diffs with line numbers, full
command-line display, permission prompts (yes / always / no), input
history, message queueing while busy, mouse-wheel scroll, ctrl+o
expand/collapse, live token counters (ctx x/y | out n), turn timer,
editor status ("In file.rs, 7 lines selected"). - Guardrails enforced in code: read-before-edit, full-read-before-overwrite,
duplicate-tool-call breaker, output size caps. - Context management: auto-compact at 2/3 of the window, context-limit
recovery (compact and continue),/compact,/clearwith system-prompt
rebuild, max-turns safety pause. - Memory and recap: project memory in
.thoth/memory.md(viaremember,
loaded into the system prompt), session recap in
~/.thoth/projects/<encoded-path>/last-session.md,/recap,/memory. - Project awareness: startup scan (top-level files + stack detection),
instruction files (THOTH.md>AGENTS.md>CLAUDE.md, short pointer
files followed automatically),/initgenerates THOTH.md. - VS Code integration via the companion extension (thoth-for-vscode):
active file, selection with line numbers, and Problems injected as
context; Windows window-title fallback without the extension. - One-shot mode (
thoth -p "..."),/status,/model,/models,
config file + env vars + CLI flags, GitHub Actions CI (Linux/macOS/
Windows), rustls TLS (no OpenSSL requirement). - Release automation: tag-triggered workflow building binaries for five
targets (Linux x86_64/arm64 musl, macOS Intel/Apple Silicon, Windows),
with install/uninstall scripts for curl | sh and irm | iex.
Full Changelog: https://github.com/thoth-coder/thoth/commits/v0.1.0