Skip to content

Releases: ljedrz/nachalnik

kamchatka 0.17.0

Choose a tag to compare

@github-actions github-actions released this 02 Oct 06:35

added

  • kamchatka reconcile a.json b.json -o merged folds several forks of one session into one.
    The items the forks share are kept whole, and from each fork's own part only the notes the
    agent wrote down for itself, numbered past every identifier any fork handed out. A shared item
    takes the most included state any fork left it in; one fork's revision of it is kept, and two
    that differ are refused. Notes that share a label are all kept, and a pinned system item where
    the forks parted names which came from which fork and which labels collide. The calibration is
    summed where every fork's log names the same model, which -r then carries on with, and
    dropped otherwise; a parameter the forks disagree on is left out. Forks whose session names
    differ are refused, since every way a fork comes about keeps the name. What it writes is a fresh log, beginning with the resume of the
    shared part, and a snapshot -r carries on from; it writes over nothing, and sends nothing to a
    model. reconcile is the first command, so kamchatka reconcile alone is no longer a first
    message, and there is no help command beside it.
  • A context the compactor cannot help is said to be full. ToolTrimmer takes tool results
    alone, so a conversation that is its own bulk grew in silence until the kernel refused a
    request over the limit. The session now says, once, that the context is full and nothing more
    may be taken automatically - with room still under the limit - and what is left is for the
    person to /exclude or the model to exclude through context; and says so again when there is
    room. Shedder::wants_room is its threshold alone, so an unpriced pin does not read as full.
  • The model is told the context is full, where it can act on it. A note naming the context
    tool goes into the context before the next request - inside the turn that filled it - so a
    headless run can make its own room rather than run into the limit. Where the model could not
    make room with it - the tool turned off, exclude and elide refused by a rule, or refused by a
    headless run for want of somebody to ask - the note asks the model to tell the person instead.
    It is worded again, a copy standing in the context included, whenever that changes, and it is
    excluded again when there is room. App::refresh_full_notice is new.
  • A sentence sends the model only to a call it could make now. A refusal from fs named
    shell, grep and write, and answers from setup, log, fork and context named
    context, log and setup, whether or not the session would let the model make that call - and
    a model sent to a tool it did not have named one that does not exist. Each is said now only where
    the tool is on offer and the call would not be refused, by a rule or by a headless run's
    --on-ask deny. Careful::reachable, Careful::permits, Careful::withdraw and
    Careful::unanswered are new; /tools toggle tells the policy what it took away.

changed

  • Requires nachalnik 0.8, nachalnik-mcp 0.8 and nachalnik-providers 0.7. The runtime's
    Event::ContextFull and Kernel::set_full_notice are how a full context is said and how the
    model is told; Compactor::wants_room is what Shedder answers it with; a compaction report's
    uncounted_before and uncounted_after are what its line says it could not price; and
    ModelInfo::with_context_limit is how a limit from the environment reaches the model.

  • An item that is out of the request is excluded, whoever took it out. The context tab
    drew the whole of a shortened output as ▫ archived and an item a newer one replaced as
    ~ superseded; both are - excluded now, and the row's note says why, as it always did.
    /load excludes what was in the context rather than archiving it, an undone note is excluded,
    and context's budget and search say excluded where they said archived. A session saved
    with the old words loads, and reads them as excluded.

  • A projection over the line limit has its longest lines cut down, rather than leaving the
    session unattachable.
    One message larger than MAX_LINE in the context made Message::Attached
    itself that long, and a projection cannot be skipped the way an oversized record is, so every
    client was refused - and a browser's project after any reshape of its chat was answered with
    failed, leaving the chat stale. Nothing is cut until the frame would be too long; then each line
    over the longest length that lets the whole projection fit is cut to it, measured as JSON, and
    the shorter lines are left whole. A line that was cut says how many bytes it lost in
    Line::clipped, a field new to the wire and null everywhere else, and --connect,
    examples/attached.rs and the browser page say so on the end of the line. Where the line is an
    item they also say where the rest is: ?N, or the context view, where an answer can carry it,
    and /save's snapshot where the line alone is longer than any answer. A projection still too
    long once its lines are cut is refused by name as before.

  • The compactor is Shedder, and it is one for a long session rather than a full one. It sheds
    what the conversation is done with, by two rules. Before every request, however empty the
    context is, a tool result whose turn is over goes to a marker, and so does a picture or a document
    the model has been shown, once its exchange is over. Once the context passes --compact, the oldest exchanges go
    whole - the person's message, the turns answering it, their results and a file attached for it -
    until it is down to a target, which --compact now takes as a second fraction - --compact 0.8,0.6, [0.8, 0.6] in a settings file - at most the first and derived from it as before when
    left out. The two are one setting so that a --compact typed beside a file that names both
    replaces both, rather than being held to a target written for another threshold; the shipped
    file says [0.8, 0.6], and a settings file's compact may be either form. Neither rule touches the turn
    in progress, a pin, or a note the model wrote for itself, and a standing summary says how many
    exchanges have gone. Nothing either rule takes is unread, so the marker says the model had read
    it - a model reading one that said only "compacted" disowned its own summaries of the files
    behind it - and that holds for a request that would not fit the limit too, which is not sent
    rather than sent without what the model has not seen. A turn that fills the context is told so by
    the full notice, inside the turn; what goes then is the model's or the person's to say.
    A result the person brings back is left alone until the context is full, where it was elided
    again before the next request: /restore and space leave a note saying who restored
    it. Breaking for the library: tools::ToolTrimmer is tools::Shedder, and it takes more than
    tool results; Setup has compact_target, and Args::compact and Settings::compact are
    the new config::Compact.

  • A /endpoint that keeps the model's name is in the record. It was in no record at all,
    since what the kernel compared to announce a switch carried no address; now model.changed
    names both, and the trace shows model at address either side where the address is what moved.

  • -r says where the record was talking, where that is not where this run is pointed. It
    says so, with the /endpoint that carries on there, and does not follow it: the address is
    where this run's key would go, and a snapshot is a file anybody can hand somebody.

  • A served session has one client at a time, and the newest wins. Every attached client could
    submit, interrupt and answer questions, and a second one typing during a turn took the first
    one's queued line. A client that attaches now takes the session, and the one that had it is sent
    Failed { about: "replaced" } and its connection is closed: --connect exits saying so rather
    than reconnecting, and the browser page stops its own reconnection. The newest rather than the
    first, because the usual second connection is the same client coming back while the session
    still holds its old, half-open one. A connection that never attaches replaces nobody. Breaking
    for a client of the protocol that kept two connections to one session: the first is let go of.

  • Messages typed into a running turn wait in a queue, each for a turn of its own. There was
    room for one, and a second took its place - a person at the desk lost their line whenever a
    client typed after them. Each now waits in order, drawn at the end of the conversation, and
    every turn that ends takes the oldest in; none is merged with another. up takes the newest
    back out. A session resting with some still waiting - after a turn was stopped - puts a new line
    behind them and the oldest in. Attached::queued_behind lists the ones after queued.
    Breaking for the library: App::queued is an iterator over all of them.

  • A command that waits on an endpoint no longer holds the session. /models, /compact and
    y at its question work out what they ask for in a task, as /model and /endpoint already
    settled their switch, and are finished when it comes back - so a served session goes on
    answering its client, a drawn one goes on redrawing, and --deadline, ctrl+c and a client's
    interrupt stop a listing or a pass where they used to wait for the endpoint. A switch is still
    let finish. The line after one of them waits until it is back - the line after /models is
    often the /model it listed - and a client's is answered queued at once and handed in after.
    Down a pipe and from a client, /compact is taken when its pass comes back and the /models
    list is said rather than paged. Breaking for the library: Outcome::Returned is new and is
    handed to `App...

Read more

kamchatka 0.16.1

Choose a tag to compare

@github-actions github-actions released this 29 Sep 19:37

security

  • ./kamchatka.json is read only after a yes at a terminal. The file found in the working
    directory may set mcp, no-sandbox, allow, allow-server, on-ask and the sandbox lists,
    and it was applied as if typed, so running kamchatka inside a repository somebody else wrote
    started their servers and applied their permissions. It is now asked about on standard error
    before anything else, the question naming which of those keys it sets, and anything but y or
    yes runs without it. A run with no terminal on standard input and standard error is refused, and
    --config-file kamchatka.json reads the file as before. A file this program would refuse is
    refused for what is wrong with it before it is asked about. The file under the config directory
    is read without asking.

added

  • --check PATH reads a session's record and says what does not add up. PATH is the log,
    the snapshot, or their name without the suffix, and whichever of the pair is there is read. It
    reports:

    • lines that are not records, and events this version does not know, by name;
    • a session written in a later format;
    • records missing, or numbered twice;
    • calls asked for and never finished, and calls finished that were never asked for;
    • the snapshot's own problems();
    • a snapshot whose items and states disagree with the log replayed to the record it was taken
      at.

    Nothing is started, and anything found makes the exit status non-zero. It is a flag rather than a
    subcommand because the first word on the command line is already a message. check::check is
    the reader behind it, and works on the text of the two files.

  • A release binary can be reproduced, and shows where it came from. The release build remaps
    CARGO_HOME out of the binary. Before, every dependency's source path was compiled in, so only a
    machine with the runner's home directory could rebuild the same bytes. The release workflow
    builds each binary a second time, from another checkout path and another CARGO_HOME, and
    fails if the two differ. On a tag, the archive and the binary get a build provenance
    attestation; gh attestation verify ARCHIVE --repo ljedrz/nachalnik checks it.
    CONTRIBUTING.md has the command to reproduce a binary.

  • shift+enter puts a new line in the prompt. The screen asks the terminal for the kitty
    keyboard protocol, which is what lets shift+enter arrive as anything but enter. A terminal
    that does not speak it sends enter, and alt+enter is still the new line there.

changed

  • A list of items put inside an item is refused as that. Some models write a list as an
    object of one item key, as markup has it - ids: {"item": ["4", "5"]} - and context refused
    it as not a list, or select as not a selector, without saying where the list had gone. Both
    now name the wrapper and show the call as it should be, ids: [4, 5].

  • An edit with nothing to replace says how to add text. A call with a new and no old
    was told only that old is required, and an empty old to give the text to replace. Both now
    say that edit replaces old with new, so adding text means putting a line it goes next to in
    old and that line with the addition in new.

  • A call whose action is true or false is refused in the words every refusal uses. It
    said the action was "a truth value", where every other refusal naming a value's kind says
    `true` or `false`.

  • setup says a long tool result is truncated, as everything else does. Its sentence about the
    output limit said the result "is cut", where the marker in the result, the archived copy's note
    and /help all say truncated.

  • The chat tab reads a tool result only as far as the lines it shows. Deciding whether to mark
    a result as cut short counted every line of it, on every frame, for every result in the
    conversation. It now stops one line past the six it draws.

  • shell says where every call starts. Its description now names the working directory and
    says that each call starts there afresh, so a cd lasts only for the command it is in. Before,
    it said only "in the working directory". A model that could not see where it was, or was trained
    where a shell's directory drifts between calls, began nearly every command with a cd to a path
    it had guessed. Under --advise that cd was one more stage to ask about, and it moved relative
    paths away from the place the confinement note judges them from.

fixed

  • A command at rest writes the session's snapshot once, not once per item it changed. An
    exclusion or a pin over a range, or a compaction, announces every item it moves, and each
    announcement rendered, synced and renamed the whole snapshot again - the same file every time,
    since the kernel had made every change before the first announcement arrived. A snapshot is now
    written only when something was logged since the last one.

  • A call whose arguments were not JSON is refused as that. Where the permission policy
    refused it, having no operation to judge it by, the refusal said the call "names no operation",
    and a model whose arguments had stopped at {"call": read that as being about something else and
    sent the same call again. The refusal now says the arguments were not JSON, shows what arrived,
    and asks for the call again as one JSON object.

  • A link to a directory is not counted by grep or glob. A walk leaves one alone, since it
    reaches the files under it by their own names, and it was counted among the path(s) that are not files as if something had been withheld. A link to a pipe or a device still is.

  • The /compact question counts its items as the rest of the screen does: 1 item, 3 items,
    rather than the item(s) spelling kept for what a model reads.

  • endpoint::configured_limit is None for a KAMCHATKA_CONTEXT_LIMIT of 0, as its
    documentation says, rather than a limit every request is refused against. The program reads
    checked_limit, which already refused it.

  • --serve beside a settings file saying on-ask: Deny serves. The value is read without
    regard to case, and the check that lets the default through beside --serve compared it
    exactly, so Deny was refused as if it said allow.

  • A served session says a line was replaced only when one was. With a line waiting for a
    turn to end, any client's command, or a line said into a session that had gone idle, was
    announced to every client as having replaced the waiting one, which was still there. The
    notice now comes only when the line took the waiting one's place.

  • shell keeps the last line of a command's output when it has no newline and came late. A
    line the command had started, then gone quiet on for longer than the tool's read waits, was
    thrown away when the output ended, and nothing said bytes were missing. It is now kept, as a
    last line without a newline is when it comes all at once.

  • Headless passes a message over when the request it would go in is too long to send. Once a
    turn was refused for a request longer than the model takes, every message a script sent after it
    still went into the context, making the request longer each time, and all of them went out
    together once something made room. While the refusal stands and the next request would be
    refused too, messages are now passed over unsent and said to be once, as they are past a
    --spend ceiling; commands are still read.

  • -r numbers nothing again that the log past the snapshot numbered. A run killed in the
    middle of a turn leaves its snapshot from when the turn began and its log running past it, and a
    session carried on from that snapshot numbered records, items, calls and permissions again from
    where it was taken, so two files of one session named different things with one identifier. The
    numbering is now carried past the log beside the snapshot, the rewrites replayed from it stop at
    the snapshot, and the resumed session says how many records it carried on without.

  • A --connect client told the session is finished leaves when the socket resets. A session
    that ends while a client's last command is still unread closes the connection with a reset, and
    the client took that for a dropped connection: it said so and tried to reattach for a minute to
    a session that had just told it it was over. After session.finished, a reset is the end, as a
    clean close already was.

  • --spend counts a response the broadcast dropped. What a session spent was added up from
    the model.finished events it read, and a subscriber that fell behind a fast stream loses events,
    so a ceiling could be passed by whatever a lag took. It is now read off the log, which drops
    nothing, on the next event that arrives - which is still before the turn asks again.

  • A turn that panics is a turn that failed, and the session goes on. The turn ran on a task
    whose panic sent no outcome, so busy stayed set for good: a headless run waited forever, no
    later line was read, and the record stopped at tool.started. The turn now runs on a task of
    its own, and the task that reports the outcome awaits it and turns a panic into
    Outcome::Failed, saying what the panic said. A tool's panic no longer reaches here, since the
    runtime answers it as a failed call with tool.panicked in the record; what does is a panic in a
    provider or in the kernel itself.

  • A served session out of file descriptors says so once, however it drains. One connection
    taken on the way out of a shortage ended the run of failures, and the next failure was said
    again. Descriptors come back one at a time, so that next failure often came, and
    a_session_out_of_descriptors_says_so_once failed whenever CI was slow enough to show it. A run
    now ends only once a whole second goes by past the retry with nothing fai...

Read more

kamchatka 0.16.0

Choose a tag to compare

@github-actions github-actions released this 28 Sep 09:34

breaking

  • Linux only, on x86_64 and aarch64. Any other target is refused with a compile_error!
    saying so. What makes the shell worth handing a model - Landlock, openat2 beneath a directory
    and the network gate below - is Linux's, and on macOS and Windows the shell ran unconfined behind
    a question read off the command's name. 0.15.1 is the last version that builds on macOS and
    Windows
    , and the one to start a port from; POSTPONED.md says what a port would need. The Mac
    binary is gone from the release, and an aarch64 Linux one is attached beside the x86_64 one.
  • sandbox::Confinement::Unsupported is Confinement::Off, which is what it had come to
    mean: nobody asked for a confinement, under --no-sandbox or before anything was probed.
  • Confinement::complaint is gone. Nothing called it; the Display of a Confinement is what
    every view says it with.
  • Path rules compare names exactly. .ENV and .env are two files here, so the folding that
    made .env* catch .ENV on macOS and Windows went with those platforms. So did reading a
    backslash as a separator: it is a character a name may hold, as it already was to fs and the
    settings file, so secrets/ no longer covers secrets\key, and a rule with a backslash in it is
    taken rather than refused.
  • stopping::Terminated takes SIGTERM and SIGHUP, and no longer has a Windows half.
  • Sandbox::network is a sandbox::Network, where it was a bool: Open, NoTcp (Landlock
    alone, as before), Shut and Asked, the last two behind the network gate below. A confinement
    written by hand that said network: false says Network::NoTcp.
  • sandbox::available hands back a Probed, the Confinement it returned before beside
    whether the gate holds. available(&program).confinement is the old answer.
  • Sandbox has devices and closed fields, and Shell and Setup a devices one: the
    devices under /dev a confined command may read and write - sandbox::DEVICES unless somebody
    says otherwise - and the ports it may not connect to whatever its network is, which
    Sandbox::of fills with the ones this process serves a session on. A confinement written by hand
    says devices: DEVICES.iter().map(Into::into).collect() and closed: Vec::new(), and
    --confine-and-run takes both lists after the read-only paths, each counted the same way.

security

  • A --sandbox-device that leads out of /dev is refused, and not granted if it gets that far.
    The check compared components without settling them, so /dev/../home/you passed as a device
    and the shell could read and write everything beneath it, with no screen saying so. A device is
    now resolved, .. and links alike, before it is checked and again before the ruleset grants it,
    and /dev itself is refused.
  • A settings file found underfoot is said on standard error before anything it asks for is
    done.
    It was said only in the conversation, which exists after the MCP servers it names have
    been started - so a kamchatka.json in a cloned repository could run a program before anything
    said the file had been read. The conversation still says it.
  • A --sandbox-read path that is not there yet is held to the same rule as one that is. The
    check that refuses a read-only path inside a writable one skipped a path it could not resolve,
    so --sandbox-read build in a working directory with no build yet was accepted, and once
    anything made it, it was writable while every screen called it read-only.

added

  • /undo and /redo. Undo was u and U on the context tab and nothing else, so a
    --headless run, a --connect client and a browser had no way to take a change back - while
    /exclude, /load and the over-budget line all told somebody to press u. The two commands
    are the same two calls into the same function, and the messages now name both.
  • fs read takes from and lines, and a long file is read in parts. A file past the output
    limit came back cut at a byte, with [... N bytes truncated ...] and nothing saying where it had
    broken off, so reading on meant guessing a sed -n through shell - an exec:run call, for a
    file the session could already read. read now stops at the last whole line under the limit
    (/limit fs:read) and its first line names those lines and the from to read on with. A file
    past 8 MiB, which was refused, is read a part at a time like any other, and says that more
    follows where it would take too long to count how much. A file that fits, read whole, comes back
    as it always has.
  • A command is asked about the network when it tries, rather than for what it is called. On
    Linux, on x86_64 and aarch64, the confined child installs a seccomp filter - the new gate
    module - that holds every socket() for AF_INET or AF_INET6 until the program answers: from
    net:reach where it says allow or deny, and from the person, once per command, where it
    says ask. A running command's question stands in the prompt's place, headed a command wants the network, and takes y, a and n; a headless run answers it with --on-ask while the
    turn is still running, a --connect client once its input has closed, and a served session
    sends it to every client as reaching and takes reach back. The answer is a policy.ruled
    for net:reach, and the tool result the model reads says whether the command was let out or
    refused, and whether by an answer or by a standing rule - not who answered, which may have been
    --on-ask.
    Where the gate holds, Careful stops reading a command for program names, so an allowed
    exec:run runs git status unasked and a script that opens a socket is asked about. Where it
    cannot be installed nothing changes. App::reached, App::decide_reach and Careful::reaching
    are how an embedder with a loop of its own reaches the question.
  • A refused network is refused for UDP too, where the gate holds. Landlock can refuse TCP and
    nothing else, so a confined command could send a datagram with net:reach denied; the gate
    refuses the socket, and the words that promised no TCP there say no network.
  • The permissions tab says whether the network is gated, after the confinement on the shell's
    line, so a session that fell back to reading command names says so.
  • --sandbox-device and sandbox-device name the devices under /dev a confined command may
    read and write
    , replacing the usual five rather than adding to them; the shipped
    kamchatka.json lists those five. A path outside /dev is refused where it is given. Naming
    /dev/ptmx and /dev/pts has ptys back, and every other terminal of the person's with them.
  • /params KEY null takes a parameter away. There was no way back from one once set, short
    of /restart; a null was sent as a null. The log records the parameters left, as it does when
    one is set.

changed

  • glob answers in the shape of ls -R. Every matching path came back whole on a line of its
    own, so a pattern matching most of a tree - the usual one - spent most of its answer writing the
    same directories again. Each directory is now written once, as ls -R writes it, followed by
    the names in it that matched. The directories come in the order ls -R walks them, so src
    and the files in it come before src/app, where the paths were alphabetical before. The tool's
    description says the shape. glob's description also says to name the extension being looked
    for, *.rs rather than *: a live model listed a whole tree, met the cap, and guessed a path it
    had not been shown.
  • fs says it does not go through a shell, rather than that there is no shell here. The note
    on ~ read as a fact about the session, and a live model offered shell beside it said it had
    no way to run a command. The walks' "with no shell in front of them" is now "not through find
    or grep", and the refusal of a ~ path says fs does not go through a shell, for the same
    reason.
  • fork says its answer comes back to the model alone, and that forks in one turn are
    independent.
    It said "nobody has read what it said", which a person watching the fork stream
    would dispute; and a model asking two forks together had nothing telling it the second would not
    see the first's answer.
  • glob stops one past its cap. It walked the whole tree after its 200th path to say how many
    there were, which on a large tree was most of the call; it now stops at the 201st and says there
    are more, which is what a model acts on whatever the number.
  • --deny fs:write makes the --sandbox-allow paths read-only for shell too. fs was
    refused a write there and a command was not, so the two tools disagreed about one path, and the
    one that could write was the one whose writes nothing checks. They stay readable; a socket among
    them, which connecting needs write access to, is out of reach with them, and the shell's
    description names them read-only.
  • A confined command cannot set up an io_uring, where the gate holds. A ring opens a socket
    without calling socket(), so io_uring_setup is answered ENOSYS, which is what a program
    that can use one falls back from.
  • context budget no longer explains elide and exclude under its table. The same tool's
    schema says what each does, on every request; the footer keeps what its fifth column adds up to
    and what giving up a row that holds something frees.
  • A change refused for want of a reason is one sentence. It says nothing was done and which
    operation to call again; what a reason is for is in the argument's description, which every
    request carries.
  • context's description no longer says that each operation describes itself.
  • A file's value is refused with the file named, wherever the refusal happens. Everything
    after the merge -...
Read more

kamchatka 0.15.1

Choose a tag to compare

@github-actions github-actions released this 24 Sep 13:05

security

  • A server rule naming no server this run starts is refused. --deny-server filess beside
    --mcp files=… matched nothing, and files was left to the question - which a headless run
    with --on-ask allow answers yes - under a rule that read as given. An --allow-server naming
    no server is refused the same way, as a rule about a domain no tool declares already is.

  • A --sandbox-read path inside a writable one is refused, rather than drawn as read-only and
    written to.
    Landlock only adds to what a process may do and fs checked the writable roots
    first, so a read-only path inside the working directory could be written by both tools while
    every screen called it read-only.

  • A session named with a path is refused. The name is the stem the record and /save DIR/
    write under, and a resumed snapshot whose name was ../../escaped had its log and snapshot
    written outside the private directory they belong in.

  • A chain of links too long to follow is refused by fs, not taken for a file to be made.
    The path check follows links up to a bound, and a chain dangling past it had its last link read
    as a file about to be created in the directory it sits in - allowed, and then followed out of
    the directory by the open, on a platform with no openat2 to stop it there (macOS, Windows, and
    Linux older than 5.6).

added

  • Careful::offers: what one of the session's own tools declares, so that a call to it naming
    no operation is refused saying so. wiring tells it about every tool a session starts with.
  • mcp::reinstall: the tools of servers that are already running, installed into another
    kernel under the rule attach holds, with a sentence for each server left out.
  • wiring::Flagged: where a provider was pointed when the session began, and restore to put
    it back - for a loop of its own that relaunches a session with Setup::relaunch.

fixed

  • A call whose arguments are a level too deep is read as the call. {"call": {"item": {"action": "read", ...}}} named no operation as written, so it was judged against everything
    the tool does and refused, or told action was missing; a model that wrote every call that way
    spent whole sessions refused. A wrapper whose one entry is an object naming an action is read
    through, by the rules and by the tool alike, as a wrapper written as text already was.
  • A call refused by a rule because it named no operation is told so. A call whose action
    could not be read declares everything its tool does, so --deny fs:write refused a call that
    meant to read, and the model was told only that a standing rule refused it - and retried the
    same call. The refusal now says first that the call named no operation, and what it held
    instead, so arguments one level too deep read as the mistake they are.
  • The context tab with f on projects the context once a frame. It built the request a second
    time to decide which rows to list, beside the one the frame had already built for the rest of
    the screen.
  • /restart leaves out an MCP server whose tools would take another tool's identifier, and says
    so
    , as starting the run refuses it. The servers' tools went back into the new kernel one after
    another, and a server whose list had changed since the run began could displace another's tool
    without a word.
  • A message sent into a turn that fails, or into a /step that comes to rest, goes into the
    context when it ends.
    It stayed waiting, so the next message sent ran first and the waiting
    one went in after that turn, or never.
  • --connect asks a question again when its answer went down with the connection, instead of
    forgetting it and leaving the session waiting on an answer the client could no longer give.
  • A shell call dropped before its command ends removes the command's temporary directory, as
    it already killed the command's process group. It was left behind with whatever the command
    wrote there.
  • A confined command whose temporary directory could not be made is given no TMPDIR, as
    documented, rather than the directory it could not be made in.
  • An undo of several steps can walk past a pin it put back itself. Which pins were the
    model's was read once before the walk, so walking back a pin and the restore after it
    stopped halfway, left the item pinned, and called the model's own pin the person's.
  • context with look and a select marks what it cannot price. The rows a selector matched
    showed an image as costing nothing, with no + and no word that the total was a floor, where the
    full listing said both.
  • mcp::attach installs nothing until every server has answered. A name collision was refused
    only after the second server's tools had replaced the first's, and a spec that failed part way
    left the earlier servers' tools registered while their servers went with the Err.
  • A streamed fragment no longer copies the whole context list to find the item its line
    follows.
    Only a fragment that starts a new line asks, and it reads the last identifier in
    place.
  • A search open on the context tab no longer matches every item on every frame. Each frame
    made the whole text of every item and matched it again, which in a long session redrawn several
    times a second during a turn cost more than the rest of the frame. What a query said about an
    item is kept until the query or the item changes.
  • A key on the context tab no longer filters the whole context again. Every key recomputed the
    rows the frame had just drawn, which with a search open doubled what a key cost in a long
    session. The first key after a frame counts rows in what that frame drew.
  • A long chat stops copying every answer into every frame. The answers kept from one frame to
    the next were copied whole into each frame and all but a screenful thrown away, which cost more
    than drawing the rest of the chat. A frame now points at the answers it keeps and makes lines
    only of the rows on screen.
  • context request sizes a message by what it sends. The bytes column counted the text a
    message said, so a turn that only called a tool read as 0 however large its arguments, its
    thinking was not counted, and a picture was as long as its name. It now counts the content, the
    calls and the reasoning, as the token count beside it does.
  • A shell command that leaves something running on its standard output still answers.
    sleep 60 & echo started held the call for a minute, and a server started with & held it
    until somebody pressed escape: the output was read to its end, and a background job keeps it
    open. A quiet moment after the command itself has ended is now the end of what it said, and the
    result says the output was still open.
  • A local advisor slow to load is waited for rather than closed. The warm-up ran under the
    thirty seconds one question may take, so an engine that needed longer to load its checkpoint was
    closed before its first question and stayed closed for the session. The warm-up now has ten
    minutes; a question asked meanwhile waits a question's thirty seconds for its turn and goes
    unrated if it does not get one, leaving the advisor open. Dropping the advisor mid-load lets go
    of the engine at once.
  • log refuses an action that is not text, and a kinds entry that is not a name. The
    first was read as a bare read, and the second was dropped, which narrowed the filter to the
    names left beside it.
  • fs and shell refuse an argument of the wrong kind rather than guess. A grep or glob
    whose path was not text searched the working directory, a grep whose glob was not text
    searched every file, and a shell call whose action was not text ran its command.
  • fork ask refuses to leave out an item that is not there. It asked the copy with the whole
    context and paid for the request, as though it had been an ablation.
  • log does not give an item that never existed a history. Asked about an id the session has
    never held, it said the item was already in the context before the log began; it says there is
    no such item.
  • The context tool's counts and search say what the context holds. look counted as going
    into the next request every item in a state that sends, including ones the projector repairs
    away, where the table under it said otherwise. search read an item's text and not what a turn
    thought or called a tool with, so it answered "no line says it" about an argument the model
    had passed. And revise to the words an item already had reported a rewrite, and journalled one
    for undo, though nothing had changed.
  • A served session leaving takes away only its own socket. It removed whatever file was at the
    path it had bound, so a socket removed by hand and bound again by a second session was taken
    away when the first one left.
  • --connect answers each question once. An answer typed at the client left the question at
    the front of its queue until the record saying it was decided came back, so a second y typed
    before then answered the same question again, was refused, and left the next one unanswered.
  • /load waits for calls set aside to be run or cancelled. It was refused only while a turn
    ran or a question waited, so after /step had left calls in Ready it loaded, archived the turn
    that asked for them, and the next step ran them against the loaded context.
  • An interrupt holds, and one with nothing running is not saved up. A /spend lowered under
    what a turn had already spent stopped nothing, and the turn went on to its next request. An
    interrupt sent while nothing was running - a client's ctrl+c between turns - stayed set, and the
    next message's turn spent it and ended without asking anything. And a stop that landed while a
    turn's answered questions were being carried on was und...
Read more

kamchatka 0.15.0

Choose a tag to compare

@github-actions github-actions released this 23 Sep 16:44

breaking

  • --advise rates commands and decides nothing, and it is shell-advisor's. The advisor was
    asked about every call the rules were going to allow, and could refuse it or turn it into a
    question: a setup call in a session started with --allow setup went to a person, and
    answering always added a setup:policy row that changed nothing. An allow is the user's
    decision, so nothing reopens it now - what the rules allow runs unasked and nothing about it is
    sent. The advisor is asked only where each shell command a question is about lands on the
    rubric, which is drawn in the question. Feature advise is the System One client alone and no
    longer adds --advise; shell-advisor does. Advised::said is gone with the sentences it
    returned.

  • remote::Serving::last takes the Serving and hands back a future to await. It says the
    last lines, as before, and the future is the session waiting up to two seconds for its
    connections to write what they still owe and close. A loop of its own that ends a served
    session calls serving.last(&app).await where it called serving.last(&app).

  • introspect::install hands back an introspect::Installed, and App::introspect holds
    one, where both were an Arc<Kernel>. It keeps the same handle the tools reach the kernel
    through, and Installed::forked says what forks have been charged, for the spend ceiling to
    count. A caller that held the value to keep the tools alive holds it the same way.

added

  • A command the advisor could not rate says so. Where the rating would be, the question says
    the advisor could not rate this and why, on the screen and in a projection's new unrated
    list for a client to draw. The failure is not random: the advisor's endpoints sit behind a
    firewall that refuses requests by what is in them, and what it refuses - /etc/shadow, a secret
    piped to curl - is the command a colour is most for, which drew no line at all and read as one
    nobody had anything to say about. Advised::why_unrated and App::why_unrated answer it, and
    the browser page draws it.

  • The readme shows the program. Two screenshots, of the chat and context tabs, are woven into
    what it says about the tools a session manages itself with. They are linked by address and kept
    out of the published crate.

  • sandbox::confines_unix_sockets answers whether this kernel refuses a confined command a
    connection to a unix socket outside what it may write. Landlock grew the right for it in ABI 9,
    which is Linux 7.1, and below that there is none to ask for. It is a question rather than an
    assumption because handling a right the kernel does not have costs the whole ruleset its Full
    status, and a sandbox that calls itself partially enforced over a right it was never going to
    enforce is worse than one that says what it does.

  • A question about every operation a tool has says why it is. A call that names no operation
    declares all of them, which is the strictest reading of a call nobody can place and the right
    one - but nothing said so, and a session started with --allow fs:read was then asked about
    fs:write with no way to see that its rule and the call could never meet. The permission panel
    and a headless run now say which it is: a call that named nothing, or one that named something
    the tool does not do, and what a rule that would have answered looks like. App::widened
    answers it for a client of its own.

  • A release attaches a Mac binary as well. aarch64-apple-darwin, built on macos-latest,
    which is that runner's own host - so it is a second entry in the matrix rather than anything
    cross-compiled. It is unsigned: Gatekeeper quarantines a download and the first run is
    refused until xattr -d com.apple.quarantine kamchatka, which is what a paid Apple Developer
    account would fix and there is not one. It runs the shell unconfined, as any kamchatka off
    Linux does.

changed

  • Requires nachalnik 0.7, nachalnik-mcp 0.7 and nachalnik-providers 0.6. The runtime's
    undo and redo answer a Result, which is how u and U learn a turn is under way;
    record_rule and provider_changed are how the rules and the switches reach the record;
    Snapshot::problems is what -r and /load check a snapshot with; and the local advisor
    renders its request with system1::render.

  • The rubric draws its line at the working directory. The middle level was "leaves something
    changed that could be put back", and nearly everything could be: chmod -R 777 /, a global
    npm install, git config --global and an edit to ~/.bashrc were all drawn yellow beside
    cargo build. Yellow is now a change to files inside the working directory, the way git or a
    rebuild could undo, and red is anything outside it - system files, permissions, the home
    directory, the machine, another account, a remote service - as well as what cannot be got back
    and what is sent off the machine. The bottom level names reading, listing, searching and
    changing directory, so a cd is no longer one answer away from yellow. The destructive claim
    asked beside the rubric draws the same line, and the panel's band names say it.

  • What a tool keeps of one call has a ceiling: 8 MiB, tools::KEPT. The output limit decides
    what the model is shown and the whole is archived beside it, so the whole was whatever arrived -
    a yes nobody stopped grew the process, the archive and every save after it without end, and so
    did the chat tab showing it stream in. Past the ceiling shell goes on reading each stream to
    its end and lets it go, a line that never ends included, and the result says how many bytes
    went under the status line, where no output limit cuts it; fs refuses a file past it with a
    sentence pointing at grep and at head, tail and sed -n.

  • A command the model runs is not handed this program's keys. KAMCHATKA_API_KEY,
    OPENROUTER_API_KEY, OPENAI_API_KEY, KAMCHATKA_SYSTEM1_API_KEY and TYPESAFE_API_KEY are
    taken out of the shell tool's environment, confined or not, so printenv no longer puts the
    key paying for the session into the context, the request and the record. The list is
    endpoint::KEYS. An MCP server and a local advisor still inherit them: each is a program the
    person chose, running unconfined, and a server with an API of its own may read one of these
    names for it. A command that needs one of these names is handed it on purpose, in a file it
    reads, rather than inheriting the session's.

  • tests/remote.rs is tests/remote/, the way tests/screen/ and tests/introspect/
    already are: one binary named for the directory, main.rs holding what every file in it reaches
    for - a served session, the Peer that speaks the protocol by hand, the two tools that answer
    slowly - and a file for each thing it is about, along the seams its section banners already
    drew. Not one test changed.

  • The sandbox suite is two: tests/sandbox.rs for what needs a process, tests/boundary.rs for
    what does not.
    The first is #![cfg(target_os = "linux")], as it was; the second is
    #![cfg(unix)], and holds the six tests that never spawn anything - which paths Reach admits,
    what a refusal names, the ~ refused in words rather than expanded, what a confinement travels
    as on a command line, what the scratch directory may be made through, and which errors
    Sandbox::note_for will claim. All six were behind the Linux gate because the file was, so the
    macos-latest column in CI passed without running any of them. Nothing moved in src.

  • The docs say that tools::Careful::new asks about everything rather than allowing reads, that
    an MCP tool is judged as its server where it declares mcp:call, that config::Settings has
    two keys with no argument behind them, border and tools, and that a file attach has no
    name for is read as text and refused if it is not text.

fixed

  • The empty permissions tab says which answer puts a row there. It named a and n, and only
    a records anything; y and n answer the one call. It also said the path rules bind read,
    write and edit, where they bind every fs operation handed a path, grep and glob
    included.

  • A chain's destructive link is underlined when the whole command reads as destructive too.
    The fold kept the first reading to reach the worst band, and the whole command's comes first, so
    a tie - which is what cargo build && rm -rf ~/.ssh is, since a destructive link makes the whole
    read as destructive - pointed at nothing. The stage is kept on that tie now, and the whole
    command only where no stage reaches its band.

  • A running turn says esc stops it once. The chat tab said it in its footer and on the
    status line beneath; the status line is on every tab, so it says it there alone.

  • up in an empty prompt scrolls when there is nothing to put back. It recalls the last
    message, and before one had been sent it did nothing at all - the one prompt from which the key
    that scrolls at the top of any other did not. It scrolls there now.

  • log refuses an ids entry that is not an item number. It kept the numbers it could read
    and dropped the rest, so ids: [12, -1] answered as a filter on item 12 alone; only a list with
    no number in it at all was refused. It reads ids the way context does now, refusing the call
    and naming the entry.

  • A client that ends a served session is told it did. The connections were tasks nobody
    waited for, so a host that exited on a /quit could take the answer to it and session.finished
    with it. The client that typed /quit read the closed socket as a drop and went looking for a
    session that was gone. The session now waits for its connections before it returns.

  • A client whose session has gone gives up on it. An attempt that...

Read more

kamchatka 0.14.0

Choose a tag to compare

@github-actions github-actions released this 23 Sep 16:44

breaking

  • The assisted-shell feature is shell-advisor. The last of the "assist" spelling, which
    is a word about the other kind of model. One mechanism with two names is a bug by this
    repository's own rules, and the advisor is what this one is called everywhere else.

    Nothing about what it does moved: --advise in a build carrying it still asks for the verdict
    and the rating both.

  • The advisor's three variables are KAMCHATKA_SYSTEM1_*, where they were
    KAMCHATKA_TYPESAFE_*: KAMCHATKA_SYSTEM1_API_KEY, KAMCHATKA_SYSTEM1_BASE_URL and
    KAMCHATKA_SYSTEM1_MODEL. A session exporting the old names loses its advisor with the message
    that asks for a key, rather than failing quietly.

    TYPESAFE_API_KEY, without the prefix, is unchanged and still read - it is TypeSafe's own
    documented variable and not this program's to rename. The three that moved are the ones this
    program invented, and they name the kind of model rather than the company selling one: the
    question types a System One engine answers are the category's, the open ones arriving now have
    the same three, and KAMCHATKA_SYSTEM1_BASE_URL has always been able to point somewhere else.

    endpoint::advise::Account follows, with TypeSafe and OpenRouter becoming Dedicated and
    Borrowed - which is what the two cases were about all along. One is a key held for the
    decision service, wherever that is pointed; the other is the conversation's own key, and it is
    the only one carrying a rule about where it may go. That rule is unchanged, and so is the test
    that holds it.

  • tools::Rated has a worst field and is #[non_exhaustive], which it should have been from
    the start - it is a struct this crate answers with and nothing outside it builds. Reading it is
    unchanged; building one by hand is what stops compiling, and protocol::Judged grew the same
    field under the same rule.

  • protocol::Message has an Oversized variant, and protocol::write is split into framed and
    write_frame - which is what lets a caller ask how long a message is before sending it. Adding
    the variant is not a break on the wire, because Message::Unknown is what an older client reads
    it as and reads as nothing; it is a break for anything in Rust matching on the enum, which is not
    #[non_exhaustive].

  • endpoint::connect and endpoint::gemini::connect take Option<&str> rather than anything that
    becomes a String, because a session may now have no model. None builds the client and skips
    the probe, since there is nothing to ask a context limit about; an embedder that always names one
    wraps the argument in Some.

  • wiring::Setup has a record field, true by default, which is what --no-record turns off.
    A Setup built with ..Default::default(), the shape its own docs give, is unchanged; one
    written out field by field names it now.

added

  • wiring::record and Setup::relaunch: the safety net is the library's. Every run of the
    program writes its session out when it ends, and /restart writes the old one out on the way to
    the new one. Both were private to main.rs, so an embedder driving an App with a loop of its
    own had neither, examples/phone.rs included. record writes the log and the snapshot under
    the temporary directory, 0700, beside rather than over a session of the same name, and says
    where; Setup::relaunch ends the session it is handed, records it unless Setup::record is
    off, and wires a fresh one out of the same settings under a fresh name. main.rs calls both
    where it used to have its own.

  • /restart writes the session out and puts a fresh one in its place, which is what quitting
    and running the program again would do, without quitting. The old session's record is the same
    one the end of a run writes - a session somebody restarted is a session that ended, and a run
    abandoned halfway is the case that safety net is most for - and the line naming it is the first
    thing the new session says, because the new one is the only place left to say it in.

    It goes back to the flags, not to where the session had got to: the model is --model
    again, the permissions are --allow and --deny, a tool /tools toggle switched off is back,
    --system and --files are re-read, and the context is empty. What carries over is what cannot
    be rebuilt cheaply or at all - the provider connection, the MCP servers, whose tools are
    installed into the new session rather than their processes respawned, and the sandbox, which is
    a ruleset that cannot be lifted once applied. A session started with -r restarts into an empty
    one: the snapshot is where the run began, and the command is somebody saying they are done
    with it.

    A running turn is stopped rather than the command refusing until it ends, because a model that
    has found a loop is the commonest reason to type it.

    It works in all three loops. A piped run keeps reading the same input, so the lines after it run
    in the new session - one reader for the run rather than one per session, or whatever the old one
    had read ahead goes with it. Clients attached over a socket are disconnected, since a watermark
    is a record number in a log that is gone; examples/browser.html retries and comes back into
    the new session by itself.

  • SYSTEM1_ADVISOR_COMMAND: an advisor running on this machine, and nothing leaves it. Set
    it to a command line and it is used instead of the three KAMCHATKA_SYSTEM1_* variables,
    which are then not read at all. kamchatka/contrib/laya_advisor.py is the one to point it at:

    $ pip install laya
    $ export SYSTEM1_ADVISOR_COMMAND="$HOME/ai/venv/bin/python kamchatka/contrib/laya_advisor.py"
    $ kamchatka --advise

    This is the honest fix for the disclosure --advise is careful about. Everything
    tools::advice says is about a third party reading a tool's arguments - which for a write is
    the text being written and for a shell call is the command line - and pointed at a local
    engine there is no third party. The flags still say what they say, because what a flag turns
    on must not depend on an environment variable, but the thing it guards against is not
    happening.

    It is a command and not a path because laya ships
    no executable: no HTTP server, no CLI, no python -m laya. It is a library, so the smallest
    thing that reaches it from another process is an interpreter and a script, and the shim is
    that script - about forty lines, most of them comments, because laya's question dicts and
    answers already use the same three types under the same names as the hosted engine. The
    protocol is therefore not a new one: it is the body Jev::render builds, one JSON object per
    line, read back by the same Answers the HTTP path uses.

    Both of the child's streams are held rather than inherited, which matters most on a first run:
    laya downloads a checkpoint and says so at length, and a child sharing the terminal writes
    over the screen ratatui is drawing. Piping it is not enough on its own - a pipe nobody reads
    fills and blocks the writer, so the engine would stop answering while writing its own
    diagnostics - so stderr is drained for the life of the child and the last twenty lines are
    kept, hung on the end of whatever failure they explain.

    The process is started once and kept. A 421M-parameter checkpoint costs seconds to load and
    milliseconds to run, so loading one per question would put that wait in front of somebody
    deciding whether to press y - which is the one thing this kind of model was chosen for not
    doing. It is killed with the session, and its stderr is inherited rather than piped, so what
    it says about itself reaches the terminal instead of filling a buffer nobody reads.

    Every failure closes the pipe: a command that is not there, a child that stopped answering, a
    line that did not parse, a question that took longer than 30s. Not because one lost question
    matters, but because the next read off a doubtful stream is the answer to the question
    before it - an advisor confidently rating the wrong command is worse than no advisor. After
    that the standing rules decide alone, which is what an unreachable hosted advisor already did.

  • The released binary is built with shell-advisor. It is the one feature that has to be
    compiled in to exist at all, and somebody who downloaded a binary cannot add it afterwards - so
    the download was a --advise that the readme documents and the artifact did not have.

    Nothing is sent anywhere without --advise on the command line, which is the disclosure and
    always was. What the feature gate buys is keeping a third-party dependency out of builds that
    do not want one, and a released binary is by definition not one of those. What it costs is that
    the two opt-ins become one for anybody using this binary: in a build carrying both, --advise
    turns on the verdict and the rating, and the rating is asked about every command the model
    writes rather than only the ones the rules would allow. The flag's own --help names that
    wider disclosure; building from source with --features advise alone is still the narrower
    one.

  • A command joined at its |, && or ; is rated stage by stage and drawn as its worst
    stage, with that stage underlined.
    One score for a whole command line is the reading a long
    chain is worst served by: three quarters of cargo build --release && cargo test && rm -rf ~/.ssh is the work an agent does all day, and it is the last quarter a person answering has to
    see. Every stage is a question in the same request, evaluated on its own against the same
    state, so a twelve-stage pipeline costs the round trip a one-stage command costs and no stage's
    answer can be moved by another's.

    The fold is over what each stage is shown as and not over...

Read more

kamchatka 0.13.0

Choose a tag to compare

@github-actions github-actions released this 19 Sep 06:46

added

  • --advise borrows KAMCHATKA_API_KEY where no dedicated key is set and the session's own
    requests already go to OpenRouter
    . jev is served there as well as by TypeSafe, so in that one
    configuration the key already paying for the conversation can pay for the questions too, and a
    session that was not given a second key is no longer a session that cannot have an advisor.
    KAMCHATKA_TYPESAFE_API_KEY is still checked first, so a session holding both pays TypeSafe.

    The condition is the whole of what makes it safe, because a key is an OpenRouter key by virtue of
    being sent to OpenRouter and not by virtue of the variable it was read from. A session pointed at
    ollama, at Google with --gemini, or at a gateway of somebody's own holds a key that service
    issued, and spending it here would hand a third party a credential with no business with them -
    which is what the old refusal to fall back was protecting, pointing the other way. Those sessions
    are refused, and told which address the refusal was about rather than being told they need a key
    while holding one.

    What the fallback does widen is who is told: the arguments --advise already sends off the
    machine go to OpenRouter as well as to the model behind it, which is why it is in --help beside
    the variable rather than left to the readme.

    The endpoint and the model follow from whichever key was found rather than being read
    independently, since three settings that can disagree are three ways to send a key to a service
    it is not for. KAMCHATKA_TYPESAFE_BASE_URL and KAMCHATKA_TYPESAFE_MODEL still override each,
    and pointing one at the other service means setting the other too.

  • Feature assisted-shell, which has the advisor place each command a question is about on a
    three-level rubric - it only looks; it changes something that could be put back; it destroys
    something that cannot be got back or sends something off this machine - and draws that in the
    question in green, yellow or red. It is on top of advise rather than beside it and needs
    --advise at runtime as well. The line sits in the header, above the arguments and inside the
    region that does not scroll, because a warning that can be paged out of sight is one nobody has
    to have seen; the confidence is printed beside it, since a band is not a fact about the command.

  • The rating decides nothing: it is never folded into a verdict, so a session with the feature
    on refuses and allows exactly what the same session without it does. Two rules keep it honest in
    the other direction. A score is read by the level it is nearest rather than the one it has
    passed, so a command mostly on the top level is drawn as being on it; and a rating the advisor
    was not sure of is never drawn in green and never drawn safer than it scored, on the grounds
    that a distribution spread across a safety rubric is the advisor saying it could not tell, which
    is not the same as saying a command is safe.

  • What it costs is a wider disclosure than advise alone, which is why it is a second opt-in and
    not part of the first. advise sends a call's arguments only where the standing rules were
    going to allow it - in a default session, not one command, since exec:run is a question. A
    rating is asked for where they were going to ask, which is every command the model writes. A
    call heading for a refusal is still sent nowhere: it has no question to colour, and rating one
    would hand over the arguments of a call that was never going to run.

fixed

  • u and U work on a context pane a filter has emptied, which is where they are most needed.
    The keys that pick a row need one, so the handler returned early with nothing listed and took
    those two with it: with f on, hiding the last row on the screen removed the row and the key
    that would put it back, and the way out - press f first - is written nowhere.
  • The four tab shortcuts reach past an open search box. The box takes the keys while it is open
    and read every character without ctrl as one of its own, alt included - so alt+2 typed a
    2 into the query instead of going to the context tab, which /help promises it does
    everywhere. The handler's own note said a modifier means somebody reaching past the box.
  • The context header counts what f is holding back rather than what a search is also hiding. The
    figure was every row the list dropped, which with a query running is the two filters together,
    under a label naming f as the reason - so items going into the request were reported as not
    being sent. The empty pane has a note about not making that claim; this was the same claim one
    branch further on.
  • A pipe inside a table cell stays inside it. Every pipe was read as a column boundary and every
    pipe was trimmed off both ends, so \| moved each value after it one column left and the row
    was then cut to the header's width, dropping whatever fell off - and a row opening with an empty
    cell lost it. A table drawn from an answer has to say what the answer said.
  • A long trace detail wraps instead of running off the right edge. It was wrapped against the whole
    pane and then had the clock and the name column put in front of it, which is the one pane whose
    promise is that a detail wraps rather than being cut.
  • setup tools says whether what is cut is kept, rather than promising it always is. policy
    reads that setting and this stated the opposite, so a session run --forget-truncated got two
    answers from one tool two actions apart - and the one it was likelier to read sends a model
    looking for content the session was told to drop.
  • setup tools says what cuts a tool that has no row in the limits table. Those are keyed by
    subject and a tool from a server declares none of them, so a session offering nothing else read
    Nothing here cuts an answer short while the kernel cut every one of them at its own ceiling.
  • A settings file is answered before an endpoint is reached. Setup::check is what wire asks
    first and what main asks before it builds a provider, so a file naming contxt, or a path rule
    nothing can match, stops the program by name rather than after a round trip - or, with no API key
    anywhere, rather than being reported as a missing key.
  • A second message sent into one running turn says that it replaces the first. The slot holds one
    and the newest wins, which is a decision - being silent about it is not, since the waiting
    message is drawn at the end of the conversation, so the second took that row away and put its
    own there with nothing said. up reaches what is waiting, never what it replaced.
  • /step with a message, typed into a running turn, puts nothing into the context. The text was
    pushed before anything had said whether it could step, so it landed in a turn already running -
    the shape a plain message is held back from, because an item between a call and its result is
    one most of these APIs refuse - and the step was then declined in silence.
  • log's take counts records, which is what it says it counts. It counted rendered lines, and
    whole prints a replaced item's old text entire - so take: 1 against a record holding three
    lines of it handed back the last of those lines, with no sequence number and no event name in
    front of it, under a header calling that one record.
  • fork and setup refuse an argument the action they name does not read, which the other tools
    have done since the last release. without belongs to ask, so a draft carrying one bought a
    request whose answer read as the experiment the caller asked for and was not one - an ablation
    nobody performed is read as evidence. shell asks the same question now, from the same table its
    schema is built from.
  • A fork counts the caller's items rather than its own. The count was taken after this tool pushes
    the copy's system instruction, and after ask pushes the question, so the figure moved with
    which operation asked for it - and it is there to be compared between runs.
  • setup permissions says what an undecided rule means, which is not the same for the two kinds.
    One sentence said both stop and ask "whatever the rows above say": true of a server, which is
    consulted beside the rows, and false of a domain, which an exact rule answers for. A model told
    that a read it is allowed will stop does not try it.
  • A fork that was given no without says nothing of the caller's was taken away, rather than that
    the copy saw everything. The projector repairs the unfinished call out of the copy, so the
    stronger sentence was not true; what the line is for is telling an ablation from a question that
    merely asks the copy to disregard something.
  • An undo does not walk back over a decision the person has made since. Every other move in
    context asks whether an item is theirs to move - a system instruction, a pin they put on - and
    this one went straight to the kernel, so a pin made after the model elided an item came off again
    on the model's next undo. Silently, and against the one thing the word promises. What it leaves
    alone is named in the answer.
  • Walking a move back is one undo for the person, however many items the move named. It set each
    item's state in a call of its own, so undoing what this tool reported as one change left three
    checkpoints on their stack; the items are grouped by the state they return to. The report names
    each of those states rather than the first one for all of them.
  • ids refuses a number that is not an item number, and reads the same number twice as one item.
    It dropped whatever it could not read, so [-1] arrived as no items at all - which is how a call
    naming none arrives - and search with one bad number searched the whole context. The refusal
    for naming items twice reads ids and select as given now, rather than as what they came to:
    a call with both, one of wh...
Read more

kamchatka 0.12.0

Choose a tag to compare

@github-actions github-actions released this 17 Sep 20:55

added

  • /compact, which asks the compactor by hand and shows its answer before applying it: every item
    it would take, with the identifier, what it is and what it is holding, and then y or n. The
    question stands in the prompt's place like a tool's and is pinned rather than modal, so the
    context tab is a keystroke away while it waits and p there keeps a row out of the pass. Saying
    yes works the pass out again rather than applying the list that was shown, so a pin made while
    reading it is honoured rather than refused after the fact.

    It is also the only way out of a context too big to send. context is the model's own tool and
    reaching it costs a request - the request that is failing - so until now a session that had run
    out of room could only be pruned by hand, item by item.

    --headless prints the list and takes it, there being no key to press down a pipe. The opposite
    of what --on-ask does with a tool's question, because they are different questions: a tool's is
    the model asking to do something nobody vouched for, and this one is a line the operator typed.

  • The context tab says how far over the limit the next request is, when it is over. The corner
    turns red and says the compactor runs first, and has no room for the figure that decides what to
    do next; the difference between "over" and "over by two thousand" is the difference between
    reading forty rows and taking one of them out.

  • grep answers with the files its matches were in when the lines will not fit, rather than with
    the first few thousand bytes of them. A capped lines answer is filled from wherever the walk
    started and says so, which is why it already advised files_only; this takes that advice rather
    than printing it. The saving is the shape that prompted it: four broad searches in one turn, each
    capped at a hundred matches and each still filling its byte limit with context lines, put forty
    thousand tokens into a context in one step.

  • A request the model refused for its length says what that length was and how much of it has to
    go. The sentence above it is the server's, and every vendor writes those two numbers in a
    different order; this says what they mean here. It matters because the only other figure in
    front of somebody at that point is the corner, which is the estimate that has just turned out to
    be wrong - and which the same refusal is correcting.

    A request refused here gets its own sentence, because it is a different fact. The endpoint's
    is the model's own tokenizer reporting on a request it read; this one is an estimate of a
    request nobody has seen, made by the counter whose being wrong is the reason any of this exists,
    and it says so. "The model read that request as" a number the model was never shown is the kind
    of confident wrong sentence this corner of the program exists to stop.

  • --send-oversized, and send-oversized in a settings file: send a request that looks too long
    for the model anyway, and let the endpoint be the one that says no. Both halves of that check
    can be wrong - the figure is an estimate and the limit is whatever the endpoint advertised.
    liquid/lfm-2.5-2.6b is quoted at 65,536 tokens on OpenRouter and routes to a provider whose
    own window is twice that, so a session holding itself to the smaller number refuses requests
    that would have been answered. The flag costs a round trip and buys the endpoint's own count of
    the request, which is worth more than any guess made here.

  • X-OpenRouter-Categories: cli-agent,programming-app beside the referer and title, which is what
    puts an app in the marketplace rather than only in the rankings.
    Sent only to OpenRouter; KAMCHATKA_NO_ATTRIBUTION turns all attribution off. An unrecognised
    category is dropped silently by OpenRouter, so nothing here checks the spelling.

  • y on the context tab, and /copy [N], which hand what an item says to the terminal
    for its clipboard. A screen is a rectangle and a selection over one is a rectangle too, so a
    mouse dragged across the chat pane takes the frame down both sides of every line with it, the
    wrapping of whatever width the window was, and none of what has scrolled past - and getting a
    model's answer out of here meant deleting a │ from the front and the back of forty lines. This
    is the answer as the context holds it, unwrapped and whole. /copy with nothing after it is the
    last thing the model said, which is the one people are usually reaching for, and a command as
    well as a key because on the chat tab a bare y is a y typed into a message.

    It is OSC 52, an escape sequence rather than a dependency, so it works over ssh - the terminal
    at the far end is the one holding the clipboard. What it cannot do is find out whether it
    worked: there is no reply, and a terminal that does not implement it drops it silently. So the
    line says what it did rather than that the clipboard now holds it, and the byte count is the
    receipt. App::clipboard is the seam: the app sets the text, the loop that owns a terminal
    writes the sequence, and an embedder gets the text to do its own thing with.

  • context with look takes a select, which lists the items a class comes to without moving
    any of them. Until now the selector grammar could only be resolved by using it: elide with
    select: "tool:shell" said what it had taken after taking it. The listing carries what that
    class is sending and holding against what the whole request carries, and marks the items a move
    would refuse - the person's pins, a system instruction, the turn being spoken in - off the same
    function the move consults, so the preview and the move cannot come apart. 483 bytes a request.

fixed

  • ids and select in one call are refused rather than half done. select won and ids was
    dropped without a word, so elide with both moved whatever the selector matched and reported
    exactly that - an ordinary answer, with nothing in it saying the numbers were never looked at.
    It is the failure the argument wrapper already refuses one level out, where a call puts
    arguments inside call and beside it, for the same reason. The refusal quotes both back and
    says how to spell either.

    The schema cannot say it: mutual exclusion is oneOf or not, and neither is a keyword both
    dialects one schema goes out in accept - Google's Schema is a closed set of fields with
    neither in it. anyOf does not say it either, since a branch per argument still matches a call
    carrying both. So each description says it in words and the tool enforces it.

  • Four things a tool description left a model to find out by spending a call. shell said "long
    output is cut off at the end" and now says at how many bytes, read off the limits table so that
    /limit exec:run moves the sentence too. context's search and log's take say what
    happens when they are left out, which every other argument that has a default already did.
    log's kinds said "spelled as the summary spells them", which is the vocabulary behind the
    thing you need the vocabulary to ask for, and now spells two. And select gives three of its
    forms where a call is written rather than only in the tool's description above. 198 bytes a
    request, against four ways to spend a turn learning what a sentence could have said.

  • A call asked permission for the one thing it does again, rather than for everything its tool can
    do. context reading its own items declared every one of its subjects, so look at three
    items asked to be allowed revise and elide too, and --allow context:look on its own could
    not look. The same reading fault had the policy miss the two arguments it consults: a curl was
    no longer judged against net:reach, and a path rule no longer matched the path a call named.
    Both failed towards allowing more, and neither is visible from anywhere but a live session.

  • The permissions tab says what a rule covers. --allow fs writes a rule about a whole domain, and
    its row read "nothing registered needs it" while the five fs:* rows above it each named fs; a
    server rule read the same, with every tool it covers sitting above it. A row was only filled
    where a tool declared that exact capability, and no tool declares a bare domain or a server name.

  • The permissions tab draws one row per decision, where a domain rule drew one for every operation
    under it as well. --allow log in a settings file read log allow log:read above
    log:read allow log, each row naming the other in the column beside it, and --allow setup
    put four more of them on the screen. Every capability a registered tool declares is a subject
    here, and a rule about the domain above one makes it a decided subject, so one answer was
    listed as the rule and again for every operation it reaches. The operations are in the domain
    row's own column, which is where a rule is read. An operation somebody answered about separately
    keeps its row - --allow fs --deny fs:write is two decisions - and one nobody has decided is
    still counted along the bottom.

  • What a rule covers is never wider than the rule. An operation's row named the tools that declare
    it, which with one tool to a domain is the subject's own first half read back: fs:glob allow fs is a rule about one operation reading as an answer about everything the tool does, next to a
    log allow log:read that reads the other way round. It names the operation now - a domain
    names the operations in it, an operation names itself, and a server or a path rule names the
    tools it binds.

  • mcp:call is not counted among the subjects a session will stop and ask about. Every tool from
    a server declares it and Careful::judges puts the server's own name in its place where this
    program spawned it, so the figure that says how much is still undecided included a subject
    n...

Read more

kamchatka 0.11.0

Choose a tag to compare

@github-actions github-actions released this 15 Sep 16:39

added

  • grep answers with the files that matched, when that is the question. files_only is what
    grep -l is for: every file that matched and how many matches it has, most first, instead of the
    lines. Found by watching a live model open with grep tools over the whole tree and spend its
    entire turn on what came back.

    Measured against this repository, that pattern costs 3,223 tokens of lines - and because the
    walk is alphabetical and the cap fires at a hundred matches, every one of them comes from files
    beginning with .github/. It never reaches the file the question was about. The same search with
    files_only costs 894, sees all 205 files, and puts kernel/mod.rs and command.rs near the
    top. So a capped answer now names files_only first among the four things to try, because it is
    the one that answers the situation rather than working around it.

    Ranked by how much each file matched rather than by path, which is the one thing here that does
    not answer in walk order: the question is where does this live, and the file with twelve
    matches is the answer to it far more often than the file with one. The path breaks a tie, so it
    is still the same answer twice for the same tree. Everything else is the search it already was -
    the same walk, the same skips, the same path rules, the same first line accounting for what was
    not read.

  • /note: something the model should know, without asking it to answer. Everything a person
    could say to a model went in as a message, and a message starts a turn - so telling it a fact
    it will need in four turns' time cost a request, an answer, and an "understood" nobody wanted.
    Saying it with the next question buries it; saying it afterwards is too late. /note the CI runner has no network puts it in the context and stops there.

    It is /attach with a message instead of a file, down to what it does not do: nothing is sent,
    nothing is said about the item (the chat derives that line from the item itself, so there is one
    account of it rather than two), and it is not pinned, because what is worth keeping from
    compaction is a judgement about the note rather than about notes - p is one key on the row.

    It goes in as ContextItem::memory, a reference whose source is memory, which is the
    constructor the runtime already had for exactly this and which nothing here was calling. The
    difference from a user message is not the wire - a reference projects as a user-role message
    just as an attached file does, so a note and then a question is two user messages either way.
    It is what the item is to everything that reads it: the model gets note: in front of the words
    and can tell a fact it was handed from a thing it was asked, /exclude memories names every one
    of them and nothing else, and the chat draws it as what went in rather than as a line somebody
    spoke.

  • grep and glob: finding things costs read rather than shell. There were four tools and
    none of them could look for anything, so every "where is this defined?" went through shell -
    which subsumes every other capability. A session that only wanted to be asked about a
    repository had to hand over the one permission that answers for everything, and a read-only run
    was not possible at all. These two declare read, and the path rules that bind read bind them.

    Underneath is ripgrep's own engine - grep-searcher, grep-regex, ignore, globset - linked
    in rather than shelled out to. The rg binary is those libraries plus a printer, so this is not
    a search written here; it also needs no rg on the machine, and adds no second process for the
    sandbox to account for. The printer is the half worth writing:

    3 match(es) in 2 file(s) · 205 file(s) searched
    skipped: 1 file(s) a path rule says to ask about, 2 binary file(s)
    src/ui/mod.rs:413:fn draw_search(frame: &mut Frame, app: &mut App, area: Rect) {
    

    The cut is at matches, not at bytes. A byte limit takes the tail of the last file searched
    and leaves the model believing it has seen the rest - the thing that makes an agent run the same
    search three times. A hundred matches, a line cut at two hundred characters, and a first line
    that says it stopped and what to do about it. The byte limit is still there underneath, listed by
    /limit like every other, as the backstop for one pathological line rather than as the thing
    that shapes an answer. It leads rather than trails because an output limit cuts from the end,
    which is what shell's exit line already knew.

    And the answer accounts for what it did not read. The count of files searched is what tells
    "it is not there" from "nothing was opened", and the skipped: line names each reason. The one
    that had to be built rather than linked is the path rules: Careful matches them against the
    path in the call, and a search names a directory - so .env*: ask bound read and would have
    waved a walk straight through. A rule that is not allow now stops the walk opening that file,
    since "ask me first" is not a thing nine hundred files can honour, and the nearest honest thing
    to it is not to read them and to say how many.

    Three walking rules, each of them a thing a model would otherwise conclude something false from.
    What a .gitignore hides is skipped and .git always, because it is a database rather than
    anything anybody wrote. Hidden files are searched, because a model that cannot find
    .github/workflows concludes the file does not exist, and an absence it cannot account for is
    worse than a few extra files opened. And the order is sorted rather than the parallel walker's,
    because the answer becomes a context item: two identical searches differing only in their order
    are two items nobody can diff and a budget pays for twice.

    A symbolic link is read where it points inside the working directory and counted where it points
    out, which is Reach::allows answering - the same call, with the same answer, that read makes
    about the same path. That is the second rule this had: the first was to skip every link, and a
    live run in this repository argued it down, where five crates each carry a LICENSE-MIT link to
    the file at the root and every answer to every search led with skipped: 5 symbolic link(s).
    Noise on a line whose whole job is to be rare, and a claim that something was withheld when
    nothing was.

  • A shell result says how it went in colour. Every line of a tool result is drawn quiet, which
    is right for the wall of output and wrong for the one line somebody is waiting for: the first
    line of a shell result is what the command exited with, and it read the same as the output
    under it. It is now green where the command reported success, red where it reported a failure,
    and yellow where it never got to report - stopped at the person's request, killed by a signal,
    or a status that could not be read at all. The output under it is untouched, since a line that
    stands out only works while the ones around it do not.

    The middle of those three is a unix reading. Windows has no signal to report, and the shell
    there hands a killed child back as an ordinary exit code - kill -9 arrives as 2304 - so a
    command that was killed reads red, and nothing can tell it from a command that really exited
    2304. What is yellow on every platform is the stop somebody asked for, which this program does
    itself and does not have to read off a status.

    Three and not five. The fourth thing somebody would want - a 1 from grep meaning no match
    rather than a fault - is deliberately absent, because telling those apart means knowing what the
    command was, and a guess that paints a working pipeline red or a real failure yellow is worse
    than the number on its own. Yellow is therefore not "a small failure", which nothing in a status
    line can tell: it is the command never reported.

    The colours are the three this program already uses for this question - allow/ask/deny on
    the permissions tab, and the budget bar as it fills - rather than a fourth vocabulary to learn.
    And the reading of the line lives in tools::shell beside the writing of it, as Exit, because
    a colour worked out at the drawing end from a string it does not own is a second opinion about
    what a result means; the two drift the first time the wording changes. A debug_assert in the
    tool checks that what it wrote reads back as what it meant, so every test in the crate that runs
    a command is also a test of that.

  • A resumed session reads back what its items used to say. A Snapshot carries items and not
    events, and the text a rewrite replaced is an event - context.replaced, the one event in the
    runtime that carries content, which is the whole reason it does. So -r came back with every
    v1 page empty, while the words were sitting in a file next to the one it was reading: /save
    writes the .jsonl of every event and the .json snapshot together, under one name.

    App::recall now opens the .jsonl of the same name beside whatever path -r was handed, and
    walks its context.replaced records through the same App::remember the live path uses - so a
    resumed v1 is the v1 the session had, eight deep, in the order it happened, and an item
    whose rewrite was undone is deduped by the line that dedupes it live. Only for items that came
    back: a rewrite of something an undo removed before the snapshot was written is history for
    an item this session does not have, and a resumed kernel has no undo stack to bring it back on.
    The resume line says what it picked up, beside the item and token counts it already said.

    Best effort, deliberately. The session has resumed by the time the log is read, so nothing
    found there is worth failing over: a record that is absent, unreadable or cut off mid-line
    costs a page...

Read more