Skip to content

Releases: gorillanobakaa-dot/Gorilla.Fieldkit

Fieldkit 0.1.0: tested tools that let a small AI model do real work

Choose a tag to compare

@gorillanobakaa-dot gorillanobakaa-dot released this 30 Sep 07:07

Fieldkit 0.1.0: tested tools that let a small AI model do real work

Layman track first, developer track below.

Install guide, step by step (ask your AI to do it, or do it yourself; Gorilla OpenCode and LM Studio set-up; making a small model aware of the tools): INSTALL.md.

Date: 2026-09-29


Why This Release Exists

An AI coding helper that runs on your own laptop is usually a small model. On its own, a small model guesses: it reads file after file, loses track, and sometimes makes things up. Fieldkit gives it tools that do the hard part in tested code and hand back a plain answer. The model only has to pick the right tool. On one laptop, the same small model (Gemma) went from one of five tasks right to five of five, and finished in about half the time.

What You Will Notice

Finding where something is in a project

  • Before: The model reads files one by one and can report a near match as the answer.
  • After: The model asks Fieldkit and gets DEFINED at file:line, or NOT DEFINED, so it has nothing to guess.
  • Affects: everyone

Sending a Word, Excel, PowerPoint or PDF file

  • Before: People's names hide in the file's properties, and Word writes them back on every save.
  • After: fieldkit office deliver FILE checks the file, removes the names, checks again and scans for private data. It answers SAFE TO SEND or NOT SAFE with the reason.
  • Affects: everyone

Changes the AI makes to your files

  • Before: The model changes files and reports its own version of what happened.
  • After: Fieldkit shows a preview first, keeps a backup, checks the result and can undo it. Changes to the system, and anything that cannot be undone, need your approval, and the model cannot approve on your behalf.
  • Affects: everyone

Publishing a release

  • Before: Release notes can claim things nobody tested.
  • After: fieldkit release check refuses a release until every claim in its notes has a passing proof. It has no override switch.
  • Affects: some users

Deliberately Not Done

  • A cloud service or an account — Everything runs on your own computer. Nothing is uploaded.
  • A --force switch on the release gate — A tired person at 2 a.m. would use it. A failing gate means you fix the release or fix the check.
  • Separate Windows and Linux versions — One Python codebase runs on both, so the two cannot drift apart.

Privacy & Security

Fieldkit has no telemetry and no account, and it reports nothing to anyone. It uses the network only when you run a command that needs it: fieldkit gather downloads the GitHub repositories you list, fieldkit release check reads a release from GitHub, the Debian kernel pipeline downloads the kernel from kernel.org, inventory.py github lists your GitHub repositories, and fieldkit exam talks to the model server on your own computer (LM Studio by default). Your private words, such as your name and email, stay in fieldkit.local.json, which Git ignores. The privacy scan never prints a secret in full.

How to Install

Before you start:

  • Python 3.11 or newer
  • Git
  • About 50 MB of disk space

Step 1: Download Fieldkit.

git clone https://github.com/gorillanobakaa-dot/Gorilla.Fieldkit

✓ You have a new folder called Gorilla.Fieldkit.

Step 2: Go into the folder and install it with its test tools.

cd Gorilla.Fieldkit && python -m pip install -e ".[test]"

✓ The last line says Successfully installed fieldkit-0.1.0 together with the other packages.

Step 3: Run the tests on your own machine.

python -m pytest -q

✓ The last line shows passed tests and zero failed. Some tests say skipped, each with the reason, because they need tools you have not gathered yet.

Step 4: Check what Fieldkit sees on your computer.

fieldkit host

✓ It prints your system, for example Windows 11 or Debian, and your Python version.

To go back: Run python -m pip uninstall fieldkit, then delete the Gorilla.Fieldkit folder. Fieldkit changes nothing else on your computer.

If Something Goes Wrong

fieldkit is not recognised as a command.
The folder where Python puts commands is not on your PATH.
What to do: Use python -m fieldkit instead of fieldkit. It does the same thing.
Status: deferred

Some tests say skipped: pfind not gathered.
Those tests need tools that Fieldkit copies in from GitHub.
What to do: Run fieldkit gather, then run the tests again.
Status: fixed in 0.1.0 (the skip explains itself)

The Debian kernel pipeline stops at every stage on Windows.
Those stages only run on Debian, and Fieldkit reports that as not-this-platform instead of pretending.
What to do: Run the kernel pipeline on Debian. Its full build has not run yet, so report anything that fails.
Status: investigating

Common Questions

Q: Does this make my small model as good as a big one?
A: No. It makes the small model good at jobs that a tool covers, such as finding code, checking files and naming a build failure. Designing something new or reasoning about an unfamiliar bug is still up to the model.

Q: How do I use it with Gorilla OpenCode or LM Studio?
A: Add Fieldkit as an MCP server that runs fieldkit mcp. The AI then sees six tools: discover, describe, run, undo, next and readiness. The exact settings are in the README.

Q: Can the AI break my computer with it?
A: Changes to files show a preview first and keep a backup that undo restores. System changes and anything that cannot be undone wait for your approval. Fieldkit reads command text and files; it is not a sandbox, so keep your normal care.

Q: Were the results measured or guessed?
A: Measured, on one laptop, with five tasks. That is a small test, and the same author wrote the tools and the tasks. You can run the same test yourself with fieldkit exam run --model <your model>.

Bottom Line

Fieldkit 0.1.0 is a toolbox that does the hard, repeatable parts of a coding job in tested code, so a small model on your own laptop can do useful work. In one measured test, the same model went from one task in five to five in five, and used fewer tokens doing it. It is honest about its limits: the test was small, the Debian kernel build has not run for real yet, and it is not a sandbox. It costs nothing, needs no account and uploads nothing. If you run a small model on ordinary hardware, try it and run the exam on your own model.

Claim Sources

Claim Basis Evidence
Gemma went from one of five to five of five 📄 stated in input Gemma with Fieldkit tools: 5 of 5 right, 410 s, 6,630 tokens.
finished in about half the time 🤖 model inference (none — model judgment)
the model only has to pick the right tool 🤖 model inference (none — model judgment)
nothing is uploaded 📄 stated in input Everything runs locally; nothing is uploaded; no account; no cloud.
network use is limited to gather, release check, the kernel download, inventory github and the local exam server 🤖 model inference (none — model judgment)
it is not a sandbox 📄 stated in input it is not a sandbox
the test was small 📄 stated in input five tasks, one fixture, written by the same author as the tools

How to verify this document:
📄 stated in input — the model's phrasing of something your source text said.
Find the matching line in the original to verify.
🤖 model inference — the model's own judgment or synthesis. Treat as opinion,
not measurement. Re-run on the same input and check whether specific numbers
stay consistent between runs.

Auto-generated plain-language release notes.


Developer track: Fieldkit 0.1.0: deterministic tools, an MCP agent interface and a release gate, one Python codebase for Windows 11 and Debian

Version: 0.1.0
Date: 2026-09-29


Summary

Fieldkit moves the repeatable parts of agent work out of the model and into tested Python. Every command returns JSON and uses fixed exit codes (0 fine, 1 error, 2 bad usage, 3 findings). An agent interface (discover, describe, run, undo) enforces one sequence for anything that changes state: it refuses draft cards, validates inputs, runs a preview, requires approval where the card demands it, backs up the scope, applies, verifies and restores the backup if verification fails. The same interface is served over MCP (JSON-RPC 2.0 on stdio, protocol 2024-11-05) by fieldkit mcp, and the MCP surface has no approval argument, so a model cannot approve its own irreversible or system changes. This first public release also adds office deliver, lifecycle and cross-machine release evidence (release prove).

Known Alternatives Considered

Separate Windows and Linux script sets were considered and rejected in favour of one cross-platform Python codebase, so that the two cannot drift apart. The release gate deliberately has no --force flag: the design comment reads "a failing gate means fix the release or fix the check". Claims about another machine use a recorded git tree id instead of file hashes, because line-ending conversion changes file bytes between Windows and Linux while the tree id stays the same.

Architecture Impact

First public release. The registry is split: fieldkit/desk/tools.yaml and imports.yaml are public; local/tools.yaml, local/imports.yaml and local/tests/ are git-ignored and merged at load time, so an owner's private tools never enter the repository. Pipelines, triage signatures and tool cards are data (YAML); the engine code reads them.

Toolchain

Python 3.11 or newer. Runtime dependencies: `pyyaml>=6`, `lxml>=5`, `python-docx>=1.1`, `openpyxl>=3.1`, `python-pptx>=1.0`, `reportlab>=4`, `pypdfium2>=4`. Optional extras: `[test]` (`pytest>=8`), `[readers]` (`markitdown`, `docling`). Install with `python -m pip install -e ".[test]"...
Read more