Skip to content

Releases: nicotem/exegete

v0.14.1-alpha: qualcoder-mcp is now Exegete

Choose a tag to compare

@nicotem nicotem released this 01 Oct 01:36

Pre-release. Full details: CHANGELOG, INSTALL, PRIVACY, NOTICE.

Exegete, until now called qualcoder-mcp, is a qualitative analysis application in its own right, used in conversation with an AI assistant and compatible with QualCoder. It works as it did, and every earlier spelling is still accepted until v1.0, so nothing you set up stops working. Two things to do: with Claude Desktop, open the extension file from this release; and for participants' data, use Claude Desktop's chat with the extension, on the conditions under "Participants' data". Also new: a README front page, an Experimental route for OpenAI's apps, and a plain statement of which assistants can open your files by themselves, outside Exegete. This remains experimental research software: back up your projects (Exegete also does, before every write unless a call asks it not to).

Install or update

Claude Desktop, in one click: new or updating

  1. Fully quit any other program that uses Exegete (a Claude Code session, LM Studio's server).
  2. Download exegete-0.14.1-alpha.mcpb from this release's Assets and open it (double-click it, or drag it onto the Claude window).
  3. New: press Install, then Install again when Claude says it needs to fetch a few things, and follow the README's "Start here".
    Updating: this updates your qualcoder-mcp extension, with its two settings, rather than adding a second one; Claude Desktop lists it as Exegete. The first start may take longer and needs the internet, as Claude may fetch Exegete's libraries again, and Claude may ask again before using each tool. Your projects folder (~/QualCoder projects unless you changed it; ~ is your home folder) does not change.

The extension is not signed. On a personal Claude plan it installs like any other; a computer or Claude account your university or employer manages may block unsigned extensions, or extensions altogether, and INSTALL.md says what you then see. Neither updating nor removing it touches your projects.

SHA-256 of exegete-0.14.1-alpha.mcpb: cc805cad78e4bf6d9c7b405a9075504dcfc21aeda9d367aa217aca7ec0093e8a

PyPI or a copy of the source

Quit first every AI app that starts Exegete (its host): Claude Desktop: Cmd+Q; Claude Code: end the session; LM Studio: toggle the server off. A server left running while its files change fails when it first needs a part it has not loaded, and the first start after the update moves Exegete's folder, best done with no older version running.

Keeping the old name: pip install --upgrade qualcoder-mcp, pipx upgrade qualcoder-mcp, uv tool upgrade qualcoder-mcp and uvx qualcoder-mcp bring the current Exegete (the old name's package is released beside every release until v1.0; its last release will say so), and the qualcoder-mcp command and python -m qualcoder_mcp.server start it, saying so in one line of the log. Keep a host entry named qualcoder, and do not add an exegete entry beside it: that would start two servers and show every tool twice.

Moving to the new name: install it, change the command in your host's configuration to exegete, and only then remove the old package: pip install exegete, then pip uninstall qualcoder-mcp; or pipx install exegete or uv tool install exegete, then pipx uninstall qualcoder-mcp or uv tool uninstall qualcoder-mcp (here the other order leaves your host with no server, as the Exegete the old package brought goes with it). With uvx: uvx exegete.

A copy of the source (git): quit your host, then git pull and pip install -e . as always. python -m qualcoder_mcp.server and the qualcoder-mcp command keep working through a small stand-in; the new forms are -m exegete.server and exegete. Pointing the copy at the new address is optional: git remote set-url origin https://github.com/nicotem/exegete.git.

What stops working: your own code that imported Exegete's inner modules under the old name (qualcoder_mcp.database and the like).

Then restart your host fully, so it reloads the tool descriptions, and run the check. INSTALL.md's "Coming from qualcoder-mcp" has every route.

Afterwards: the check

  • Extension: it can leave the link at ~/.qualcoder_mcp and one old log file. The extension has no exegete command, so type uvx exegete@latest --check-transition in a terminal (it needs uv; the @latest makes uv fetch this release rather than reuse a copy it may have kept from before). If uvx is not found, those two are harmless and can stay.
  • PyPI or source: type exegete --check-transition in a terminal.

It lists what the move left (the link, the old package, a host's entry still starting the old command, an extension older than Exegete, Claude Desktop's log under the old name, the earlier projects folder) as numbered steps in the order to take them, with full paths ready to paste. It changes nothing, and ends with exit code 0 when nothing is left. If the old package came through uv tool or pipx and there is no exegete command yet, its first step installs Exegete; for a copy of the source, it names the folder and says to quit your host before updating it. Then --tidy removes the link (never a folder), only when nothing it can find could still use it or start an older version; with it, --tidy-old-logs removes the old logs. Projects, backups, the name files in projects and your hosts' configuration files are never touched.

Participants' data: assistants that open files by themselves

Exegete's protections (the ##### mark that keeps a memo's private part out of the conversation, your approval before anything is written, the backups) cover what passes through its tools. Some assistants also open files on your computer with tools of their own, and what they read that way goes to their provider without passing through Exegete.

For participants' data, the documents now suggest Claude Desktop's chat with the extension, on the conditions below; where your institution needs commercial terms, on a Team or Enterprise account.

  • Claude Desktop's chat with the extension, as far as Anthropic's pages say, opens no file by itself when computer use is off, no other extension that reads files is installed, and no folder that holds your projects or transcripts is connected to it (your home folder, Documents or a whole drive included).
  • Codex, in its "Ask for approval" and read-only modes alike, reads files well beyond the folder it works in, without asking: on a Mac or Linux, any file your account can read; on Windows, at least everything in your home folder but a few folders that hold keys. Exegete's answers tell it where your project is. A folder of its own keeps a study's files out of the folder where Codex works and makes changes without asking, but not out of its reach.
  • Claude Code, under any plan or terms, reads the folder it starts in without asking; its read-only commands, such as cat, read outside that folder without asking too; and in its usual starting mode its own file tools can also read outside it after one question the first time, as PRIVACY.md says. INSTALL.md's steps now make an empty folder for Claude Code, register Exegete there and start it there, never in your home folder or a folder that holds a study.
  • Cowork reads in the folders you connect to it.

README, INSTALL.md and PRIVACY.md give this advice; PRIVACY.md's "Assistants that open files by themselves" covers Codex, Claude Code, Cowork, Claude Desktop's chat and LM Studio, with each maker's page and the date it was read. See also "Practising on the same computer" under Known limits.

For earlier users: what the new name changes

  • Names. The program, its PyPI package, command and Python module are all exegete (pip install exegete; exegete --version answers exegete <version>); it calls itself Exegete in the conversation and its log. The GitHub address is https://github.com/nicotem/exegete; the old one redirects.
  • Exegete's own folder moves by itself. At the first start, ~/.qualcoder_mcp becomes ~/.exegete, whole: the secret key, the coding sessions, the last-project hint and the privacy run records. A link left under the old name lets an older version on the same computer keep using the same folder and key. If a backup or sync rule of yours names the old folder, change it.
  • Settings (environment variables) now start EXEGETE_ (for example EXEGETE_TOOLSET, EXEGETE_WORKSPACE). The earlier QUALCODER_MCP_... spellings and QUALCODER_PROJECT_PATH are still read until v1.0, and the log says so; if both spellings of one setting are set to different values, Exegete does not start and names the two.
  • The projects folder, for PyPI and source installs. With no workspace set, it is now ~/Documents/Exegete projects: new project copies and new projects go there. ~/Documents/Qualcoder MCP Projects is never moved or emptied, and its projects are still found; set EXEGETE_WORKSPACE to it to keep using it. The extension's folder is unchanged.
  • The AI coder name (the name the assistant's coding is recorded under) moves to the project's exegete.json the first time the name is stored. The earlier qualcoder_mcp.json stays beside it, flagged so that qualcoder-mcp 0.12 to 0.14 refuse to write under a name you have since changed; if one of those says the file "was written by a newer version", update that qualcoder-mcp.

What Exegete covers today, and what still needs QualCoder

Exegete covers creating a project (Experimental), cases and attributes, bringing in text, coding with suggestions you decide one by one, the codebook, memos, annotations and a journal, searching, reports and exports, comparing coders, replacing names, and backups ...

Read more

v0.14.0-alpha

v0.14.0-alpha Pre-release
Pre-release

Choose a tag to compare

@nicotem nicotem released this 29 Sep 14:12

v0.14.0-alpha: one-click install for Claude Desktop, creating a project, and texts that say what the server does

Pre-release. Full details: CHANGELOG, INSTALL, PRIVACY, NOTICE.

v0.14 makes the server easier to start and more honest to use. It was built in six parts: a workbench for the tests, privacy, creating a project, existing projects, what the server tells the assistant and the researcher (in three parts, after a claims audit of every tool and text), and a one-click extension for Claude Desktop. Each part was reviewed before it merged, by a QA gate and a Security gate and re-verification until a check found nothing major (the workbench, which changed nothing under src/, by one check). One tool is added, create_project, in a new opt-in tool set: 73 tools in full, 21 in core, 74 in lifecycle.

Install in one click (Claude Desktop)

For Claude Desktop on macOS or Windows there is now nothing to type.

  1. Download qualcoder-mcp-0.14.0-alpha.mcpb from this release's Assets.
  2. Double-click it (or drag it onto the Claude window), click Install, and Install again when Claude says it needs to fetch a few dependencies. Claude Desktop fetches uv if your computer has none, and uv fetches Python 3.13 and the server's libraries exactly as locked; nothing is compiled on your computer.
  3. Look at the two settings (Settings, Extensions, qualcoder-mcp): Tool set, lifecycle by default (every tool, creating a project included), and Folder for projects, ~/QualCoder projects by default, outside Documents, which iCloud or OneDrive may sync.

The extension is not signed. On a personal Claude plan it installs like any other extension; an organisation that allows only signed extensions blocks it, and INSTALL.md says what you then see. To update, install a newer .mcpb the same way; neither updating nor removing it touches your projects.

SHA-256 of qualcoder-mcp-0.14.0-alpha.mcpb: ae2d16bf754ea1bd31ff2981f78e775626d458b699d438f0899f7cc3d1e96c19

It was built from the release commit on Linux, Windows and macOS by CI, which found the three builds identical, validated the manifest with the official MCPB tool, and installed and started it as Claude Desktop does.

Other hosts (Claude Code, LM Studio, Claude Desktop configured by hand) keep the Terminal route: pip install --upgrade qualcoder-mcp, then restart the host fully so it reloads the tool descriptions. The dependency floor is unchanged (mcp>=1.17.0,<2).

What is new

  • Creating a project from the conversation (Experimental, opt-in). create_project makes a new, empty project in QualCoder 4.0's format, as 4.0's own New Project makes it, in one transaction, so a failure leaves nothing committed; then selects it. It asks for your own QualCoder coder name last, or accepts "not known" with a warning. It lives only in the lifecycle tool set, which the extension uses by default; the Terminal route keeps full unless you set QUALCODER_MCP_TOOLSET=lifecycle. QualCoder 4.0 opens such a project without a message when your QualCoder coder name is the one you gave, or when you gave "not known" (it then records yours); if your QualCoder is set to another name, it asks whether to keep it or switch. 3.8.2 opens it but shows sub-codes as ordinary codes.
  • A folder for projects. QUALCODER_MCP_WORKSPACE names where create_project and copy_project_to_workspace work; the extension's "Folder for projects" sets it. Unset, it stays ~/Documents/Qualcoder MCP Projects on the Terminal route.
  • The project memo can be written: set_memo takes project as a target. Only the public part is replaced; the private part after ##### survives. A coding session now starts with the memo's public part, as QualCoder 4.0 hands the memo to its own assistant.
  • Every tool says what kind of tool it is, in MCP's four marks (reads only; can replace or remove work; a repeat changes nothing; reaches nothing beyond this computer). What each host does with them differs by host and mode; INSTALL.md says what Anthropic documents, mode by mode. For work on real data, keep your host asking.
  • A reading in place of the confidence score. Each coding suggestion is now marked explicit (the passage states what the code names) or interpretive (the code rests on what the passage implies rather than on what it says; the reason names the words it rests on, and it may draw on what the same participant says elsewhere or on your study's framework as the project memo states it, never on outside facts or assumptions). The 0 to 1 score was a model's rating of itself, not a measurement, and it went into the research record as if it were one. Nothing sorts or totals by the reading, and you can change it at review. An applied coding's memo says it in words.
  • A coding session starts from your answers. Before starting one, the assistant is told to ask three things: what to look for, as a lens (your own codes, topics, people's own words, actions, feelings or values, or other) and whether to point out passages no code fits; how long a coded passage should be; and whether a passage may carry more than one code. The answers are to be the session's instruction, now required. The server refuses a session without an instruction but cannot tell whether it holds your answers: check the instruction the session records (get_coding_session_info shows it).
  • The review shows the file's own text. Session files keep no text around a passage any more. The review reads it from the file each time and shows, in a transcript, the nearest earlier turn by another speaker (found by speaker labels, and saying so), then the paragraph or turn with the coded words marked, then the code, the reading and the reason.
  • A decided suggestion can be reopened (update_suggestion_status's reopen), and delete_coding marks the suggestion a session applied as removed, so it can be applied again.

What changed

Privacy.

  • The pseudonymisation run no longer returns a plain fingerprint of each file's text before the run, which, beside the rewritten text, could confirm a guessed name. The run record is now format 3, with keyed fingerprints. Records written by v0.12 and v0.13 still hold the plain ones: keep them private, or delete the ones you do not need.
  • Error answers and the server's log lines carry the kind of error, never SQLite's message, which a crafted or damaged project could make quote a note, private part included. The log names no project, file, code, case or path (the MCP library's own lines can quote a malformed request; see PRIVACY.md). What a host records in the same log file is the host's (Claude Desktop's server log records every request and answer), and PRIVACY.md and INSTALL.md say so.
  • A coder hidden in QualCoder while a conversation is under way is filtered from the next read, without selecting the project again.

Existing projects.

  • Backups are taken with SQLite's own online backup, so a backup made while QualCoder was writing no longer carries a half-written journal; a backup that does is marked unclean and refused by restore_backup. During a backup QualCoder's own saves wait (about 1.2 seconds per gigabyte of database).
  • delete_category and merge_category remove the category's own node from QualCoder 4.0's saved graphs, so the Graph window no longer reports a missing category on every load.
  • PDFs with no text layer, and PDFs QualCoder 3.8.2 stored as the file itself, are named in reads, never returned, and refused by the coding tools, with the way forward (OCR outside this server). Region codings on PDFs and images and audio or video codings are counted where a read leaves them out.
  • Backups are dated and pruned by the time in their names, not by their folders' dates, so a restore's safety backup is no longer offered for pruning as old.

What the tools say.

  • An argument a tool does not declare is refused by name, and nothing runs. A misspelt safety flag used to fall back to its default in silence: with a project pseudonyms file, one letter short in apply_project_pseudonym imported the real names.
  • Text holding #####, QualCoder's private-note marker, is refused where it used to be cut, which emptied or deleted notes.
  • Reads refuse a wrong id, name or coder instead of answering as if nothing were there; case-insensitive search works beyond A to Z; search_memos searches every kind of note; attribute queries compare numbers only and say what they left out; the co-occurrence window is a distance, as QualCoder measures it; three exports say what they hold; query_by_attribute answers an object.
  • "Saturation" and "prominent themes" are gone from every text the server sends: get_coding_frequencies counts codings, not participants or importance.
  • pseudonymise_source's texts now say that with rewrite_memos on, a run rewrites a name in notes across the whole project, so two people who share a name get one pseudonym in their notes, whatever the order of the runs. The safe route is in the tool's description and PRIVACY.md.

Deprecated, removed in v0.15. Each still works and says so in its description and its answer: read_pseudonym_list (QualCoder's Pseudonyms dialog shows the same list without sending it to the AI provider); export_refi_qda, both forms (use QualCoder's own REFI-QDA export, which keeps each coder, the cases, the notes and the media); export_code_report (use get_coded_segments); cleanup_old_sessions (use delete_coding_session); merge_proposals; four help topics that repeat the tools' descriptions; and, when used, delete_code's cascade, owner on apply_codings and import_text_file, attributes on journal entries, search_files' `...

Read more

v0.13.0-alpha

v0.13.0-alpha Pre-release
Pre-release

Choose a tag to compare

@nicotem nicotem released this 25 Sep 16:32

v0.13.0-alpha: pseudonymisation one file at a time, notes rewritten, cases and files renamed, and the LGPL

Pre-release. Install or upgrade with pip install --upgrade qualcoder-mcp, then restart your MCP host fully so it reloads the tool descriptions. The dependency floor is unchanged (mcp>=1.17.0,<2). Full details: CHANGELOG, PRIVACY, NOTICE.

v0.13 follows up the pseudonymisation tool that 0.12 introduced. It was built in four parts: a housekeeping batch, the count of names left in the file text, the two rename tools, and the rewriting of notes with the mapping kept. Each part went through a QA gate, a Security gate and re-verification until clean. Three tools are added, rename_case, rename_file and read_pseudonym_list: 73 in the full toolset, 21 in core. And the licence changes: from this release qualcoder-mcp is licensed under the GNU Lesser General Public License, version 3 or any later version (LGPL-3.0-or-later), which is QualCoder's own licence.

Pseudonymisation: one file per call, and every count read two ways

  • One file per call. pseudonymise_source now takes one file_id (it took a list, file_ids). A mapping that is right for one participant is applied to that participant's file, so two people who share a name can be given two pseudonyms in the file text. With rewrite_memos on, a run rewrites that name in notes across the whole project, so when two people share a name keep it off and change their notes by hand (a fix is planned). To pseudonymise a project, run it file by file. A file the tool cannot rewrite (a PDF, a media file, a source with no stored text) is refused with the reason.
  • The names left in the file text are counted. The preview's residue report now also counts, as occurrences, the names left in the text of every file after the run: the file this call rewrites, every file it does not touch, and the PDF sources, which are never rewritten. Every file in which a name still shows is named. A name inside a longer word (Thomas_P01, Thomasson) is reported and never substituted; on a mapping you type, those longer words are listed so you can add an exact entry for one.
  • Every count is two readings. Each residue count is now {"wide": N, "whole_word": M}. The whole-word reading is what this run's own rule matches. The wide reading is a deliberately generous heuristic: anything a person reading the text would see, including inside a longer word and in any letter case. A wide count far ahead of its whole-word count is the sign of a short name.
  • A compact report by default. The preview gives full detail for the file you are rewriting and one short row (id, name, the two counts) for each other file that still shows a name, up to 1,000 files; residue_detail="project" gives full detail for up to 200. Nothing is left out silently, and a file the report could not afford to read is listed as not checked, never reported clean.
  • A pseudonym that contains a real name is caught. {"Smith": "Jones", "Thomas Smith": "Alex Smith"} would put the real surname back. The preview now warns about such a pseudonym, and it is withheld from the run record and the journal entry, which promise never to carry an original name.

Notes rewritten, and the mapping kept

  • rewrite_memos, off by default, also rewrites the public part of every note (twelve kinds, from codings and annotations to codes, files, cases and journal entries), across the whole project, by the same whole-word rule. A note's private part, from its ##### marker, is carried across unread: a name there is still there, and nothing in this server can report it.
  • A mapping you type must be kept somewhere. A typed mapping is half of the key that reverses a pseudonymisation, so the run asks where it is kept: saved into the project's own pseudonyms.json, in QualCoder's own format (save_mapping_to_project), or kept by you (researcher_keeps_mapping). Without either, the execute is refused and nothing is written.
  • The name list is a tool of its own. get_current_project says whether pseudonyms.json is present and how many entries it holds, never a name. read_pseudonym_list returns the names themselves; its description says first that this sends the real names to the AI provider, so a host that asks approval tool by tool asks for it apart from everyday reads.
  • The mapping's last copy is warned about. A restore_backup that makes the project's pseudonyms.json appear, disappear or change says so and names the safety backup that holds the one the project had. A prune_backups preview names the backups whose removal would take the only lasting copy this server knows of, counting QualCoder's own _BKUP_ backups as no lasting copy, because QualCoder deletes old ones when a project closes. If that set changes between the preview and the execute, the prune is refused and removes nothing.
  • The run record is format 2, an audit record of which rows and notes a run changed. It is not a way back: the backup taken before the run is.

Renaming cases and files

rename_case and rename_file rename a case or a file's entry the way QualCoder's Manage Cases and Manage Files ("Rename database entry") do, so a label named after a participant (Thomas_P01, Thomas_interview.txt) can be changed without leaving the conversation. They change the name and nothing else, take a backup by default, and say where the old name stays: QualCoder's saved graph labels, table displays and filters, other files whose names hold it, an imported file's stored copy in the project folder (which keeps the old name and, for a document, the original text), and every backup. rename_file refuses names QualCoder or Windows would trip over (path characters, device names, names over 200 bytes, a clash with a file in the project's documents/ folder) and an ending change QualCoder acts on. import_text_file now follows the same name rules.

Also changed

  • The inert confirm argument, announced for removal in 0.12, is gone from the six token-gated tools. A call that still passes it gets the preview, as in 0.12.
  • Every failure after a backup was taken now names that backup, and says whether the write was rolled back (nothing to restore) or may be part-written.
  • A preview token that does not verify is said to be one, rather than described as a changed project.

The licence

From 0.13.0, qualcoder-mcp is licensed under the GNU Lesser General Public License, version 3 or (at your option) any later version (LGPL-3.0-or-later), QualCoder's own licence. The licence texts ship as COPYING.LESSER and COPYING; the MIT LICENSE file is gone. qualcoder-mcp is a separate program that reads and writes QualCoder project files, but it contains a small number of routines and values taken from QualCoder so that its results match QualCoder's exactly, and NOTICE lists every one, with the QualCoder file and lines it comes from and why, together with the facts of QualCoder's file format the code restates and the tests that carry QualCoder's code.

Nothing changes for anyone who installs and runs the server. The licence's conditions apply to someone who distributes it, and in practice they matter for a modified version: whoever distributes one must make its source available under the same licence.

Every release up to and including 0.12.1 was published under the MIT License, and this project's own code in those releases remains available under those terms. Those releases also contained some of the QualCoder-derived items NOTICE lists; those items were always under QualCoder's licence, LGPL-3.0-or-later, whatever those releases declared. Nothing is withdrawn: the earlier releases stay on PyPI and GitHub as published.

What pseudonymise_source still does not do

Said plainly, because it decides what you may send afterwards:

  • PDFs are never rewritten. Their stored text is counted with every other file's, so the residue report says where a name remains in one.
  • Media files and QualCoder 4.0's ai_data/ folder are out of scope: neither rewritten nor scanned. QualCoder's AI search index keeps the previous text until QualCoder reopens the project and re-indexes, and its chat history may quote it.
  • The labels are counted, never rewritten by this tool: case, file, code, category, attribute-type and journal entry names, and attribute values. A case or file label can now be renamed with rename_case or rename_file; the others are for you to change.
  • The private part of a note is never read: a name after a note's ##### marker stays there, and nothing in this server can report it.
  • The backup it takes holds the real names, and so does every earlier backup beside the project, and pseudonyms.json is the reverse key in plain text. A project is not pseudonymised while its backups and that file sit beside it.
  • Three more places keep the real names. An imported document's stored copy in the project's documents/ folder keeps the original text; QualCoder's exports include that copy, and the residue report never reads it. This server's session files keep the excerpts of coding suggestions made before the run, and the run never deletes a session. speakers.json and speaker_regex.json can hold names too; the preview says whether they are present and never reads them.
  • Do not share the run record. The run record, and the tool's own result in the conversation, carry the length and a plain SHA-256 fingerprint of each file's text as it was before the run. With the pseudonymised text, these can confirm a guessed name. PRIVACY.md says more; a keyed fingerprint is planned for v0.14.
  • **Pseudonymised data is still personal data.*...
Read more

v0.12.1-alpha

v0.12.1-alpha Pre-release
Pre-release

Choose a tag to compare

@nicotem nicotem released this 21 Sep 21:31

v0.12.1-alpha: the in-QualCoder acceptance checks, run and recorded

Pre-release, documentation only. No code, no tool behaviour and no dependency changed from 0.12.0-alpha; only the version, the citation file and the version canary move. Upgrade with pip install --upgrade qualcoder-mcp if you want the corrected documentation, or stay on 0.12.0-alpha, which behaves identically. Full details: [CHANGELOG](https://github.com/nicotem/qualcoder_mcp/blob/main/CHANGELOG.md).

v0.12.0-alpha shipped with the six in-QualCoder acceptance checks for pseudonymise_source recorded as not run. They have now been run, and this release records what they found.

What was checked, and how

The six checks are the ones only QualCoder itself can make: that the highlights land on the pseudonyms, that the reports print the refreshed quotes, that the case links still cover what they should, that edit mode and its undo behave after a rewrite, and that QualCoder 4.0's AI search index re-indexes the changed text.

They were run headlessly, not by driving a window: QualCoder's own widget code under QT_QPA_PLATFORM=offscreen, with the result read back from the objects QualCoder itself populates (the character ranges and formats of the text document, the report dialog's results, the case file manager's underline ranges, the coding, annotation and case rows after edit mode, and ai_data/search.sqlite). That reads exact character offsets rather than pixels, which is stronger evidence than a screenshot of a highlight.

Result: all six on master 9bddf17, five of six on 3.8.2, and no failure anywhere is attributable to this server. On master the tool passed all six. On 3.8.2 check 6 does not apply, and two of the remaining five record a failure, both of them the upstream defect described below.

  • Highlights sit exactly on the pseudonyms and on nothing else: 39 assertions per build, including that nothing is highlighted inside Thomasin and that a hidden coder's rows appear only once that coder is made visible.
  • The coded-text report prints the refreshed quote for all eighteen codings, including the hidden coder's, because QualCoder's report reads the base table rather than the visibility view. A search for the real name returns nothing; a search for the pseudonym returns seven segments.
  • The annotation report shows the pseudonym at the planted positions. The whole-file case link covers the whole rewritten text and the partial link ends exactly after the pseudonym.
  • On master, every row moves by the inserted length when edit mode is left, and the undo restores the pre-edit positions.
  • On master with AI enabled, the search index re-indexes on reopen: the stored text hashes change as predicted, the real name falls to zero full-text hits, and Thomasin survives.

What the control measures, stated honestly

Every check was also run against the untouched original project. Roughly half of each step's assertions are bound to the rewritten project and fail there, which is what makes them discriminating; the rest are project invariants that hold either way. The per-step counts are in the acceptance report rather than summarised as one number, because a score of 39 out of 39 should not be read as meaning every assertion discriminates.

One sub-step could not be run headlessly and is recorded as such: pressing undo inside edit mode cannot be exercised, because QualCoder repopulates the editor's undo stack with formatting commands on every undo, in both builds.

Known limitation, upstream and not in this server

QualCoder 3.8.2 deletes whichever coding an edit leaves touching the new end of a file, whenever edit mode is left after any change to the text, and the undo cannot restore it.

ed_update_codings deletes any row whose new end is >= len(text), evaluated against the text after the edit. That is wider than "a coding at the end of the file", and it makes an everyday operation dangerous: trimming the tail of a transcript destroys whichever coding is left nearest the cut, even one that was nowhere near the end before. Measured in the acceptance run: deleting the last 193 characters of a 614-character file destroyed a coding at 400 to 421, which had sat 193 characters clear of the end. Master kept it.

It affects code_text only, in any file; annotations and case links are unaffected, and leaving edit mode without changing anything is harmless. This is upstream 3.8.2 behaviour, present whether or not a project has ever been pseudonymised: the same loss occurs on a project this server never touched, which is how the acceptance control attributes it. The QualCoder 4.0 line has fixed it, clamping such a coding instead of deleting it, with the comment that a coding ending at the file length is valid.

If you use edit mode on QualCoder 3.8.2, this is worth knowing regardless of this server.

Quality gates

The acceptance harness was itself reviewed adversarially, which found and fixed two defects in it before these results were accepted: one assertion whose condition was a literal and so could never fail, and a control claim that overstated what it measured. The release wording was then reviewed again, which corrected the two statements above: the earlier draft understated the 3.8.2 defect and overstated what was run. Both QualCoder source trees were verified byte-identical to genuine upstream at tag 3.8.2 and commit 9bddf17 before any attribution was made.

Suite at the release commit: 3035 passed, 3 skipped, 0 failed, on Python 3.13.5 and 3.11.13; six-platform CI green. Support and bug reports: GitHub Issues only. This remains experimental research software; back up your projects (the server also does, before every write).

v0.12.0-alpha

v0.12.0-alpha Pre-release
Pre-release

Choose a tag to compare

@nicotem nicotem released this 16 Sep 20:42

v0.12.0-alpha: pseudonymisation that keeps the coding, and the AI coder name is yours

Pre-release. Install or upgrade with pip install --upgrade qualcoder-mcp, then restart your MCP host fully so it reloads the tool descriptions. This release raises the dependency floor to mcp>=1.17.0,<2; pip upgrades mcp with the package (see "Upgrading from 0.11.x"). Full details: CHANGELOG, INSTALL, PRIVACY.

v0.12 comes out of a ground-truth study of QualCoder 4.0 (master at pinned commit 9bddf17, and the 3.8.2 tag where schema v14 parity matters): seven design dossiers, owner rulings on every question only the owner could answer, then two batches, a flagship and one follow-up, each through a QA gate, a Security gate and re-verification until clean. Three tools are added, pseudonymise_source, set_project_ai_coder_name and compare_coders: 70 in the full toolset, 21 in core. Everything schema-dependent is still detected by probing the project database, never by version string.

The flagship: pseudonymise a project that is already coded

pseudonymise_source replaces names with pseudonyms in the stored text of the text sources you choose and moves every coding, annotation and case link with the text, all files in one transaction. QualCoder pseudonymises only at import, from pseudonyms.json, and has never had a way to pseudonymise a source that is already coded; this tool is that way.

What it does, in plain words:

  • You give it the mapping: real names to pseudonyms, with variants (nicknames, inflections) folded onto one pseudonym. Nothing is detected or guessed. The rewrite replaces whole words only, using QualCoder's own boundary rule, so "Ann" is not touched inside "Anna" while "Tom" is replaced inside "Tom's". Case-sensitive by default; a case-insensitive mode and a case-preserving heuristic (called a heuristic wherever it is named) are offered.
  • It previews first. The preview shows, per file, how many replacements, what happens to every coding, annotation and case link, any rows that would collide, and a residue report. It returns a preview token bound to that exact operation and to the rows it covers; the run executes only when you call again with the token, and a project that changed in between is refused rather than acted on. A mandatory backup is taken before the first write.
  • A coding that cut into a name is a choice, and both answers are offered. The default policy treats the pseudonym as the same token as the name: a coding on the name now sits on the pseudonym, a coding that cut into a name grows to contain the whole pseudonym, and nothing is ever deleted. The edit-parity policy reproduces the walk QualCoder's own coding-view editor applies to this tool's exact edit list, which deletes a coding sitting exactly on a name and trims one that touches it; the preview counts what each would do before you choose.
  • The residue report says where names may remain. Memos (twelve fields), journal entries and their names, case, file, code, category and attribute-type names, attribute values, and whether pseudonyms.json, speakers.json or speaker_regex.json is present. Those counts are a deliberately generous heuristic: they report anything a reader would see, including inside a longer word (Thomas_P01) and in any letter case, so they over-report rather than under-report. Read a high count as a list of fields to check, not as a count of names; an entry for a very short name ("Ed") makes them generous and the preview warns about it.
  • Hidden coders. On a project that hides coders, their rows move with the text like everyone else's, and the preview reports what the run would do to them as counts, never names: shifted (moved, same length), substituted (a coding that was exactly the name and is now exactly the pseudonym, whatever the two lengths), resized (a coding that contained a whole name and changed length only because that name did), snapped (a coding that cut into a name and had that boundary moved: grown to contain the whole pseudonym under the default policy, cut back to exclude it under edit parity), deleted (edit parity only) and clamped (a damaged row whose stored end lay past the end of the text and was pulled back to it). The override, allow_hidden_coder, is required when snapped, deleted or clamped is non-zero; a shift, a substitution and a resize change no coding decision and need nothing.
  • Every run leaves an audit trail that carries no original name: a journal entry in the project (as QualCoder's own PDF restructure does) and a run manifest under ~/.qualcoder_mcp/pseudonymisation/ recording pseudonyms, the replacement spans, the row ids and the old and new offsets of every moved row, enough for a later release to reverse a run exactly. In this release, reversal is restore_backup.
  • The import rider. import_text_file(apply_project_pseudonyms=true) applies the project's own pseudonyms.json to a text on the way in, as QualCoder does to every text file it imports. Default off.

What it does not do, said plainly because it decides what you may send afterwards:

  • It does not rewrite memos, journal entries, case, file, code, category or attribute-type names, or attribute values. It counts them.
  • It does not touch PDFs, media files or QualCoder 4.0's ai_data/ folder: they are neither rewritten nor scanned. QualCoder's own AI search index keeps the previous text until QualCoder reopens the project and re-indexes; its chat history may quote it.
  • It does not remove pseudonyms.json, the reverse key in plain text at the project root, and it never writes one.
  • The backup it takes holds the real names, and so does every earlier backup beside the project. A project is not pseudonymised while its backups sit next to it.
  • This server's own session files keep the excerpts they recorded before a run; the result lists the affected sessions and deletes none.
  • Pseudonymised data is still personal data. Removing names does not make a transcript safe to send anywhere: role, locality, events and phrasing re-identify. PRIVACY.md says what to check with your ethics board and DPO.

Also new

  • The AI coder name belongs to the project, and you choose it. The first write that needs a name stops and asks; your answer (set_project_ai_coder_name) is stored in qualcoder_mcp.json beside data.qda, so it travels with backups and copies and two hosts agree on it. Reads never ask. Earlier rows keep the name they were written under; the project remembers the names it has used, which is what makes a later comparison between two models possible. QUALCODER_MCP_AI_CODER_NAME is now this host's declaration, offered as the first quick pick; it never writes a row by itself.
  • Preview tokens on the six destructive tools. merge_codes, delete_code, delete_category, merge_category, restore_backup and prune_backups preview first and execute only with the token the preview returned, bound to the tool, the effect-deciding arguments, the project and a fingerprint of the rows it would touch. Every cascade preview says whose work is at stake: this project's AI rows, a per-owner breakdown of the rest, hidden coders as a count, and the private notes that would die with their rows.
  • compare_coders. Per code, how much of the text in scope each coder coded, how much they agreed, and two agreement coefficients, both always present: kappa_qualcoder, which reproduces QualCoder's own "Kappa" column expression for expression, and kappa_cohen, the textbook statistic. Read-only, full toolset only.
  • Ask what is not yet coded, and page through the answer. exclude_code_ids on the searches drops passages already coded under those codes (QualCoder 4.0's own rule, restricted to the codings you can see); the search and segment tools return cursors that survive a host recycling the server; get_coded_segments samples by strategy under a character budget.
  • Colours snapped onto QualCoder's 120-colour palette, with QualCoder's own matcher, and the result says when a colour was snapped. Idempotent creates: a code, category or case name that already exists, ignoring letter case, spacing and Unicode form, answers created: false with the existing row and makes no backup; no-op writes answer changed: false.
  • Methodology vocabulary and grounding rules in the analysis tools' guidance, a qualcoder://guidance/methods resource, and an instructions string in the MCP initialize handshake. Language, not enforcement: nothing replaces your approval of each suggestion.
  • qualcoder-mcp --version, and a plain-language notice when the server is started by hand in a terminal.

What changed

  • export_frequencies_csv names visible coders only in its result (the exported file is unchanged), and omits the coders key when the coder-visibility table cannot be read.
  • Content searches report true file positions (a fix to a one-character offset after U+0130) and now use the regex engine's case folding, which differs from str.lower() in exotic cases only.
  • Tied rows in search_coded_text and get_coded_segments have a defined, total order.
  • Coder names may not carry an invisible formatting character (Unicode category Cf, ZWNJ and ZWJ excepted). Coder visibility is described as a capability of QualCoder 3.8.2 and 4.0 (schema v14 and later), not a 4.0 feature, and needs the whole view set: a project with the column and a missing view fails closed.
  • The decision about who may be named is re-read from the project each time it is made, so a coder hidden after this server connected is treated as hidden by every preview, comparison and export listing; which table the r...
Read more

v0.11.0-alpha

v0.11.0-alpha Pre-release
Pre-release

Choose a tag to compare

@nicotem nicotem released this 07 Sep 15:08

v0.11.0-alpha: QualCoder 4.0 interop conventions

Pre-release. Install or upgrade with pip install --upgrade qualcoder-mcp, then restart your MCP host fully so it reloads the tool schemas. Full details: [CHANGELOG](https://github.com/nicotem/qualcoder_mcp/blob/main/CHANGELOG.md), [INSTALL](https://github.com/nicotem/qualcoder_mcp/blob/main/INSTALL.md), [PRIVACY](https://github.com/nicotem/qualcoder_mcp/blob/main/PRIVACY.md).

QualCoder 4.0 (currently a 4.0-Beta pre-release) ships its own AI assistant and defines conventions that live inside the project itself. This release makes qualcoder-mcp follow them, so a project touched by both tools behaves consistently. Schema-dependent conventions are detected by probing the project database, never by version strings; projects from QualCoder 3.8.x keep their previous behaviour. No tool was added or removed (67 in the full toolset, 20 in core); the changes are new optional arguments, new result fields, and a few changed defaults.

Highlights

  • Memo privacy (#####). Everything from the first ##### marker in a memo is the researcher's private zone. Every read now strips it silently, searches match only the public part (with no way to probe the private zone through result counts), and every memo write preserves an existing private zone verbatim. Exported files (REFI-QDA, codebook, coded-segments report) keep full memos for parity with QualCoder's own exports and say so.
  • Private notes and deletions. Deleting a coding or annotation whose memo carries a private note requires confirm_private_note_deletion=true, and a backup is always taken for such rows even when backups are switched off. Cascade previews report how many private notes an operation would remove.
  • Attribution. New optional QUALCODER_MCP_AI_CODER_NAME. The default stays AI Coding Assistant; set it to AI Agent to group this server's work with QualCoder 4.0's built-in assistant in its visibility, undo, and reports. All rows this server creates now carry the configured name (previously some carried the project's coder name or "MCP Import").
  • Coder visibility on 4.0 projects. Reads and analytics honour QualCoder's per-coder visibility by default and disclose when hidden-coder filtering shaped a result (a count, never names). A coder argument reads one coder's rows from the full data. Writes that target a hidden coder's row by id are refused unless allow_hidden_coder=true, and echoes for such rows carry ids only.
  • Backups and copies. Workspace copies now apply QualCoder's own backup ignore set (no more duplicating the plaintext search.sqlite index or live sidecars). Backups and copies skip symlinks that point outside the project, dangle, or loop, and report what they skipped; a copy that fails part-way cleans up after itself.
  • Detecting an open QualCoder 4.0 window. 4.0 removed the lock file this server relied on. select_project, get_current_project, analyze_for_coding, and the restore_backup preview now report qualcoder_gui_signals: heuristics from database sidecars, recent AI-index and chat activity, and a guarded local process scan. They warn and ask ("appears to be open"); they never refuse on their own. An open 4.0 window will not display external writes until the project is reopened.
  • Recovery hint. Every "no project selected" error names the last project used on this machine, so a host that recycles the server process (observed with LM Studio 0.4.12) recovers with one select_project call. The selection is never restored automatically.

Also in this release

British English is now the house spelling for all prose: documentation, tool and prompt descriptions, and runtime messages. Identifiers such as analyze_for_coding, sanitize_formulas and recolor_code are unchanged. The repository gains CITATION.cff (citation metadata with the maintainer's ORCID) and CONTRIBUTING.md (how to report issues, how changes are reviewed before merge, style rules and scope). The LM Studio recipe in INSTALL.md now records a functional verification against LM Studio 0.4.22; model-quality evaluation for local models is still planned work.

Upgrading from 0.10.x

No migration step. Project files gain no tables or columns; session files are unchanged; the only new on-disk state is ~/.qualcoder_mcp/mru_project.json. Every new argument is optional and 0.10 call shapes keep working. Behaviour changes on every project: private memo zones are no longer returned to the AI; non-coding rows are attributed to the configured AI coder name; private-note deletions need confirmation and always back up; copies skip outward symlinks and the search index. Behaviour changes only on 4.0 projects: visibility-honouring reads with a possible coder_visibility block, hidden-coder write refusals, and memo carry-over on merge_category. Messages now use colons, semicolons, or commas where they used em dashes. The deprecated session_id duplicate in session-tool responses is still emitted and will be removed in a later release; read coding_session_id.

Quality gates

Built against QualCoder master at pinned commit 9bddf17. Passed a QA round, a Security gate, and two re-verification rounds, each with adversarial verification of every finding; suite at the release commit: 1534 passed, 45 skipped, 0 failed; CI green on macOS, Ubuntu, and Windows with Python 3.10 and 3.13.

Not in this release

pseudonymise_source, deferred from 0.10, is planned for a later 0.11.x or 0.12 release. Support and bug reports: GitHub Issues only. This remains experimental research software; please back up your projects (the server also does, before every write).

v0.10.1-alpha

v0.10.1-alpha Pre-release
Pre-release

Choose a tag to compare

@nicotem nicotem released this 28 Aug 08:57

Patch release: session tools now work behind middleware that reserves session_id.

Field testing found that some MCP middleware (verified with the remote-devices bridge; the collision class applies to any gateway that reserves the name for its own routing) strips a tool argument named session_id before it reaches the server, breaking the entire suggest/review/apply and proposal pipelines. Direct stdio connections were unaffected.

  • All 14 session tools rename their parameter session_id → coding_session_id (AI clients adapt automatically from the schemas).
  • Responses emit coding_session_id and keep session_id as a deprecated duplicate for one release.
  • On-disk session files are unchanged — no migration; existing sessions work as-is.

If your session tools were failing with "Session None not found" style errors behind a gateway or bridge, upgrade:

pip install --upgrade qualcoder-mcp

1233 automated tests; CI green on macOS, Linux and Windows. Full changelog: CHANGELOG.md.

v0.10.0-alpha

v0.10.0-alpha Pre-release
Pre-release

Choose a tag to compare

@nicotem nicotem released this 25 Aug 09:21

Connect QualCoder to your MCP host of choice for AI-assisted qualitative data analysis, with a human approving every write.

This is an alpha. Please work on copies of your projects, never originals.

Install or update:

pip install --upgrade qualcoder-mcp

New in this release

  • QualCoder v14 through v17 project support, including sub-codes. Support is determined by inspecting the project database itself (capability probes), not version numbers, so projects from released QualCoder 3.8.x and from the unreleased 4.0 development builds work interchangeably. Sub-codes (a code nested under another code) are fully supported: create, move, merge and delete follow QualCoder's own recipes exactly, and listings, reports, the codebook and REFI-QDA export all preserve the nesting. Verified against the QualCoder 4.0 Beta source at pinned commit 7b074d2; claims will be re-verified at the 4.0 release.
  • Write protection for QualCoder 4.0 builds. The 4.0 development builds no longer use a lock file, so an open 4.0 window cannot be detected. Text-anchored writes now re-verify inside the write transaction that the file text still matches what was validated, rolling back cleanly if an editor raced the write. Do not run writes while any QualCoder window has the same project open.
  • Multi-host support (Experimental). A reduced core toolset (QUALCODER_MCP_TOOLSET=core, 20 tools) for smaller-context hosts, setup recipes for LM Studio (fully local: participant data never leaves your machine) and Claude Code with an Anthropic API key, and a four-rung data-governance ladder in PRIVACY.md quoting official terms verbatim. Experimental: functionally tested at the server level, but no local model has been capability-evaluated with this server yet.

1233 automated tests; independent QA and security review; CI green on macOS, Linux and Windows (Python 3.10 and 3.13).

Support & privacy

Bugs, questions and ideas via GitHub Issues (SUPPORT.md).

One important note: by design this tool sends your project content, including interview text, to whichever AI provider your host uses (none, with a fully local host). Please use synthetic or consented data and check your ethics and GDPR position before pointing it at real participant data. PRIVACY.md explains exactly what flows where.

Full changelog: CHANGELOG.md. MIT licensed, no warranty.

v0.9.0-alpha

v0.9.0-alpha Pre-release
Pre-release

Choose a tag to compare

@nicotem nicotem released this 30 Jul 10:07

QualCoder MCP is now on PyPI. Install with one command:

pip install qualcoder-mcp

(or pipx install qualcoder-mcp / uv tool install qualcoder-mcp for an isolated install). Full setup — including the Claude Desktop / Claude Code configuration — is in the README and INSTALL.md.

This is an alpha. Please work on copies of your projects, never originals.

What's in this release

This is a packaging and hardening release — no new analysis features (those land in 0.10). It exists to make the tool easy to install and update, and to close what a full security review found.

  • PyPI packaging — pip install qualcoder-mcp, and updates become pip install --upgrade qualcoder-mcp. No more git clone.
  • Whole-codebase security audit — an adversarial review of all 67 tools. It found the architecture sound (no attacker-reachable vulnerability) and produced three fixes: the write-cleanup guarantee extended to three older tools, an export path-resolution hardening, and SHA-pinned CI actions.
  • Critical dependency fix — capped the mcp SDK at <2. Its 2.0.0 removed the module this server is built on, which was silently breaking fresh installs; capped and verified.
  • Upgrade guide for existing testers — if you installed via git clone before 0.9, see INSTALL.md → Upgrading from an earlier (git) install. You can stay on git or switch to the PyPI install; your projects and AI-coding sessions are untouched by upgrading, and jumping 0.6/0.7/0.8 → 0.9 in one step is fine (no data migration).

Requires

Python 3.10+, a Claude client (Desktop or Code), QualCoder 3.8.x with at least one project. macOS, Linux, or Windows.

Support & privacy

Bugs, questions, ideas → GitHub Issues (SUPPORT.md).

One important note: by design this tool sends your project content — including interview text — to Claude (Anthropic) for analysis. Please use synthetic or consented data and check your ethics/GDPR position before pointing it at real participant data — PRIVACY.md explains exactly what flows where.

Full changelog: CHANGELOG.md. MIT licensed, no warranty.

v0.8.0-alpha

v0.8.0-alpha Pre-release
Pre-release

Choose a tag to compare

@nicotem nicotem released this 26 Jul 09:29

Connect QualCoder to Claude for AI-assisted qualitative data analysis — explore your coded data conversationally, code with a human-in-the-loop approval gate, and now: let the AI propose new codes you review before they exist.
This is an alpha. It works and it's been tested hard, but it's early — please work on copies of your projects, never originals.

New in this release

Inductive / open coding — the AI can propose brand-new codes from your data, with supporting evidence. You review, rename, recolour, merge or reject every proposal; nothing exists until you approve it, and creation is all-or-nothing with a backup first.
One-word span control — every coding suggestion carries precomputed shorter/longer span alternatives (sentence ↔ full speaker turn), so "longer on #3" fixes a quote instantly. Born directly from early tester feedback.
Report exports — codebook, coded-segments report, and frequency/matrix CSVs, with counting rules matched against QualCoder's own source so the numbers agree with your Reports screen (any deliberate divergence is stated in the output). Optional sanitize_formulas flag for spreadsheet safety.
Cases, attributes & annotations — create cases, define and set attributes (with correct numeric semantics), write annotations; plus merge_category and backup retention (prune_backups).
Editable suggestions — adjust a pending suggestion's span or code before approving (edit_suggestion).
Works in Claude Desktop, Claude Code, or any MCP client (new docs section).
Tool surface: 48 → 67. Test suite: 1130 automated tests, cross-platform CI (macOS/Linux/Windows, Python 3.10 & 3.13). Every feature implemented against QualCoder 3.8.2 source ground truth and gated through independent QA and security review. This release also fixes two subtle v0.7 bugs the ground-truth research exposed (numeric attribute queries matching unset values; case-link duplicate detection).

Safety

Read-only by default. Every write is preceded by an automatic backup, verified against QualCoder's format, and refused while QualCoder has the project open. Destructive operations preview first and take a safety backup.

Requires

Python 3.10+, a Claude client (Desktop or Code), QualCoder 3.8.x with at least one project. macOS, Linux, or Windows.

Support

Bugs, questions, and feature ideas via GitHub Issues — see SUPPORT.md. Early adopters shape the roadmap — this release's span controls and multi-code guidance came straight from the first tester.
Full changelog: CHANGELOG.md. MIT licensed, no warranty.