Skip to content

v0.12.0-alpha

Pre-release
Pre-release

Choose a tag to compare

@nicotem nicotem released this 16 Sep 20:42
· 322 commits to main since this release

v0.12.0-alpha: pseudonymisation that keeps the coding, and the AI coder name is yours

Pre-release. Install or upgrade with pip install --upgrade qualcoder-mcp, then restart your MCP host fully so it reloads the tool descriptions. This release raises the dependency floor to mcp>=1.17.0,<2; pip upgrades mcp with the package (see "Upgrading from 0.11.x"). Full details: CHANGELOG, INSTALL, PRIVACY.

v0.12 comes out of a ground-truth study of QualCoder 4.0 (master at pinned commit 9bddf17, and the 3.8.2 tag where schema v14 parity matters): seven design dossiers, owner rulings on every question only the owner could answer, then two batches, a flagship and one follow-up, each through a QA gate, a Security gate and re-verification until clean. Three tools are added, pseudonymise_source, set_project_ai_coder_name and compare_coders: 70 in the full toolset, 21 in core. Everything schema-dependent is still detected by probing the project database, never by version string.

The flagship: pseudonymise a project that is already coded

pseudonymise_source replaces names with pseudonyms in the stored text of the text sources you choose and moves every coding, annotation and case link with the text, all files in one transaction. QualCoder pseudonymises only at import, from pseudonyms.json, and has never had a way to pseudonymise a source that is already coded; this tool is that way.

What it does, in plain words:

  • You give it the mapping: real names to pseudonyms, with variants (nicknames, inflections) folded onto one pseudonym. Nothing is detected or guessed. The rewrite replaces whole words only, using QualCoder's own boundary rule, so "Ann" is not touched inside "Anna" while "Tom" is replaced inside "Tom's". Case-sensitive by default; a case-insensitive mode and a case-preserving heuristic (called a heuristic wherever it is named) are offered.
  • It previews first. The preview shows, per file, how many replacements, what happens to every coding, annotation and case link, any rows that would collide, and a residue report. It returns a preview token bound to that exact operation and to the rows it covers; the run executes only when you call again with the token, and a project that changed in between is refused rather than acted on. A mandatory backup is taken before the first write.
  • A coding that cut into a name is a choice, and both answers are offered. The default policy treats the pseudonym as the same token as the name: a coding on the name now sits on the pseudonym, a coding that cut into a name grows to contain the whole pseudonym, and nothing is ever deleted. The edit-parity policy reproduces the walk QualCoder's own coding-view editor applies to this tool's exact edit list, which deletes a coding sitting exactly on a name and trims one that touches it; the preview counts what each would do before you choose.
  • The residue report says where names may remain. Memos (twelve fields), journal entries and their names, case, file, code, category and attribute-type names, attribute values, and whether pseudonyms.json, speakers.json or speaker_regex.json is present. Those counts are a deliberately generous heuristic: they report anything a reader would see, including inside a longer word (Thomas_P01) and in any letter case, so they over-report rather than under-report. Read a high count as a list of fields to check, not as a count of names; an entry for a very short name ("Ed") makes them generous and the preview warns about it.
  • Hidden coders. On a project that hides coders, their rows move with the text like everyone else's, and the preview reports what the run would do to them as counts, never names: shifted (moved, same length), substituted (a coding that was exactly the name and is now exactly the pseudonym, whatever the two lengths), resized (a coding that contained a whole name and changed length only because that name did), snapped (a coding that cut into a name and had that boundary moved: grown to contain the whole pseudonym under the default policy, cut back to exclude it under edit parity), deleted (edit parity only) and clamped (a damaged row whose stored end lay past the end of the text and was pulled back to it). The override, allow_hidden_coder, is required when snapped, deleted or clamped is non-zero; a shift, a substitution and a resize change no coding decision and need nothing.
  • Every run leaves an audit trail that carries no original name: a journal entry in the project (as QualCoder's own PDF restructure does) and a run manifest under ~/.qualcoder_mcp/pseudonymisation/ recording pseudonyms, the replacement spans, the row ids and the old and new offsets of every moved row, enough for a later release to reverse a run exactly. In this release, reversal is restore_backup.
  • The import rider. import_text_file(apply_project_pseudonyms=true) applies the project's own pseudonyms.json to a text on the way in, as QualCoder does to every text file it imports. Default off.

What it does not do, said plainly because it decides what you may send afterwards:

  • It does not rewrite memos, journal entries, case, file, code, category or attribute-type names, or attribute values. It counts them.
  • It does not touch PDFs, media files or QualCoder 4.0's ai_data/ folder: they are neither rewritten nor scanned. QualCoder's own AI search index keeps the previous text until QualCoder reopens the project and re-indexes; its chat history may quote it.
  • It does not remove pseudonyms.json, the reverse key in plain text at the project root, and it never writes one.
  • The backup it takes holds the real names, and so does every earlier backup beside the project. A project is not pseudonymised while its backups sit next to it.
  • This server's own session files keep the excerpts they recorded before a run; the result lists the affected sessions and deletes none.
  • Pseudonymised data is still personal data. Removing names does not make a transcript safe to send anywhere: role, locality, events and phrasing re-identify. PRIVACY.md says what to check with your ethics board and DPO.

Also new

  • The AI coder name belongs to the project, and you choose it. The first write that needs a name stops and asks; your answer (set_project_ai_coder_name) is stored in qualcoder_mcp.json beside data.qda, so it travels with backups and copies and two hosts agree on it. Reads never ask. Earlier rows keep the name they were written under; the project remembers the names it has used, which is what makes a later comparison between two models possible. QUALCODER_MCP_AI_CODER_NAME is now this host's declaration, offered as the first quick pick; it never writes a row by itself.
  • Preview tokens on the six destructive tools. merge_codes, delete_code, delete_category, merge_category, restore_backup and prune_backups preview first and execute only with the token the preview returned, bound to the tool, the effect-deciding arguments, the project and a fingerprint of the rows it would touch. Every cascade preview says whose work is at stake: this project's AI rows, a per-owner breakdown of the rest, hidden coders as a count, and the private notes that would die with their rows.
  • compare_coders. Per code, how much of the text in scope each coder coded, how much they agreed, and two agreement coefficients, both always present: kappa_qualcoder, which reproduces QualCoder's own "Kappa" column expression for expression, and kappa_cohen, the textbook statistic. Read-only, full toolset only.
  • Ask what is not yet coded, and page through the answer. exclude_code_ids on the searches drops passages already coded under those codes (QualCoder 4.0's own rule, restricted to the codings you can see); the search and segment tools return cursors that survive a host recycling the server; get_coded_segments samples by strategy under a character budget.
  • Colours snapped onto QualCoder's 120-colour palette, with QualCoder's own matcher, and the result says when a colour was snapped. Idempotent creates: a code, category or case name that already exists, ignoring letter case, spacing and Unicode form, answers created: false with the existing row and makes no backup; no-op writes answer changed: false.
  • Methodology vocabulary and grounding rules in the analysis tools' guidance, a qualcoder://guidance/methods resource, and an instructions string in the MCP initialize handshake. Language, not enforcement: nothing replaces your approval of each suggestion.
  • qualcoder-mcp --version, and a plain-language notice when the server is started by hand in a terminal.

What changed

  • export_frequencies_csv names visible coders only in its result (the exported file is unchanged), and omits the coders key when the coder-visibility table cannot be read.
  • Content searches report true file positions (a fix to a one-character offset after U+0130) and now use the regex engine's case folding, which differs from str.lower() in exotic cases only.
  • Tied rows in search_coded_text and get_coded_segments have a defined, total order.
  • Coder names may not carry an invisible formatting character (Unicode category Cf, ZWNJ and ZWJ excepted). Coder visibility is described as a capability of QualCoder 3.8.2 and 4.0 (schema v14 and later), not a 4.0 feature, and needs the whole view set: a project with the column and a missing view fails closed.
  • The decision about who may be named is re-read from the project each time it is made, so a coder hidden after this server connected is treated as hidden by every preview, comparison and export listing; which table the read tools go to is still settled when the connection opens, so re-select the project after hiding a coder (PRIVACY.md says exactly what is and is not re-read).
  • Session files are written atomically; the backup folder's name (which can carry a participant's) no longer reaches the server log; the preview-token secret's creation and permissions are hardened; a failed atomic write no longer leaves a temp file behind on Windows; the AI coder name file can no longer outgrow its own reader.
  • The dependency floor is mcp>=1.17.0,<2: the core toolset calls FastMCP.remove_tool, which does not exist before mcp 1.17.0, so an install at the old floor (>=1.2.0) started with the full surface refused and an AttributeError.

Upgrading from 0.11.x

  • The mcp floor is >=1.17.0,<2. pip install --upgrade qualcoder-mcp upgrades mcp with the package. An environment that cannot move mcp past 1.16.x cannot install 0.12; such an environment could start 0.11 but could not honour QUALCODER_MCP_TOOLSET=core.
  • confirm is inert on the six destructive tools. A call with confirm=true and no preview_token returns the preview with a note; nothing executes. Take the token from the preview and call again with it (the preview's execute_with spells the call out). confirm stays in the signatures for this release and is removed in v0.13.
  • Duplicate names are compared case-insensitively, reversing the v0.10 rule that "Stress" and "stress" were two codes through this server. Creates answer with the existing row; renames to a case variant of another row are refused; pairs QualCoder's own GUI made are refused with the candidates listed.
  • The first write to each project asks for its AI coder name. Choosing AI Coding Assistant keeps continuity with everything v0.11 wrote; the ask says so when the project already holds rows under a known name. QUALCODER_MCP_AI_CODER_NAME declares, it no longer attributes, and owner= on apply_codings and import_text_file no longer chooses the name.
  • A sidecar file, qualcoder_mcp.json, appears in the project folder once a name has been chosen, and in every backup and copy of it, this server's and QualCoder's own. QualCoder ignores it; restore_backup puts back whatever the backup held; deleting it makes the next write ask again.
  • No migration step otherwise; project and session files are unchanged. Scripts that read session_id from a session-tool response must read coding_session_id.

Known limits

  • The six in-QualCoder acceptance checks for the pseudonymisation tool were prepared, and are recorded as not run in this release and planned for 0.12.1. A fixture project (two files, three names, eighteen codings including a hidden coder's, two annotations, two case links, names planted in every memo and label field) was built and rewritten with the tool, and 21 database-level checks on the result pass (every coding's stored text equals the text at its positions, every position inside the file, the hidden coder's rows moved consistently with the visible coder's, the backup holds the pre-run text, the journal entry and manifest carry no name). What only a QualCoder window can show, the coding screen's highlights sitting on the pseudonyms, the coding and annotation reports, the case file manager's underlines, edit mode and undo after a rewrite, and 4.0's re-indexing of the search index, was scripted step by step for both pinned builds and has not been run. Until it is, treat the tool's on-screen results in QualCoder as verified at the database level only. The script and fixture are kept, and running them is planned for 0.12.1.
  • Detecting an open QualCoder 4.0 window is a heuristic (4.0 writes no lock file), and an open 4.0 window does not show external writes until the project is reopened. Never write while any QualCoder window has the project open.
  • The residue report reads memos, labels and attribute values, not the file text itself: a name the whole-word rewrite leaves inside a longer word in a transcript (Thomas_Smith) is not counted anywhere in this release. A look-alike letter from another script is out of scope for both the rewrite and the count.
  • On a project that gains the coder-visibility capability while this server is connected, the read tools go to the base tables until the project is re-selected; PRIVACY.md's "Coder visibility" section says what is and is not re-read.
  • QualCoder's own coding report lists a hidden coder's segments in both pinned builds, because it reads the base table; hiding a coder in QualCoder hides their work from its coding screen and from this server's default reads, not from its reports, and this server's file exports keep that same parity.
  • The edit-parity policy applies QualCoder's editing walk to this tool's exact edit list; QualCoder's own editor, fed by a diff library, may keep a coding this policy deletes.
  • The multi-host recipes (LM Studio, Claude Code with an API key) remain Experimental; no local model has been evaluated for coding quality with this server.

Quality gates

Built against QualCoder master at pinned commit 9bddf17 and the 3.8.2 tag. Batch A, Batch B and the flagship each passed a QA gate and a Security gate followed by re-verification rounds (the flagship through five fix rounds), then two follow-ups on the hidden-coder override; six-platform CI (Ubuntu, Windows, macOS on Python 3.10 and 3.13) green on every merge. Suite at the release commit: 3035 passed, 3 skipped, 0 failed, on Python 3.13.5 and 3.11.13, and again under the regenerated uv.lock. The version canary was checked against a fresh clone and a fresh virtual environment.

Not in this release

Planned for v0.13: quote-anchored writes with a fuzzy fallback; a persistent change journal with undo; media region coding; the reverse pseudonymisation run over a manifest; a count of names remaining in the file text beside the residue report, and both readings (wide and whole-word) in that report; rewriting the public part of memos under a separate switch. Support and bug reports: GitHub Issues only. This remains experimental research software; back up your projects (the server also does, before every write).