Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
68 changes: 34 additions & 34 deletions Docs/CALIBRATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,8 +10,8 @@ It is **not an accuracy figure**. Accuracy needs machine-written text to measure

- **Corpus** `signsofai-human-baseline`, fingerprint `123fa5b9ebca3f29`
- **Texts** 90 (280,221 words)
- **Engine** SignsOfAI.Core 0.2.1
- **Run** 2026-08-03
- **Engine** SignsOfAI.Core 0.3.0
- **Run** 2026-08-04
- **Target false-positive rate** 5%

Every text here was published before generative models could have written it. That is the whole basis for calling it human, and it is a stronger guarantee than any classifier offers about anything. The manifest names each source, its licence and its year, so the claim can be traced rather than trusted.
Expand All @@ -28,8 +28,8 @@ A rate that holds in English and fails in Spanish is not one number, and reporti

| Group | Texts | Median | 90th pct | Highest | Threshold for 5% | Best bound it can support |
|---|---|---|---|---|---|---|
| **en** | 65 | 6.4 | 12.3 | 23.4 | — | 5.6% |
| **es** | 25 | 7.7 | 15.4 | 18.9 | — | 13.3% |
| **en** | 65 | 5.8 | 11.8 | 23.4 | — | 5.6% |
| **es** | 25 | 7.2 | 15.1 | 18.4 | — | 13.3% |

A dash means this group has too few texts to bound that rate at all — with nothing flagged it still takes roughly seventy-five before the interval alone gets under 5%. That is a statement about the corpus, not the tool.

Expand All @@ -39,22 +39,22 @@ The reason the whole exercise exists. If this project cannot show a rate for sec

| Group | Texts | Median | 90th pct | Highest | Threshold for 5% | Best bound it can support |
|---|---|---|---|---|---|---|
| **en-anglophone-affiliation** | 21 | 7.0 | 9.8 | 15.4 | — | 15.5% |
| **en-other-affiliation** | 19 | 6.4 | 14.1 | 18.3 | — | 16.8% |
| **en-wikipedia** | 25 | 5.2 | 10.4 | 23.4 | — | 13.3% |
| **es-wikipedia** | 25 | 7.7 | 15.4 | 18.9 | — | 13.3% |
| **en-anglophone-affiliation** | 21 | 5.9 | 9.0 | 14.2 | — | 15.5% |
| **en-other-affiliation** | 19 | 6.2 | 13.6 | 18.0 | — | 16.8% |
| **en-wikipedia** | 25 | 4.9 | 10.4 | 23.4 | — | 13.3% |
| **es-wikipedia** | 25 | 7.2 | 15.1 | 18.4 | — | 13.3% |

A dash means this group has too few texts to bound that rate at all — with nothing flagged it still takes roughly seventy-five before the interval alone gets under 5%. That is a statement about the corpus, not the tool.

Across these groups the median score runs from 7.7 (**es-wikipedia**) down to 5.2 (**en-wikipedia**), a spread of 2.5 points on a scale of a hundred. The longest tail belongs to **es-wikipedia** at 15.4 for the ninetieth percentile. A tool with the defect this project criticises would show one group sitting well above the rest; on this corpus none does. It is a first indication rather than a finding — these are tens of texts, not hundreds — and the numbers move as the corpus grows, in whichever direction they move.
Across these groups the median score runs from 7.2 (**es-wikipedia**) down to 4.9 (**en-wikipedia**), a spread of 2.3 points on a scale of a hundred. The longest tail belongs to **es-wikipedia** at 15.1 for the ninetieth percentile. A tool with the defect this project criticises would show one group sitting well above the rest; on this corpus none does. It is a first indication rather than a finding — these are tens of texts, not hundreds — and the numbers move as the corpus grows, in whichever direction they move.

## Every threshold

| Score at or above | Human texts flagged | Rate | 95% interval |
|---|---|---|---|
| 5 | 75 / 90 | 83.3% | 74.3% – 89.6% |
| 10 | 18 / 90 | 20% | 13% – 29.4% |
| 15 | 7 / 90 | 7.8% | 3.8% – 15.2% |
| 5 | 61 / 90 | 67.8% | 57.6% – 76.5% |
| 10 | 16 / 90 | 17.8% | 11.2% – 26.9% |
| 15 | 6 / 90 | 6.7% | 3.1% – 13.8% |
| 20 | 1 / 90 | 1.1% | 0.2% – 6% |
| 25 | 0 / 90 | 0% | 0% – 4.1% |
| 30 | 0 / 90 | 0% | 0% – 4.1% |
Expand All @@ -79,31 +79,31 @@ Every rule below fired on text no machine wrote, so each hit is a false positive

| Rule | Texts it fired on | Share | Total hits |
|---|---|---|---|
| `rhet.rule-of-three` | 45 | 50% | 100 |
| `stat.burstiness` | 25 | 27.8% | 25 |
| `lex.moreover` | 23 | 25.6% | 67 |
| `lex.furthermore` | 23 | 25.6% | 55 |
| `rhet.in-order-to` | 22 | 24.4% | 60 |
| `lex.just` | 19 | 21.1% | 30 |
| `lex.facilitate` | 17 | 18.9% | 29 |
| `rhet.in-terms-of` | 16 | 17.8% | 30 |
| `lex.comprehensive` | 14 | 15.6% | 36 |
| `lex.simply` | 14 | 15.6% | 24 |
| `lex.ademas` | 12 | 13.3% | 28 |
| `lex.utilize` | 11 | 12.2% | 28 |
| `rhet.not-only-but` | 11 | 12.2% | 19 |
| `lex.robust` | 10 | 11.1% | 22 |
| `lex.notably` | 10 | 11.1% | 17 |
| `rhet.regla-de-tres` | 9 | 10% | 14 |
| `lex.actually` | 9 | 10% | 11 |
| `rhet.in-conclusion` | 9 | 10% | 11 |
| `syn.serves-as` | 9 | 10% | 10 |
| `lex.importantly` | 8 | 8.9% | 11 |
| `rhet.weasel-attribution` | 8 | 8.9% | 11 |
| `rhet.with-regard-to` | 8 | 8.9% | 11 |
| `lex.crucial` | 8 | 8.9% | 9 |
| `rhet.in-terms-of` | 9 | 10% | 22 |
| `rhet.not-only-but` | 8 | 8.9% | 16 |
| `rhet.in-order-to` | 7 | 7.8% | 33 |
| `lex.furthermore` | 7 | 7.8% | 24 |
| `lex.robust` | 7 | 7.8% | 19 |
| `lex.just` | 7 | 7.8% | 16 |
| `lex.simply` | 7 | 7.8% | 16 |
| `lex.utilizar` | 7 | 7.8% | 15 |
| `rhet.in-this-article` | 7 | 7.8% | 15 |
| `lex.notably` | 7 | 7.8% | 14 |
| `syn.superficial-ing` | 7 | 7.8% | 12 |
| `rhet.with-regard-to` | 7 | 7.8% | 10 |
| `lex.crucial` | 7 | 7.8% | 8 |
| `syn.serves-as` | 7 | 7.8% | 8 |
| `lex.moreover` | 6 | 6.7% | 35 |
| `lex.comprehensive` | 6 | 6.7% | 27 |
| `lex.utilize` | 6 | 6.7% | 22 |
| `lex.facilitate` | 6 | 6.7% | 16 |
| `rhet.rule-of-three` | 6 | 6.7% | 16 |
| `lex.importantly` | 6 | 6.7% | 9 |
| `rhet.weasel-attribution` | 6 | 6.7% | 9 |
| `lex.actually` | 6 | 6.7% | 8 |
| `rhet.in-conclusion` | 6 | 6.7% | 8 |
| `rhet.important-note` | 6 | 6.7% | 7 |

A rule near the top is not automatically wrong. Some tells genuinely appear in human academic prose and the catalog says so. But a rule firing on most human texts is measuring the genre rather than the machine, and should be reweighted or retired.

Expand Down
22 changes: 20 additions & 2 deletions src/SignsOfAI.Cli/Program.cs
Original file line number Diff line number Diff line change
Expand Up @@ -183,7 +183,11 @@
findings = result.Findings.Select(f => new
{
f.RuleId, category = f.Category.ToString(), severity = f.Severity.ToString(),
f.MatchedText, f.Message, f.Suggestion, f.Evidence
f.MatchedText, f.Message, f.Suggestion, f.Evidence,
// Reported so a consumer can tell the two apart: this one matched, and it matched at a
// rate people write at, so it counts for nothing. Leaving it out would make the findings
// list and the score disagree with no way to see why.
f.AtHumanRate
}),
// Kept in its own object rather than folded in with the findings: these are characters at
// offsets, not judgements about prose, and a consumer should not have to tell them apart.
Expand Down Expand Up @@ -324,6 +328,17 @@ static void PrintBaseline(string path, BaselineReport r, bool useColor)
Console.WriteLine();
}

// The headline counts evidence, so it agrees with the category tallies and with the score. What
// matched at a human rate is named beside it rather than folded in: a reader who sees eleven
// highlights under a headline of three deserves to know why, on the same line.
static int Counted(AnalysisResult r) => r.Signals.Count;

static string AtRate(AnalysisResult r)
{
var n = r.Observations.Count;
return n == 0 ? "" : $" + {n} at a human rate";
}

static void PrintReport(string path, AnalysisResult r, int top, bool useColor)
{
string Col(string s, int code) => useColor ? $"[{code}m{s}" : s;
Expand All @@ -333,7 +348,7 @@ static void PrintReport(string path, AnalysisResult r, int top, bool useColor)
Console.WriteLine();
Console.WriteLine(Bold($" ✍ Signs of AI Writing — {Path.GetFileName(path)}"));
Console.WriteLine($" {Col($"{r.OverallScore:0}/100", scoreColor)} {Bold(r.Verdict)} " +
$"({r.Findings.Count} signal{(r.Findings.Count == 1 ? "" : "s")}, {(r.Language == "es" ? "Español" : "English")})");
$"({Counted(r)} signal{(Counted(r) == 1 ? "" : "s")}{AtRate(r)}, {(r.Language == "es" ? "Español" : "English")})");
Console.WriteLine($" words {r.Statistics.WordCount} · sentences {r.Statistics.SentenceCount} · " +
$"burstiness {r.Statistics.Burstiness:0.00} · lexical diversity {r.Statistics.LexicalDiversity:0.00}");

Expand All @@ -349,9 +364,12 @@ static void PrintReport(string path, AnalysisResult r, int top, bool useColor)
foreach (var f in shown)
{
int sev = f.Severity switch { Severity.High => 31, Severity.Medium => 33, Severity.Low => 36, _ => 90 };
if (f.AtHumanRate) sev = 90;
var head = $" {Col("●", sev)} [{f.Category}] " + (string.IsNullOrEmpty(f.MatchedText) ? "" : Bold(f.MatchedText));
Console.WriteLine(head.TrimEnd());
Console.WriteLine($" {f.Message}");
if (f.AtHumanRate)
Console.WriteLine(Col(" used here at a rate people write at — shown, not counted", 90));
Console.WriteLine(Col($" → {f.Suggestion}", 90));
}
if (r.Findings.Count > shown.Count)
Expand Down
7 changes: 6 additions & 1 deletion src/SignsOfAI.Core/AiWritingAnalyzer.cs
Original file line number Diff line number Diff line change
Expand Up @@ -72,13 +72,18 @@ public AnalysisResult Analyze(string text, string? language = null, IReadOnlyLis
Statistics = statistics,
};

var findings = _analyzers
var matched = _analyzers
.SelectMany(a => a.Analyze(context))
.Select(f => normalized.Changed ? ToSource(f, normalized, text) : f)
.OrderBy(f => f.Span.Start)
.ThenBy(f => f.Span.Length)
.ToList();

// Rules that measured the genre rather than the machine are silenced here, against rates taken
// from writing that predates generative models. Nothing is re-decided per finding: a rule is
// either used at a human rate in this text, and says nothing, or it is not.
var findings = GenreGate.Apply(matched, rulePack, statistics.WordCount);

var (overall, byCategory) = Scorer.Score(findings, statistics);

return new AnalysisResult
Expand Down
16 changes: 15 additions & 1 deletion src/SignsOfAI.Core/Calibration/CalibrationModel.cs
Original file line number Diff line number Diff line change
Expand Up @@ -25,8 +25,22 @@ public sealed record CalibrationSample

public required int WordCount { get; init; }

/// <summary>Every rule that fired on this human text — each one a false positive by construction.</summary>
/// <summary>
/// The rules that produced <em>evidence</em> on this human text — each one a false positive by
/// construction, since no machine wrote any of it.
///
/// This excludes rules the text used at a rate people write at, which are shown to a reader but
/// score nothing. Reporting those here would make the published misfire table look unchanged
/// while the scores moved, which is the opposite of informative.
/// </summary>
public required IReadOnlyList<string> RuleIds { get; init; }

/// <summary>
/// Everything that matched, including what was found at a human rate. Two things need it: the
/// derivation of the rates themselves, which must see all usage or it would measure the effect of
/// its own previous output, and the published count of how much the rates are absorbing.
/// </summary>
public IReadOnlyList<string> MatchedRuleIds { get; init; } = [];
}

/// <summary>
Expand Down
33 changes: 31 additions & 2 deletions src/SignsOfAI.Core/Model/AnalysisResult.cs
Original file line number Diff line number Diff line change
Expand Up @@ -12,10 +12,39 @@ public sealed record AnalysisResult
/// <summary>Language code actually used for analysis ("en" or "es").</summary>
public required string Language { get; init; }

/// <summary>All findings, ordered by position in the text.</summary>
/// <summary>
/// Everything that matched, ordered by position in the text — both the findings that count as
/// evidence and the ones the writer is using at a rate people write at. Highlighting works from
/// this list, and so does the rewriter, which should offer to replace a word regardless of whether
/// its rate proves anything.
///
/// <b>Do not count this to report how many signals were found.</b> Use <see cref="Signals"/>. The
/// two differ exactly when the genre gate has marked something, and a host that counts this list
/// will contradict its own category tallies and its own score.
/// </summary>
public required IReadOnlyList<Finding> Findings { get; init; }

/// <summary>Score contribution per category (0–100).</summary>
/// <summary>
/// The findings that are evidence: everything except what the text uses at a human rate. This is
/// what "N signals" means, it is what <see cref="CategoryScores"/> counts, and it is what moved
/// <see cref="OverallScore"/>.
///
/// Derived here rather than left to each host on purpose. The distinction arrived as a flag on
/// <see cref="Finding"/>, and within a day the five hosts disagreed about it: one counted
/// correctly, one reported a headline that its own category chips contradicted, and the MCP server
/// handed an agent, as "why this reads like AI", matches the engine had already ruled out. A
/// semantic that every consumer has to remember is a semantic that some consumer will forget.
/// </summary>
public IReadOnlyList<Finding> Signals => field ??= [.. Findings.Where(f => !f.AtHumanRate)];

/// <summary>
/// What matched but is not evidence: rules this text uses at a rate measured on writing published
/// before generative models existed. Worth showing — "you use 'furthermore' about as often as
/// other people do" is a useful thing to be told — and worth nothing to the score.
/// </summary>
public IReadOnlyList<Finding> Observations => field ??= [.. Findings.Where(f => f.AtHumanRate)];

/// <summary>Score contribution per category (0–100), counting <see cref="Signals"/> only.</summary>
public required IReadOnlyList<CategoryScore> CategoryScores { get; init; }

/// <summary>Overall "reads like AI" score, 0 (human) – 100 (unmistakably AI).</summary>
Expand Down
14 changes: 14 additions & 0 deletions src/SignsOfAI.Core/Model/Finding.cs
Original file line number Diff line number Diff line change
Expand Up @@ -30,4 +30,18 @@ public sealed record Finding

/// <summary>Contribution of this finding to the overall score (higher = stronger AI signal).</summary>
public double Weight { get; init; }

/// <summary>
/// True when this rule is being used in this text at a rate people write at, measured against
/// writing published before generative models existed. The finding still describes something real
/// and is still shown, but it is not evidence of a machine and contributes nothing to the score.
///
/// It is marked rather than removed on purpose. Ninety human academic papers produced a median of
/// seven flagged tells each, which is the sort of thing that teaches a reader to disbelieve the
/// eighth — but a tool whose whole argument is *showing the evidence* cannot answer that by hiding
/// evidence. "Furthermore appears once in three thousand words, which is how people write" is a
/// more useful thing to tell someone than silence, and it keeps the finding available to the
/// rewriter for a writer whose goal is removing the word rather than proving anything.
/// </summary>
public bool AtHumanRate { get; init; }
}
68 changes: 68 additions & 0 deletions src/SignsOfAI.Core/Rules/GenreGate.cs
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
using SignsOfAI.Core.Model;

namespace SignsOfAI.Core.Rules;

/// <summary>
/// Marks the findings of rules that are describing the genre rather than the machine.
///
/// Measuring the analyzer against ninety texts published before generative models existed produced a
/// number worth staring at: a median of <b>seven</b> flagged tells per human academic paper, 888
/// across the corpus, with only two of the ninety coming back clean. The score survived that — the
/// median was 6.8 out of 100 and nothing reached the recommended threshold — because the rules
/// involved carry weights of one and two. So the false-positive *rate* was never the problem.
///
/// The problem is the evidence panel, and it costs this project more than it would cost anyone else,
/// because showing the evidence instead of a percentage is the entire argument. A teacher who pastes
/// a colleague's paper and gets seven confident-looking tells learns not to believe the eighth, which
/// is the one that mattered.
///
/// <para><b>Why this marks rather than deletes.</b> Deleting was the first attempt and it was wrong in
/// three ways that are worth keeping written down. It made evidence vanish as a document grew, which
/// also handed anyone a way to bury a tell by padding. It cut the live rewriter off from words it
/// knows how to replace, for a writer whose goal is removing "utilize", not proving anything about it.
/// And it answered "our evidence panel is noisy" by hiding evidence, from a tool that exists to show
/// it. Marking keeps every finding visible and honest about what it is: present, and present at a rate
/// people write at.</para>
///
/// <para><b>What this does not fix.</b> The rate is measured per rule, and language models are tuned
/// away from repeating any single tell — the recognisable shape of machine prose is fifteen different
/// tells appearing once each, every one of them individually at a human rate. Marking rather than
/// deleting means that shape stays visible and countable, but nothing here scores it. Measuring
/// breadth rather than density is a separate question and it needs its own evidence, not a constant
/// chosen today.</para>
/// </summary>
public static class GenreGate
{
/// <summary>
/// <paramref name="findings"/> with the genre-rate ones flagged <see cref="Finding.AtHumanRate"/>.
/// Findings whose rule carries no measured human rate are returned untouched.
///
/// <paramref name="wordCount"/> must be the word count of the same text the findings came from,
/// as counted by <c>StatisticsCalculator</c> — the thresholds are derived against that counter, so
/// a different one silently rescales every comparison.
/// </summary>
public static IReadOnlyList<Finding> Apply(
IReadOnlyList<Finding> findings, RulePack pack, int wordCount)
{
if (findings.Count == 0 || wordCount <= 0) return findings;

var thresholds = pack.HumanRates;
if (thresholds.Count == 0) return findings;

var counts = new Dictionary<string, int>(StringComparer.Ordinal);
foreach (var f in findings)
if (thresholds.ContainsKey(f.RuleId))
counts[f.RuleId] = counts.GetValueOrDefault(f.RuleId) + 1;

if (counts.Count == 0) return findings;

var atHumanRate = new HashSet<string>(StringComparer.Ordinal);
foreach (var (ruleId, hits) in counts)
if (hits / (double)wordCount * 1000.0 <= thresholds[ruleId])
atHumanRate.Add(ruleId);

return atHumanRate.Count == 0
? findings
: [.. findings.Select(f => atHumanRate.Contains(f.RuleId) ? f with { AtHumanRate = true } : f)];
}
}
Loading
Loading