Skip to content

Ticket AI Reading Developer Guide

Ed Mozley edited this page Aug 31, 2026 · 2 revisions

Ticket AI Reading β€” Developer Guide

The two AI features from discussion #104: the maintained summary at the top of a ticket, and "Read it for me".

The three automatic (non-AI) behaviours are in Reading Long Tickets β€” Developer Guide. The user-facing page is Reading Long Tickets.

πŸ”’ THE INVARIANT: nothing here ever deletes content

Neither feature removes, rewrites or replaces anything. Read this before changing either of them:

  • Neither one touches emails or ticket_notes. Both read. The only table they write is ticket_ai_summaries, which is additive-only β€” INSERT and SELECT, no UPDATE of a summary and no DELETE anywhere in the feature.
  • A refresh writes a new version. It never overwrites the old one. $version = $latest ? ((int)$latest['version'] + 1) : 1; β€” and every earlier version stays readable in History.
  • The summary never replaces the conversation. It is an extra pane above a thread that renders exactly as it did before. Switch the feature off and the ticket is unchanged; the stored summaries are simply not shown.
  • It says what it has not read. last_email_id records how far it got, so the panel states "4 messages have arrived since this was written" rather than quietly presenting a stale reading as current.
  • A truncated answer is recorded as truncated, not served as though it were whole.

Why so absolute: this is the one feature in FreeITSM that could quietly become the only thing anybody reads. A summary that silently stands in for the ticket is a worse outcome than no summary, so every design decision here is arranged to keep the real conversation in front of the reader and to keep the summary visibly, checkably second-hand.

The companion guide's boundary note applies here too: stripInboundThread() at ingestion stores only the new part of an inbound reply. Nothing on this page does that β€” these features read whatever is stored.


Files

File What it holds
includes/ticket_ai.php Settings, the transcript builder, staleness helpers, both prompts
api/tickets/ai_summary.php GET current / GET history / POST new version
api/tickets/ai_read.php GET stored briefing / POST new one
assets/js/inbox.js The panel, the modal, the busy guards, the counting wait
assets/css/inbox.css .ai-summary*, .ai-panel*, .ai-history*
database/freeitsm.sql Β· includes/db_verify_schema.php ticket_ai_summaries

The table

One table, two kinds:

CREATE TABLE IF NOT EXISTS `ticket_ai_summaries` (
    `id`            INT NOT NULL AUTO_INCREMENT,
    `ticket_id`     INT NOT NULL,
    -- 'summary' = the standing panel at the top of the ticket; 'read' = a
    -- "read it for me" briefing. One table because they want the same three
    -- things β€” a version history, a record of how much was read, and a way to
    -- know the conversation has moved on since. Two tables would be two sets
    -- of those rules, and they would drift.
    `kind`          VARCHAR(16) NOT NULL DEFAULT 'summary',
    -- Numbered per (ticket, kind), so a briefing and a summary count separately.
    `version`       INT NOT NULL DEFAULT 1,
    `summary`       MEDIUMTEXT NOT NULL,
    …
    `last_email_id` INT NULL,
    `generated_by`  INT NULL,
    `truncated`     TINYINT(1) NOT NULL DEFAULT 0,
    KEY `ix_ticket_ai_summaries_ticket` (`ticket_id`, `kind`, `version`),

generated_by NULL means FreeITSM refreshed it by itself; an analyst id means somebody pressed the button. Recording an analyst on an automatic refresh would make it look as though a person had asked for it.


πŸ”‘ Staleness is a fact, not a guess

The obvious implementation compares timestamps. That is wrong twice over: two messages can share a received_datetime to the second, and a summary written at 14:03 may or may not have read a message stored at 14:03.

last_email_id is what the summary actually read, so the comparison is exact:

/* By id rather than by timestamp: two messages can share a received time to
   the second, and "newer than the last one I read" has to be exact. */
$stmt = $conn->prepare("SELECT COUNT(*) FROM emails WHERE ticket_id = ? AND id > ?");
$stmt->execute([$ticketId, $lastEmailId]);
return (int)$stmt->fetchColumn();

⚠️ The mark moves even for a message with an empty body:

// Recorded even when the body turns out to be empty: this is "how far I
// have read", and an empty message still moves that mark forward.
$lastEmailId = (int)$m['id'];
if ($plain === '') continue;

Skip that and an attachment-only message makes every subsequent summary report itself as permanently one behind.

The transcript

One builder for both features (and available to the merge summary). It had been written twice before this and the copies had not disagreed yet, which is the only reason nobody had noticed.

Read from the RECENT end

/* The most RECENT messages, not the first ones. A ticket that overflows the
   cap overflows it at the old end, and the old end is the part somebody
   picking this up cares least about. Fetched newest-first under the limit
   and then reversed, so the transcript still reads forwards. */
$stmt = $conn->prepare(
    "SELECT e.id, e.direction, e.from_name, e.from_address, e.received_datetime,
            e.subject, e.body_content, e.body_type
       FROM emails e
      WHERE e.ticket_id = ?
   ORDER BY e.received_datetime DESC, e.id DESC
      LIMIT $maxMessages"
);
$stmt->execute([$ticketId]);
$messages = array_reverse($stmt->fetchAll(PDO::FETCH_ASSOC));

$maxMessages is interpolated, which is safe only because it is clamped to an int first β€” max(1, min(500, (int)…)) at the top of the function. Keep it that way; a LIMIT placeholder does not bind portably in emulated-prepare mode.

Say when you have truncated

// Did we leave anything behind? Worth saying out loud in the prompt, so the
// model does not confidently describe a beginning it never saw.
$stmt = $conn->prepare("SELECT COUNT(*) FROM emails WHERE ticket_id = ?");
$stmt->execute([$ticketId]);
$truncated = (int)$stmt->fetchColumn() > count($messages);

…and it goes into the prompt as words:

$user .= "NOTE: this ticket is longer than what follows. You are seeing the "
       . "most recent part of the conversation only. Do not describe how the "
       . "ticket began unless the text below actually says.\n\n";

Without it the model narrates an opening it never saw, confidently and wrongly.

Bodies are stripped to text

Markup is noise that costs tokens β€” and stripping it means no HTML from a stranger's email is ever echoed back into anything FreeITSM renders.


The prompts: the rules are prohibitions

The two rules doing the real work in the summary prompt are both about what not to do (this is the prompt text as the model receives it, assembled in ticketAiSummarySystemPrompt()):

- State only what the conversation says. Never infer a cause, a fix or a
  resolution nobody wrote down.
- If something important is unclear or missing, say it is unclear. That is a
  useful answer, not a failure.

A summary that invents a fix reads exactly like one that did not. And a summary that hides its uncertainty removes the only signal telling somebody to go and read the thread themselves.

"Read it for me" is allowed to suggest, and therefore has to mark suggestions (from ticketAiReadSystemPrompt()):

What I would do next
- concrete suggestions. Begin this section with the line 'These are suggestions,
  not conclusions.'
…
- Separate what the ticket SAYS from what you are guessing, every time.

Both are told internal notes are staff-only, and never to phrase anything as if it were going to the customer.


πŸ’· Spending guards

Every one of these exists because the alternative charges somebody.

The server decides, not the browser

$auto = !empty($input['auto']);
if ($auto) {
    /* The server decides, not the browser. Two tabs opening the same ticket
       would otherwise bill twice for the same summary, and a page that
       refreshes on a timer would bill for ever. */
    if ($settings['summary_auto_after'] <= 0) {
        echo json_encode(['success' => false, 'error' => 'auto_disabled']);
        exit;
    }
    if ($latest !== null && $behind < $settings['summary_auto_after']) {
        echo json_encode(['success' => false, 'error' => 'not_due']);
        exit;
    }
}

The page asks for an automatic refresh; this decides. Both refusals happen before aiProviderChat() is reached, so a refused request costs nothing.

No background job, deliberately

An automatic refresh only ever happens when somebody opens the ticket. A nightly sweep across an open queue would bill for summaries of tickets nobody looked at, and the bill would arrive before the feature had been any use to anyone. Cost is therefore proportional to tickets actually read.

Reopening a briefing is free

ai_read.php answers GET from the table and never touches the provider. Only a re-read spends anything.

Measured: 62 seconds β†’ 0.16 seconds. The first build stored nothing, on the reasoning that a stored suggestion becomes a fact the next reader takes for a person's assertion β€” but the lived experience of that was waiting a minute for a briefing you had already read. Storing it and labelling it properly is the better trade, and the labelling is the same as the summary's: datestamp, what it read, an AI badge, and how far behind it is.

The defaults are OFF

// The maintained summary panel at the top of a ticket (idea 7).
'summary_enabled'       => 0,
…
'read_enabled'          => 0,

with a catch that falls back to the same array:

} catch (Exception $e) {
    // Defaults β€” and the defaults are OFF, so a settings table we cannot
    // read can never start spending somebody's money.
}

Every other setting in this area changes how something already free is displayed. These two spend money with somebody else's API key, so a settings table we cannot read must fail towards not spending.

Two clicks must not charge twice

async function runReadForMe() {
    if (_aiReadBusy || !_aiReadTicketId) return;
    _aiReadBusy = true;

⚠️ This one shipped broken. A reasoning model takes a minute; the panel showed a motionless "Reading the ticket…" for all of it; and a motionless wait is precisely what makes somebody click again. Every click was a fresh paid call whose answer overwrote the previous one. The guard is the fix β€” the counting wait below only stops it looking broken.

const started = Date.now();
const tick = setInterval(() => {
    const el = document.getElementById('aiReadWait');
    if (el) el.textContent = t('tickets.ai.working_secs').replace('{n}', Math.round((Date.now() - started) / 1000));
}, 1000);

Counts up rather than animating: "38s" says it is still going and roughly how long this model takes on this desk. Cleared in finally so it cannot outlive the request.

PHP must not time out mid-call

/* ⚠️ A reasoning model on a long ticket takes a MINUTE β€” measured at 62s on a
   real one. PHP's default limit would kill the request after the provider had
   already been paid and before anything was written down, which is the worst
   of both. */
@set_time_limit(300);

⚠️ The reasoning-model trap

Both endpoints hit this, and so does every other AI feature. Full write-up in AI Providers; the part that matters here:

A reasoning model spends its output budget thinking before it writes a single character. With a modest max_tokens you get HTTP 200, a full usage record and content: "". Measured on qwen/qwen3.7-plus, a one-sentence question spent 375 reasoning tokens of 383.

So neither endpoint treats an empty answer as a generic failure:

if ($text === '') {
        /* Nothing came back, and there are two very different reasons for that.
           A reasoning model that ran out of budget mid-thought is a SETTINGS
           problem with a fix the administrator can act on; anything else is not.
           Reported as itself, because "that did not work" sends somebody looking
           at their API key for a problem that is nowhere near it. */
        $ranOut = in_array($result['finish_reason'] ?? '', ['length', 'max_tokens'], true)
                  || (int)($result['reasoning_tokens'] ?? 0) > 0;
        echo json_encode(['success' => false, 'error' => $ranOut ? 'reasoning_overran' : 'empty_response']);
        exit;
}

And a truncated answer β€” worse than an empty one, because it reads almost complete and the section it loses is the last β€” is recorded rather than hidden:

/* Did it finish? A truncated summary is the worst thing this can produce β€”
   it reads almost like a complete one, and the section it lost is usually the
   last, which is where "waiting on" lives. Stored as a fact so the panel can
   say so, rather than being dropped (half a summary is still worth reading if
   you know that is what it is). */
$wasCut = in_array($result['finish_reason'] ?? '', ['length', 'max_tokens'], true) ? 1 : 0;

The real fix is not a bigger budget β€” it is System β†’ AI thinking, which turns extended thinking off per feature. Measured on a two-message ticket: 54.7s with thinking on, 6.9s with it off, and the fast answer was the better one.


Security

Both endpoints gate on tenancy before anything else:

// Multi-tenancy: a summary of a ticket you cannot open would be a novel way
// to read one.
if (!analystCanAccessTicket($conn, $analystId, $ticketId)) {
    echo json_encode(['success' => false, 'error' => 'Ticket not found']);
    exit;
}

'Ticket not found' for both missing and forbidden β€” the same wording, so the response does not confirm a ticket exists in a company you cannot see.

kind never reaches SQL from a request:

// Never interpolated from a request: an unknown kind falls back rather than
// reaching the query.
if (!in_array($kind, TICKET_AI_KINDS, true)) $kind = 'summary';

The AI namespace

Both use TICKET_AI_NS = 'tickets_reply_cleanup' β€” the same one the reply cleanup and merge summary already use, so an administrator configures "the AI that reads tickets" once rather than pasting a key into a third panel.

Testing it

Use the messy fixture, not a clean database β€” see the companion guide. Then check, in order:

  1. With both settings off: ai_summary.php returns {"disabled":true} and ai_read.php returns {"error":"disabled"} β€” no provider call.
  2. auto: true when not due: not_due; with auto_after = 0: auto_disabled. Both before any spend.
  3. A real refresh writes version + 1 and leaves the previous row alone.
  4. Reopen a briefing: served from the table, sub-second, and the network tab shows no provider traffic.
  5. The panel shows the datestamp, the message count, the version, and the "n messages have arrived since" line when the ticket has moved on.

Related pages

FreeITSM

Getting Started

Modules

Multi-tenancy (planned)

Blue sky thinking

Bugs resolved

Links

Clone this wiki locally