-
Notifications
You must be signed in to change notification settings - Fork 4
Tools Data
The src/tools/data directory holds the readers behind the data tool: three
streaming operations — inspect, select, validate — over CSV/TSV, JSON, and JSON
Lines. Every function reads the file in chunks, keeps memory bounded regardless
of input size, and returns either a JSON-serializable result or a typed
DataRefusal with a machine-readable reason. The header comment in
src/tools/data/index.ts states the contract: results are exact or sampled
(view flag), and the readers "refuse with a typed reason instead of guessing
at bytes it cannot read as the declared format."
The directory is five files. src/tools/data/index.ts holds the public entry
points (inspectData, selectData, validateData, detectDataFormat).
src/tools/data/shared.ts holds the substrate: the fatal UTF-8 decoder, the
text stream, the refusal types, and cell classification. src/tools/data/csv.ts
implements the RFC 4180 parser and its three operations. src/tools/data/json.ts
implements a hand-written streaming JSON push parser and its three operations.
src/tools/data/jsonl.ts splits lines and reuses the JSON parser per record.
The tool contract is one layer above, in src/tools/gateway/data-tool.ts
(createDataTool) and src/tools/gateway/data-surface.ts (dataToolSurface,
prepareDataArguments), which normalizes arguments, resolves paths against the
workspace, and wraps results in the observation envelope.
detectDataFormat(path, explicit?) in src/tools/data/index.ts resolves the
format of a data file in a fixed priority order. The explicit argument wins:
it is trimmed, lowercased, and checked against DATA_FORMATS (csv, tsv,
json, jsonl); an invalid explicit format returns null. Next comes the
extension through EXTENSION_FORMATS: .csv→csv, .tsv and .tab→tsv,
.json→json, .jsonl and .ndjson→jsonl. Finally a bounded content sniff of
the first SNIFF_BYTES (4096) bytes: a leading [ or { is JSON, except that
a second non-blank line opening a { or [ while the first line ends in }
reads as JSONL; a tab in the first line is TSV; a comma, semicolon, or pipe is
CSV. Null means the file is not recognizably one of the supported formats.
Each of the three entry points follows the same skeleton. They call
statDataFile(path) from src/tools/data/shared.ts, which stats the path and
refuses with not-a-file when it is not a regular file, returning { size } on
success. They then call detectDataFormat; a null result returns
unsupported(path, explicit), which builds an unsupported-format refusal
listing the supported formats and suggesting run_script for binary scientific
formats. They forward options into a format-specific reader inside a try, and
catch any thrown error into refusalFromError(error, path, signal), which maps
DataRefusalError through, an aborted signal to aborted, ENOENT to
not-found, EISDIR to not-a-file, and anything else to read-error.
inspectData forwards sampleRows and maxRows to all three formats, plus
delimiter and header for CSV/TSV. selectData forwards offset and limit
as a window, plus columns for CSV/TSV and pointer for JSON, and cross-checks
them: columns on a JSON file returns an invalid-argument refusal ("use
pointer"), and pointer on a JSONL file returns invalid-argument ("JSONL
selects records by offset and limit"). validateData forwards maxRows and
the CSV/TSV options, with no windowing. The signal option is forwarded to all
three so an abort surfaces as an aborted refusal.
A focused test in tests/contracts/data-csv.test.ts pins the resolution
priority directly against detectDataFormat: an explicit "json" on a .csv
file returns "json"; an extension beats a conflicting explicit "yaml"
returning null; a tab file reads as "tsv"; a comma file as "csv"; a
{"a": 1} text file as "json"; two such lines as "jsonl"; prose as null.
The same file pins the refusals through the real entry points: an invalid
UTF-8 byte (0xe9) in a file called latin1.csv refuses with invalid-utf8
and byteOffset 9 across all three operations; a NUL byte refuses with
binary and the offset; a missing file refuses with not-found; a directory
with a .csv name refuses with not-a-file; a notes.txt of prose refuses
with unsupported-format naming the supported list and run_script; an
explicit { format: "parquet" } refuses naming the bad format.
The streaming contract lives in src/tools/data/shared.ts.
streamTextChunks(path, options) is an async generator yielding { text, byteOffset } chunks, where byteOffset is the absolute byte offset of the
first byte of the text in the file. It creates a createReadStream with a
default high water mark of 64 KiB, and decodes every chunk through a
Utf8StreamDecoder. It performs two refusals inline. First, a binary sniff:
the first BINARY_SNIFF_BYTES (8192) bytes are scanned for a NUL byte, which
throws a binary refusal naming the absolute offset. Second, the decoder
itself throws invalid-utf8 with the byte offset at the first ill-formed
sequence. It strips one leading UTF-8 byte-order mark, records it in
stats.bom, and advances the offset. The finally block always destroys the
underlying stream, so a reader that stops early never reads the rest of the
file; an abort listener destroys the stream with an aborted refusal.
Utf8StreamDecoder.push(chunk) validates every complete UTF-8 sequence
against the Unicode 15 table 3-7, carrying an incomplete trailing sequence into
pending for the next chunk. It rejects invalid lead bytes, overlong
encodings, and continuation bytes outside 0x80–0xbf, throwing a
DataRefusalError with invalid-utf8 at the absolute offset. end() rejects
a dangling partial sequence.
The view-flag contract is encoded as DataViewFlags { exact, sampled, converted } with three constructors: exactView() (whole file verbatim),
sampledView() (scan stopped early, counts are lower bounds, rowCount is
null), and cutView() (scan complete but a returned value was cut to a budget,
marked with a $summary or $truncated placeholder). The doc comment in
shared.ts notes inspection samples are previews and exempt: a placeholder
inside a sample is noted, and exact still speaks for the counts.
Refusals are plain objects with ok: false, a reason from the
DataRefusalReason union, and a message; DataRefusalError wraps one to
carry it through a throw. refusal(reason, message, extra) builds the object,
isDataRefusal(value) narrows it, and refusalFromError converts whatever a
reader threw into the refusal it stands for.
Cell classification is shared by the CSV and JSONL inspectors.
classifyCell(text) returns one of empty, sentinel, integer, float,
boolean, date, string, using regexes for the numeric and date forms and
SENTINEL_TOKENS (NA, N/A, NaN, null, NULL, None, -) for
missing values. inferColumnType(histogram) returns the column type a
histogram of cell classes supports: only typed cells vote (empties and
sentinels are missing values), integers and floats together are float, and
any other mixture is mixed. numberLiteralPrecision(literal) judges whether
reading a number literal as an IEEE double would misstate its text, returning
null when exact or one of the PrecisionKind values unsafe-integer,
excess-digits, inexact, overflow, underflow, oversized-literal.
src/tools/data/csv.ts implements an RFC 4180 state machine as the
CsvParser class. The states are FieldStart, Unquoted, Quoted,
QuoteInQuoted, and Overflow. push(text) feeds one decoded chunk; a run
accumulator (runStart) makes chunk boundaries inside a field or a quoted
field transparent. end() finalizes a trailing record. Recovery is lenient
and reported through CsvIssue objects with a kind, 1-based record,
line, and message: a bare quote inside an unquoted field is literal text
(bare-quote), text after a closing quote joins the field (text-after-quote),
and an unterminated quote ends at end of input (unterminated-quote).
Memory is bounded whatever the input. A field is cut at CSV_MAX_FIELD_CHARS
(1 MiB) and a record at CSV_MAX_RECORD_CHARS (16 MiB); after a cap is hit the
parser enters the Overflow state and drops input until the record's next line
break, reporting field-too-large or record-too-large. A stray opening quote
therefore costs one oversized, reported record instead of swallowing the rest
of the file into one field. The doc comment on the class states this invariant
directly.
Delimiter detection is detectCsvDelimiter(sample, sampleComplete), which
parses the first records under each candidate in CSV_DELIMITER_CANDIDATES
(,, \t, ;, |), takes the modal field count, and scores by how many
records share it; a candidate that never produces two fields is excluded, ties
fall to candidate order (comma first), and a file that no candidate splits is a
one-column file read with the comma. scanCsv in the same file holds back the
first 64 KiB of text for detection when no delimiter was given, then feeds it
through the same parser so the detection sample is never parsed twice.
Header resolution is the HeaderResolver class. Under auto, the first record
waits for the second: headerLooksLikeNames returns true only when every cell
is a non-numeric, non-empty text name, so a one-record file under auto is one
data row. header: true forces the first record to be the header, and
header: false treats every record as data.
inspectCsv reports per-column facts through ColumnAccumulator and
accumulateCell: inferred type, empty and sentinel counts, cell-class
histogram, up to five sample values, min/max length, numeric range (integers
tracked exactly through BigInt, floats through doubles), and precision
issues. It also reports ragged rows, quoting issues, blank lines, and the line
ending (lf, crlf, cr, mixed, or none via lineEndingOf). Caps bound
the report: COLUMN_TRACK_CAP 10,000 columns, FIRST_ISSUES_CAP 10 issues,
SAMPLE_VALUES_CAP 5 values. Sample cells are capped by capText at
SAMPLE_TEXT_CAP (512) characters.
selectCsv returns a bounded window of rows as source text with an optional
column projection by header name or 0-based index; resolveProjection refuses
with unknown-column for a name not in the header and invalid-argument for a
negative index or a name with no header. The limit is clamped to 1–1000 with a
default of 50; hasMore is set by reading one row past the window.
validateCsv returns valid as true when every scanned row is well formed
and the scan reached the end, false when ragged rows or quoting issues were
found, and null when the scan stopped at maxRows without a fault.
A focused test in tests/contracts/data-csv.test.ts exercises these
directly. The parser tests feed 'a,b,c\r\n"x, y","he said ""hi""","line1\r\nline2"\r\n1,2,3\r\n'
and assert three records with no issues; they survive a chunk boundary at every
split point; and they recover from quoting faults, asserting the issues
text-after-quote@2:2, bare-quote@3:3, and unterminated-quote@4:4. The cap
tests cut a quoted field at CSV_MAX_FIELD_CHARS and close the record at the
next line break, and a 64 MB runaway-quote test streams in 64 KiB chunks and
asserts retained memory stays under 16 MB after a forced collection. The inspect
test pins sentinels counted but never converted: a NA cell reports
temperature?.sentinels as { NA: 1 } and the inferredType as float with
the sentinel excluded from the numeric range. The precision test asserts that
9007199254740993 is reported as an unsafe-integer with its literal text.
The select test asserts a verbatim window ["2", "bob", "NA"] with hasMore
true, and a limit clamp from 5000 to 1000. The format-detection test asserts
detectCsvDelimiter picks ; for a semicolon file, \t for a tab file, and
returns confidence 0 for a one-column file.
src/tools/data/json.ts contains a hand-written incremental JSON parser, the
JsonStreamParser class, with the stated purpose that a multi-gigabyte array
can be inspected or selected from without being loaded, and numbers keep their
source text so precision loss is reported instead of silently rounded;
JSON.parse never touches a whole data file.
The parser drives an event sink conforming to JsonEventSink, which receives
onStartObject, onKey, onEndObject, onStartArray, onEndArray,
onScalar, and captureStrings. A callback may return false to stop the
parse after that event; the parser then reports stopped: true and reads
nothing further. depth is the nesting level (root is 0). It tracks
line, column, and byteOffset for error locations and enforces
JSON_MAX_DEPTH (1024), throwing a nesting-too-deep refusal past it. String
keys are captured to KEY_CAPTURE_CAP (1024) and values to
STRING_CAPTURE_CAP (65,536); a string cut at the cap is marked truncated.
Number literals keep their source text in JsonScalarEvent.literal; a literal
longer than MAX_NUMBER_LITERAL_CHARS (512) is truncated, its value is
NaN, and the event is marked truncated. Syntax errors throw an
invalid-json refusal carrying line, column, and byte offset.
JsonValueBuilder materializes one subtree from events under a character
budget. Once the budget is spent, the subtree's root becomes a $summary
placeholder and the builder keeps counting so the summary reports the true
member count; numbers with a precision issue become $literal placeholders,
and strings cut at the capture cap become $truncated placeholders.
parseJsonPointer(pointer) parses an RFC 6901 pointer into segments with
~0/~1 escapes, with "" selecting the whole document. PathTracker and
StructureScanner track the pointer of the value about to start and the
duplicate-key and precision facts every consumer needs; duplicate-key tracking
caps at DUPLICATE_KEY_TRACK_CAP (10,000) keys per object, setting
checkedFully false past the cap.
inspectJson uses an InspectSink that reports the root type, an ordered key
list capped at KEY_LIST_CAP (200), an element- or member-type histogram, a
sample within SAMPLE_BUDGET_CHARS (8,192) per element, maxDepth, duplicate
keys, and precision. selectJson uses a SelectSink that seeks the
pointer and captures the target under SELECT_BUDGET_CHARS (1 MiB); in
windowed mode (offset/limit) over an array it reads one element past the
window to learn hasMore. It refuses with selection-too-large when the
target exceeds the budget, pointer-not-found naming the deepest existing
prefix when the pointer does not resolve, and invalid-argument when
offset/limit target a non-array. validateJson reports a syntax fault with
its location as syntaxError, and returns valid: null when the scan stopped
at maxRows before the end; duplicate keys are legal JSON and are reported,
not failed. scanJsonText parses one complete JSON text (a JSONL line) with a
TextScanSink, throwing a DataRefusalError for a syntax fault.
A focused test in tests/contracts/data-json.test.ts pins these. The parser
test replays a document as a flat event log pushed whole and one character at
a time, asserting identical logs. A table of 17 syntax faults each asserts an
invalid-json refusal with a specific message and a byte offset. The fault
location test asserts a tru literal on line 3 of a multibyte document
reports line 3 and the exact byte offset past the multibyte prefix. The
inspect precision test asserts that 9007199254740993 becomes a $literal
placeholder with precision: "unsafe-integer" while 9007199254740991 stays a
number, and that 1e400 is overflow and 1e-400 is underflow. The select
test asserts a pointer with escapes (/meta/a~1b/~0x) resolves, a windowed
array selection returns hasMore true and stops reading early, a missing
pointer (/items/9/id) refuses naming the deepest existing prefix /items,
and an oversized selection (selection-too-large) is refused with a message
suggesting an array window. The validate test asserts valid: null when
stopped early and valid: true for a document with duplicate keys.
src/tools/data/jsonl.ts reads JSON Lines, where each non-blank line is one
record scanned with the same push parser from json.ts, so precision facts,
duplicate keys, and syntax faults carry the line number they belong to. Memory
holds one line, the bounded sample, and per-key counters.
forEachLine is an async generator splitting decoded chunks into physical
lines, stripping one trailing CR, and carrying a partial line across chunks up
to JSONL_MAX_LINE_CHARS (1 MiB). Past the cap the line is dropped as it
streams and reported once its terminator arrives, so a whole JSON document
saved with a .jsonl name costs nothing to hold. The visitor callbacks
onLine, onOversized, and onBlank return false to stop.
inspectJsonl reports linesScanned (blank lines included) versus
rowsScanned (non-blank lines), blank-line count, invalid records with their
line and message, a root-type histogram, a top-level key presence histogram
capped at KEY_HISTOGRAM_CAP (500), per-key value types, maxDepth, duplicate
keys, precision issues, and a sample of valid records each within
JSON_SAMPLE_BUDGET_CHARS. selectJsonl returns a bounded window of records
each as { line, value, truncated } or { line, error }, using
SELECT_LINE_BUDGET_CHARS (1 MiB) per record. validateJsonl returns
valid: false with the first invalid line, null when stopped early, and
true for a clean file.
A focused test in tests/contracts/data-jsonl.test.ts pins these. A fixture
with eight lines (a blank line, a broken line, a record with a large integer
and a duplicate key, and a bare array) asserts rowsScanned 7, linesScanned
8, blankLines 1, and invalid naming line 5 with its message. The select
test asserts a window with offset: 2, limit: 3 returns the broken record as
{ line: 5, error: ... } in place and the large integer as a $literal
placeholder. The oversized-line test writes a 64 MB JSON array followed by two
records, asserts the array is skipped without being held (retained memory under
16 MB after a forced collection), and that the report names it with a message
suggesting format json. The select test also asserts pointer on a JSONL
file refuses with invalid-argument.
The data tool is composed in src/tools/gateway/data-tool.ts.
createDataTool(deps) returns a ToolSpec built from dataToolSurface. Its
run normalizes arguments through prepareDataArguments, validates the op
is one of inspect, select, validate, resolves the path against the
workspace via resolveToCwd, and reserves an observation envelope before any
reading: reserveObservation(DATA_OBSERVATION_SELF_CAP_BYTES, options), where
DATA_OBSERVATION_SELF_CAP_BYTES is 32 KiB from src/tools/gateway/caps.ts.
If the reservation is exhausted it returns observationBudgetExhausted. It
then commits the reservation up front (commitObservationReservation) because
the readers stream and await between reserve and finalize, so the cap is
charged up front and reconciled when the result settles. The result is
finalized through finalizeObservation with format: "json", an explicit unit
and counts from countResult (which reads returned, offset, limit, and
hasMore off the reader's own result), and a details object carrying op,
path, format, view, and viewLabel from describeDataView. A refusal is
mapped through refusalResult to an error message of the form data: {op} refused ({reason}): {message}, and a reservation is released on the refusal
and error paths.
prepareDataArguments in src/tools/gateway/data-surface.ts repairs
weak-model argument shapes: numbers as strings, columns as a comma-separated
string, header as "true"/"false", and file/file_path for path.
The tool is registered in src/tools/core-bootstrap.ts as a lazy tool: the
eager part is dataToolSurface (schema and policy), and the implementation is
loaded on first run through lazyTool with path: "src/tools/gateway/data-tool.ts"
and scope: "core". The same surface is reachable as a data capability on the
gateway tool: tests/contracts/data-tool.test.ts calls it as
{ tool: ToolNames.Gateway, args: { op: "call", capability: "data", args } }.
A focused test in tests/contracts/data-tool.test.ts pins the tool contract.
The CSV inspect test asserts the result's format, rowCount, and header,
and the details.viewLabel as "exact". The select test asserts a verbatim
window with projected columns and an observation next hint of offset=3.
The precision test asserts an unsafe integer is returned as a $literal
placeholder. The refusal test maps each reader refusal to an error that names
the op and the reason, and the argument-repair test asserts a file alias and
string numbers are repaired while a bad op is refused.
The change points this area invites:
-
A new data format is added by extending
DATA_FORMATSinsrc/tools/data/shared.ts,EXTENSION_FORMATSinsrc/tools/data/index.ts, the content-sniff branch indetectDataFormat, and theswitchin each ofinspectData,selectData, andvalidateData, plus a reader module with the sameinspect/select/validateentry points and result shape. -
A new refusal reason is added to the
DataRefusalReasonunion insrc/tools/data/shared.ts. -
A new tool argument is added to
dataToolSurface.parametersinsrc/tools/gateway/data-surface.tsand repaired inprepareDataArguments, then forwarded in therunofcreateDataTool. -
A new capability on the gateway reuses the same surface and registration
path; the
datacapability is reached viaop: "call"and is pinned bytests/contracts/data-tool.test.ts.
-
The chunk-boundary logic is load-bearing. Both
CsvParserandJsonStreamParseruse run accumulators (runStart,chunkIndex) so that a field or string split across two chunks behaves as if it were whole. The tests split text at every possible point; a change to how a run is appended must keepappendRunchecking both the field cap and the record cap (fieldRoomandrecordRoom). -
The
Overflowstate inCsvParsermust close the record at the next line break, not discard rows after it. The runaway-quote test asserts that rows after the fault still parse and that the recovery point is reached. -
streamTextChunksdestroys the stream infinally. A reader that stops early (a sink returningfalse, or amaxRowsbound) relies on this to not read the rest of the file; the JSONL line-splitting performance test exists because rescanning a chunk prefix for every line broke the bound. -
View flags describe the returned data.
exactView,sampledView, andcutViewencode whether counts are lower bounds and whether a value was cut to a budget. A change to a reader's stop condition must update itsviewso a sampled scan cannot pass for the complete dataset; thedescribeDataViewlabel indata-tool.tsdepends on the exact three-flag shape. -
Boundary rules.
src/tools/**never importssrc/interactive/**, and optional object fields are passed with the spread pattern (...(x !== undefined ? { x } : {})) becauseexactOptionalPropertyTypesis on. Relative imports end in.jsunder NodeNext. -
Test placement. A regression belongs in
tests/contracts/; a test intests/extended/**runs only underpnpm run test:full, never in CI, so it guards nothing. The memory-bound tests usev8.setFlagsFromString("--expose-gc")and a forced collection, so they are deterministic only under that protocol.
Source and generation metadata
title: "Tools data"
summary: "Streaming inspection, selection, and validation of CSV/TSV, JSON, and JSON Lines files, exposed through the `data` tool; results carry explicit view flags and typed refusals."
sources:
- "src/tools/data/index.ts"
- "src/tools/data/shared.ts"
- "src/tools/data/csv.ts"
- "src/tools/data/json.ts"
- "src/tools/data/jsonl.ts"
- "src/tools/gateway/data-tool.ts"
- "src/tools/gateway/data-surface.ts"
symbols:
- "inspectData"
- "selectData"
- "validateData"
- "detectDataFormat"
- "streamTextChunks"
- "Utf8StreamDecoder"
- "CsvParser"
- "JsonStreamParser"
- "createDataTool"
- "scanCsv"
- "forEachLine"
- "numberLiteralPrecision"
- "scanJsonText"
tests:
- "tests/contracts/data-csv.test.ts"
- "tests/contracts/data-json.test.ts"
- "tests/contracts/data-jsonl.test.ts"
- "tests/contracts/data-tool.test.ts"
invariants:
- "Nothing in the directory reads a data file whole; every result carries DataViewFlags {exact, sampled, converted} so a sampled scan cannot pass for the complete dataset."
- "Readers validate input bytes: invalid UTF-8 names its byte offset, a NUL byte within the first 8 KiB is binary, an unrecognized format names the supported ones."
- "Cells and numbers stay source text: inspection classifies them, selection returns them verbatim, and numbers a double would misstate are reported with a $literal placeholder and a precision kind."
validate:
- "pnpm run test:file -- tests/contracts/data-csv.test.ts"
- "pnpm run test:file -- tests/contracts/data-json.test.ts"
- "pnpm run test:file -- tests/contracts/data-jsonl.test.ts"Clio Coder · Repository · Website · Documentation
Wiki v0.1 · Developing implementation reference · Source snapshot: 657dce13d. Authored architecture documents define the product contracts.
- Clio Coder GUI Client
- apps / clio-coder-gui
- Apps clio coder gui server
- Apps clio coder gui tests
- apps
- Architecture
- Command-line surfaces
- Core
- Domains agents
- Config Domain
- Context Domain
- Dispatch domain
- Domains evidence
- Domains extensions
- Domains gateway
- domains
- Domains interop
- Domains lifecycle
- Domains memory
- Middleware Domain
- Domains mux
- Domains observability
- Domains plugins
- Prompt Compiler
- Domains providers
- Domains quota
- Domains resources
- Domains safety
- Domains scheduling
- Domains session
- Vendored Tool Registry and Resolution
- Engine
- Engine acp
- Engine apis
- engine
- Entry point
- Interactive
- interactive
- Interactive overlays
- Interactive renderers
- clio-coder wiki
- Scripts
- Contract tests
- Tests extended
- tests
- Tools
- Tools data
- tools
- Tools verify
- Worker runtime