-
-
Notifications
You must be signed in to change notification settings - Fork 0
FAQ
No. Rules-only mode is the default and is fully offline.
Not in the default metadata exposure. It sees column names, types, stats, and
pattern sketches. llm_exposure="sample" also sends up to 5 example values per
column: emails and phones are replaced and long strings become patterns, but
short category/text values go verbatim. llm_exposure="none" makes no network
call at all.
By design. Missingness is reported; fill_na is never auto-proposed. Add it to
the recipe yourself if imputation is appropriate.
The source column was missing from the frame. Non-strict modes warn and continue;
strict raises. Prefer suggest_update / re-plan when schema drifts.
pandas inferred the column as numeric while reading the file, so 01234
became 1234 before CleanFrame saw it — which is why no diff shows the change.
The same coercion turns literal NA/None text into a missing value and
rewrites 1e5. clean/report re-read a bounded verbatim slice and warn naming
the column, e.g.:
⚠ pandas type inference changed values while reading (Zip: '01234' lost its
leading zero(s) and became 1234). Pass text=True to read every field verbatim.
Fix it by reading verbatim:
cleanframe clean data.csv --text --out-dir out/cf.clean("data.csv", text=True)text=True is recorded as text: true in the recipe's read: section, so
apply_recipe replays the same read. Use it for any identifier-like column —
ZIP codes, account numbers, SKUs, part numbers.
Every advisory uses the CleanFrameWarning category, so one filter covers them
all:
import warnings
import cleanframe as cf
warnings.simplefilter("ignore", cf.CleanFrameWarning)Use "error" instead of "ignore" to make advisories fatal in CI. Filtering by
category leaves unrelated Python warnings alone.
Yes — commit the recipe YAML and call apply_recipe (or the CLI apply
subcommand) in the task. Fail the run on DriftError. From a shell, branch on the
exit code: 3 is drift, 4 is validation failure.
pip install "cleanframe-engine[excel]"That covers .xlsx / .xlsm. A legacy .xls needs pip install xlrd; writing
.xls is refused — write .xlsx.
pip install "cleanframe-engine[parquet]"Multi-sheet workbooks now raise if no sheet is chosen. Use
cf.clean_workbook(path) or cleanframe clean file.xlsx to clean every tab
(one recipe + diff per sheet), or pass sheet= to pick one.
Yes — cf.stream_apply(recipe, in_path, out_path, chunksize=N) (CLI
apply FILE --recipe R --chunksize N) replays row-independent recipes
out-of-core. Global ops (dedup, fill_na mean/median/…) are refused; peak
memory is bounded by the chunk size, not file size.
Read-time format auto-correction detects the delimiter and encoding by default
and pins them into the recipe. Pass --no-correct (correct_format=False) to
disable; an ambiguous delimiter raises rather than guessing.
Yes — Jinja2 autoescape is on; covered by tests.
cast to int rounds floats (pandas nullable Int64). Prefer keeping amounts
as float, or round explicitly with the round op first.
GitHub Wiki — sources live in
wiki/ in this repository.