textreuse 1.0.0
textreuse 1.0.0
This release combines several years of maintenance and feature work into a single 1.0.0 release for CRAN preparation. It improves text input, corpus construction, local alignment behavior, candidate generation, LSH workflows, matrix export, performance, and documentation reliability.
Text input and corpus construction
TextReuseTextDocument()andTextReuseCorpus()now accept anencodingargument, making it easier to read source files whose text encoding is known or differs from the platform default.TextReuseCorpus()now keeps skipped-document bookkeeping deterministic. Skipped documents are reported consistently, and skip metadata is available even whenskip_short = FALSE.- Very short documents are handled more predictably when skip n-grams are used, avoiding assertion failures and making corpus construction easier to diagnose.
Alignment and match inspection
align_local()now returns an empty local alignment instead of throwing an error when two texts have no matching words. This makes batch alignment workflows easier to run because no-match pairs can be represented directly.align_local()gainspreserve_punctuation, allowing displayed alignments to keep punctuation from the original texts when that context is useful.- New
count_matches()andmatching_tokens()helpers expose absolute match counts and the matched tokens themselves, so users can inspect what drove a similarity score rather than relying only on a ratio.
Candidate generation and comparison
- New token-index helpers find candidate document pairs from shared n-grams, giving users another way to identify likely reuse pairs before running more expensive comparisons.
pairwise_candidates()and matrix conversion now preserve all document IDs, including documents without returned candidate pairs.as_sparse_matrix()provides a sparse matrix representation of candidate results, which is more convenient for downstream modeling, graph analysis, and workflows with many documents.
Locality-sensitive hashing and performance
lsh_add()can add new documents to an existing LSH bucket cache, so users can extend an index without rebuilding it from scratch.lsh_compare()can run comparisons in parallel on non-Windows platforms whenoptions(mc.cores)is set.- Long-running C++ hashing and n-gram loops now check for user interrupts, so expensive jobs can be stopped more cleanly from R.
Compatibility, documentation, and CI
- Compatibility with current dplyr and tidyr releases has been refreshed.
- README, vignette, reference, and pkgdown examples were regenerated against current package output.
- Stale external links and documentation badges were updated so package checks and the public documentation site are cleaner.
- pkgdown site font sizing is unified across pages.
- AppVeyor is configured to run the R build path instead of MSBuild project detection.
Verification
- pkgdown site build with
run_dont_run = TRUEcompleted successfully. R CMD check --no-manual textreuse_1.0.0.tar.gz: OK.R CMD check --as-cran --no-manual textreuse_1.0.0.tar.gz: 3 NOTEs, documented incran-comments.md.