Repository navigation
DATRASextra v0.5.0
New features
-
Data extracted from ICES DATRAS now carries a record of where it came from, retrievable with the new
extraction()function. The record has one row per survey, year and quarter and reports the ICES calculation date, the extraction date, the source and endpoint, the archive file and its checksum, and the versions of DATRASextra, DATRAS, icesDatras and R that produced the object. It follows the data through the processing pipeline and is reconciled with the records still present, so subsetting narrows it instead of leaving stale entries behind.The central field is
DateofCalculation, supplied by ICES, which records when a block of records was last recalculated. Because ICES revises historical data as well as appending to it, this is the only reliable way to tell that an analysis will no longer reproduce against the current database. It is reconstructed from the data themselves when no explicit record is available, soextraction()also works on archives and objects created before this release. -
New
write_manifest(),read_manifest()andverify_extraction()describe and check a local archive.write_manifest()records a checksum and the ICES calculation date for every survey-year-quarter in a directory of exchange files, andverify_extraction()re-reads the archive and reports each entry asok,changed(contents differ locally),revised(ICES recalculated the data upstream),missingornew. Two users can compare manifests to confirm they hold identical data without hosting anything.download_datras()maintains the manifest automatically, and it can be built for any directory of exchange files regardless of how it was produced. -
write_datras()now returns the path withpayload_hash,zip_hashandalgoattributes. The payload hash is taken over the exchange file inside the archive, so it is identical whenever the data are identical; the archive hash is not, becauseutils::zip()stores the modification time of the file it compresses. -
New
reference_tables()reports the lookup tables bundled with the package -species_info,survey_info,survey_info_full_raw,spawning_infoand the internal ICES area and gear-spread tables - together with when each was generated, which script generated it, what it was generated from, and a hash that shows whether it still matches the version recorded in the package registry (inst/reference_tables.dcf). These tables are snapshots of external sources such as WoRMS and the ICES web services, so an analysis can depend on how old they are; this makes that visible. Maintainers regenerate the registry withDATRASextra:::.write_reference_registry()after rebuilding any table indata-raw/, and a test fails if the two drift apart. -
The vignette
vignette("data-processing-and-qc")gains a section on recording and verifying an extraction, and on the age of the bundled reference tables. -
New vignette
vignette("data-processing-and-qc")documenting every processing and quality-control step from download to analysis-ready object. It covers whatDATRAS::getDatrasExchange()andDATRAS::readICES()do before any DATRASextra function is called (column renaming, the-9missing value sentinel, matching of orphanCArecords, and the derivedhaul.id,LngtCm,SpeciesandCountcolumns), enumerates the filters applied byclean_datras()and how to change or replace them, documents the rule-based and percentile checks ofcheck_outliers()together with the diagnostic attributes it attaches, and lists further checks left to the user. -
read_datras()gains astrictargument, passed through to the underlying DATRAS reader. It controls howCArecords without a haul identifier are matched back to a haul: the defaultstrict = TRUEleaves records with several candidate hauls unmatched, whilestrict = FALSEassigns records with several candidate hauls to one of them at random. This changes the default behaviour ofread_datras().download_datras()takes the same argument and passes it on whenreturn_data = TRUE; it has no effect on the files written to disk, which hold the exchange data as delivered by ICES.
Bug fixes
-
download_datras()returned data for the wrong years when several surveys were downloaded in one call without specifyingyears. The per-survey year list overwrote theyearsargument inside the download loop, so the archive was read back filtered to the year coverage of whichever survey came last:download_datras(surveys = c("NS-IBTS", "BITS"))returned no NS-IBTS data from before the first BITS year. Only the returned object was affected; the files written to disk were always complete, and re-reading such an archive withread_datras()gives the full data. -
download_datras()failed for every survey and year withError in Year + (Month - 1) * 1/12 : non-numeric argument to binary operator. The ICES DATRAS field list declaresYearandTimeShotas character fields, and icesDatras 1.5.2 (released 2026-06-25) started applying that schema to downloaded data by default, so the arithmetic inDATRAS:::addExtraVariables()was handed character vectors. Reading archived exchange files was never affected, because those are parsed from CSV. The fix is in DATRAS, which now coerces the fields that must be numeric, so DATRASextra requires DATRAS >= 1.01.2. Users on an older DATRAS can work around it withoptions(icesDatras.fix_types = FALSE). -
check_outliers(action = "remove")discarded every attribute of the object it returned, because the removal step rebuilt the object withlapply()and restored only the class. Objects lostcm.breaks,swept_area_summaryandswept_area_unit, so a subsequent call to a function depending on them failed with a message about a missing spectrum. Attributes are now preserved. -
c()on twodatras_rawobjects discarded the attributes of both. This was invisible before extraction records existed, but it is the point at which survey-year files read separately are combined, so it now merges their records instead. -
clean_datras(impute_missing_depth = TRUE)failed for every input withinvalid type (list) for variable 'mgcv::s(lon, lat, k = 200)'.mgcvidentifies smooth terms by matching the bare symbols, so the namespace-qualified call in the model formula was never recognised as a smooth. The formula is now built in an environment that providess, and the basis dimension is capped at the number of unique haul positions so that imputation also works for small objects.
Minor changes
-
prune_datras()now retainsDateofCalculationinHH. It was previously dropped, which removed the only indication of when ICES last recalculated the records. This adds one integer column per haul. -
The default discrete colour palette used by the plotting functions is now sampled at equally spaced CIE L* lightness along the package colour ramp instead of taking the first
nanchor colours. Previously two or three groups were assigned neighbouring colours from the light end of the ramp (sand, algae green, sea green), which separated poorly on screen and in print, and were close to indistinguishable under red-green colour vision deficiency. The new selection spans the full ramp, keeps the light-to-dark ordering, and increases the smallest pairwise colour distance for two groups by roughly a factor of four. Plots that rely on the default palette will change appearance; passingcolexplicitly reproduces any previous colours. -
Plots with a single group now use the teal anchor as the default colour rather than the palest anchor of the ramp, which had very little contrast against a white background. This affects
plot_stratified_index(),plot_spatial_indicators()andplot_length_distribution(). -
calc_stratified_index()andcalc_spatial_indicators()now return grouping columns with their natural type.HH$Yearis stored as a factor, which was carried into the result tables, soYearcame back as a factor or a character string. Columns whose values are all numeric are converted back to numeric; genuinely categorical groups such asSurveyare unchanged. -
plot_spatial_indicators()draws connected lines and a continuous axis whenever the x variable is numeric-valued, including years stored as a factor or character. Previously such an x variable took the categorical branch, which drew unconnected points and one tick mark per level. Non-numeric x variables are still drawn as categories. -
plot_stratified_index()gainsy_scale. By default ("auto") the index is divided by a power of 1000 chosen from the data and the factor is stated in the y-axis label, for exampleIndex (per km^2) [10^9], so that wide tick labels no longer overlap the axis label. Usey_scale = "none"for the previous behaviour or pass an exponent directly.ylimis still given in the original units.