ARCHIVED

The tarproc utilities were a set of utilities written in Python for transforming and processing tar files and respecting sample boundaries within tar files representing training datasets. They have been superceded by a

Status

The Tarproc Utilities

Tarfiles are commonly used for storing large amounts of data in an efficient, sequential access, compressed file format, in particualr for deep learning applications. For processing and data transformation, people usually unpack them, operate over the files, and tar up the result again.

This library and set of utilities permits operating directly on tar files. This is faster than operating on files on file systems, and it is usually easier too.

tarcats -- concatenate tar files sequentially
tarsplit -- split a tar file by number of records or size
tarpcat -- concatenate tar files in parallel
tarproc -- map command line programs over tar files
tarshow -- show contents of tar files
tarsort -- sort tar files based on some key

The following are less commonly used utilities that are specifically useful for deep learning:

tarfirst -- extract the first file matching some criteria
targrep -- grep through files inside tar files (this will replace tarfirst)
tar2db, tar2lmdb, tar2tsv -- convert tar files to database files
tarmix -- mix tar files based on statistical sampling
tsv2tar -- build tar files based on a .tsv file plan

The utilities allow operating on stdin/stdout when necessary, allowing command line pipes to be constructed. For example:

    $ gsutil cat gs://bucket/file.tar | tarsort | tarsplit -o output

Python Interface

from tarproclib import reader, gopen
from itertools import islice

gopen.handlers["gs"] = "gsutil cat '{}'"

for sample in islice(reader.TarIterator("gs://lpr-imagenet/imagenet_train-0000.tgz"), 0, 10):
    print(sample.keys())

dict_keys(['__key__', 'cls', 'jpg', 'json', '__source__'])
dict_keys(['__key__', 'cls', 'jpg', 'json', '__source__'])
dict_keys(['__key__', 'cls', 'jpg', 'json', '__source__'])
dict_keys(['__key__', 'cls', 'jpg', 'json', '__source__'])
dict_keys(['__key__', 'cls', 'jpg', 'json', '__source__'])
dict_keys(['__key__', 'cls', 'jpg', 'json', '__source__'])
dict_keys(['__key__', 'cls', 'jpg', 'json', '__source__'])
dict_keys(['__key__', 'cls', 'jpg', 'json', '__source__'])
dict_keys(['__key__', 'cls', 'jpg', 'json', '__source__'])
dict_keys(['__key__', 'cls', 'jpg', 'json', '__source__'])

TODO

cleanup
- organize commands under top level
- use entrypoints/console_scripts in setup.py
tarmix
- implement convert and rename
tarshuffle
- implement stream shuffling with large on-disk buffer
add argo examples

Name		Name	Last commit message	Last commit date
Latest commit History 139 Commits
.githooks		.githooks
.github/workflows		.github/workflows
docs		docs
old		old
tarproclib		tarproclib
test		test
testdata		testdata
.deepsource.toml		.deepsource.toml
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
VERSION		VERSION
lines2tar		lines2tar
mkdocs.yml		mkdocs.yml
requirements.dev.txt		requirements.dev.txt
requirements.txt		requirements.txt
setup.py		setup.py
tar2json		tar2json
tarcats		tarcats
tarpcat		tarpcat
tarproc		tarproc
tarshow		tarshow
tarsort		tarsort
tarsplit		tarsplit
tasks.py		tasks.py

License

tmbdev-archive/tarproc

Folders and files

Latest commit

History

Repository files navigation

ARCHIVED

Status

The Tarproc Utilities

Python Interface

TODO

About

Resources

License

Stars

Watchers

Forks

Languages