Releases: haanvid/CL1AD
Release list
Compositionality Level 1 Action Dataset
Compositionality Level 1 Action Dataset (CL1AD) consists of 900 videos that are made by 10 subjects performing 9 actions directed at 4 objects. There are 15 object-directed action categories in total (Fig. 1). The dataset was designed so that the categorization task was non-trivial. A non-action-directed-object (non-ADO) appears along with an ADO in each video to prevent the model from inferring a human action or an ADO solely by recognizing an object in a video (Fig. 2). For each object-directed action category, 6 videos were shot for each subject with three different non-ADOs appearing in two videos each in different states (opened or closed) if possible, as they are presented in the other videos as ADOs. During the recording of the dataset, the subjects generated each action without constraints. Objects were located in random positions in the task space. However, the camera view angle was fixed, since the problem of view invariance is beyond the scope of the current study. The dataset is open for a public use.
Fig. 1. Composition of CL1AD. 15 classes were made by combination of 4 objects and 9 actions.
Fig. 2. Sample frames from CL1AD. An object that is not related to the action also appears in the scene to make the recognition task non-trivial.
Summary
- Labels: action and action-directed object pairs.
- Action-directed object categories: box, laptop, bottle, cup.
- Action categories: sweep, open, close, drink, change-page, type, shake, stir, blow.
- Number of classes: 15 classes.
- Number of videos taken per class (trials): 6 videos.
- Number of videos per subject: 90 videos.
- Number of subjects: 10 subjects.
- Total number of videos: 900 videos.
Reference
@online{1602.01921,
Author = {Haanvid Lee, Minju Jung, and Jun Tani},
Title = {Recognition of Visually Perceived Compositional Human Actions by Multiple Spatio-Temporal Scales Recurrent Neural Networks},
Year = {2016},
Eprint = {1602.01921},
Eprinttype = {arXiv},
}
Click "CL1AD.zip" to download the dataset.

