Fully Convolutional Networks for Semantic Segmentation by Jonathan Long*, Evan Shelhamer*, and Trevor Darrell. CVPR 2015 and PAMI 2016.
Switch branches/tags
Nothing to show
Clone or download
shelhamer data: prepared NYUDv2
link to our preparation of the dataset with converted depth in the format of Gupta et al. '13
Latest commit 1305c73 Apr 5, 2018
Permalink
Failed to load latest commit information.
data data: prepared NYUDv2 Apr 5, 2018
demo bundle demo image + label and save output Jan 23, 2018
ilsvrc-nets add note on ILSVRC nets, update paths for base net weights Sep 16, 2016
nyud-fcn32s-color-d make process title optional Jan 20, 2017
nyud-fcn32s-color-hha make process title optional Jan 20, 2017
nyud-fcn32s-color make process title optional Jan 20, 2017
nyud-fcn32s-hha make process title optional Jan 20, 2017
pascalcontext-fcn16s make process title optional Jan 20, 2017
pascalcontext-fcn32s make process title optional Jan 20, 2017
pascalcontext-fcn8s make process title optional Jan 20, 2017
siftflow-fcn16s make process title optional Jan 20, 2017
siftflow-fcn32s drop accidental file Jan 20, 2017
siftflow-fcn8s make process title optional Jan 20, 2017
voc-fcn-alexnet make process title optional Jan 20, 2017
voc-fcn16s make process title optional Jan 20, 2017
voc-fcn32s make process title optional Jan 20, 2017
voc-fcn8s-atonce make process title optional Jan 20, 2017
voc-fcn8s make process title optional Jan 20, 2017
README.md link to slides Feb 2, 2017
infer.py bundle demo image + label and save output Jan 23, 2018
nyud_layers.py add NYUD FCNs May 20, 2016
pascalcontext_layers.py add PASCAL-Context FCNs May 20, 2016
score.py site for fcn.berkeleyvision.org Apr 14, 2016
siftflow_layers.py add SIFT Flow FCNs May 20, 2016
surgery.py comment surgery helpers for clarity Sep 10, 2016
vis.py replace VOC helper with more general visualization utility Jan 23, 2018
voc_layers.py PASCAL VOC: include more data details, rename layers -> voc_layers May 19, 2016

README.md

Fully Convolutional Networks for Semantic Segmentation

This is the reference implementation of the models and code for the fully convolutional networks (FCNs) in the PAMI FCN and CVPR FCN papers:

Fully Convolutional Models for Semantic Segmentation
Evan Shelhamer*, Jonathan Long*, Trevor Darrell
PAMI 2016
arXiv:1605.06211

Fully Convolutional Models for Semantic Segmentation
Jonathan Long*, Evan Shelhamer*, Trevor Darrell
CVPR 2015
arXiv:1411.4038

Note that this is a work in progress and the final, reference version is coming soon. Please ask Caffe and FCN usage questions on the caffe-users mailing list.

Refer to these slides for a summary of the approach.

These models are compatible with BVLC/caffe:master. Compatibility has held since master@8c66fa5 with the merge of PRs #3613 and #3570. The code and models here are available under the same license as Caffe (BSD-2) and the Caffe-bundled models (that is, unrestricted use; see the BVLC model license).

PASCAL VOC models: trained online with high momentum for a ~5 point boost in mean intersection-over-union over the original models. These models are trained using extra data from Hariharan et al., but excluding SBD val. FCN-32s is fine-tuned from the ILSVRC-trained VGG-16 model, and the finer strides are then fine-tuned in turn. The "at-once" FCN-8s is fine-tuned from VGG-16 all-at-once by scaling the skip connections to better condition optimization.

  • FCN-32s PASCAL: single stream, 32 pixel prediction stride net, scoring 63.6 mIU on seg11valid
  • FCN-16s PASCAL: two stream, 16 pixel prediction stride net, scoring 65.0 mIU on seg11valid
  • FCN-8s PASCAL: three stream, 8 pixel prediction stride net, scoring 65.5 mIU on seg11valid and 67.2 mIU on seg12test
  • FCN-8s PASCAL at-once: all-at-once, three stream, 8 pixel prediction stride net, scoring 65.4 mIU on seg11valid

FCN-AlexNet PASCAL: AlexNet (CaffeNet) architecture, single stream, 32 pixel prediction stride net, scoring 48.0 mIU on seg11valid. Unlike the FCN-32/16/8s models, this network is trained with gradient accumulation, normalized loss, and standard momentum. (Note: when both FCN-32s/FCN-VGG16 and FCN-AlexNet are trained in this same way FCN-VGG16 is far better; see Table 1 of the paper.)

To reproduce the validation scores, use the seg11valid split defined by the paper in footnote 7. Since SBD train and PASCAL VOC 2011 segval intersect, we only evaluate on the non-intersecting set for validation purposes.

NYUDv2 models: trained online with high momentum on color, depth, and HHA features (from Gupta et al. https://github.com/s-gupta/rcnn-depth). These models demonstrate FCNs for multi-modal input.

SIFT Flow models: trained online with high momentum for joint semantic class and geometric class segmentation. These models demonstrate FCNs for multi-task output.

Note: in this release, the evaluation of the semantic classes is not quite right at the moment due to an issue with missing classes. This will be corrected soon. The evaluation of the geometric classes is fine.

PASCAL-Context models: trained online with high momentum on an object and scene labeling of PASCAL VOC.

Frequently Asked Questions

Is learning the interpolation necessary? In our original experiments the interpolation layers were initialized to bilinear kernels and then learned. In follow-up experiments, and this reference implementation, the bilinear kernels are fixed. There is no significant difference in accuracy in our experiments, and fixing these parameters gives a slight speed-up. Note that in our networks there is only one interpolation kernel per output class, and results may differ for higher-dimensional and non-linear interpolation, for which learning may help further.

Why pad the input?: The 100 pixel input padding guarantees that the network output can be aligned to the input for any input size in the given datasets, for instance PASCAL VOC. The alignment is handled automatically by net specification and the crop layer. It is possible, though less convenient, to calculate the exact offsets necessary and do away with this amount of padding.

Why are all the outputs/gradients/parameters zero?: This is almost universally due to not initializing the weights as needed. To reproduce our FCN training, or train your own FCNs, it is crucial to transplant the weights from the corresponding ILSVRC net such as VGG16. The included surgery.transplant() method can help with this.

What about FCN-GoogLeNet?: a reference FCN-GoogLeNet for PASCAL VOC is coming soon.