This is a Experimental version of OpenCL by AMD Research, we now recommend you to use The official BVLC Caffe OpenCL branch is over at Caffe branch now at
C++ Python C CMake Makefile Protocol Buffer Other
Switch branches/tags
Nothing to show
Clone or download
Latest commit 41805c8 Jul 24, 2016
Failed to load latest commit information.
cmake add find path for AMDAPPSDK3.0 and addes src/caffe/CMakeLists.txt Sep 14, 2015
data [example] image classification web demo Jul 12, 2014
docs Fix path to mnist_autoencoder.prototxt Jul 23, 2015
examples Port the softmax layer Jul 31, 2015
include/caffe Switch to using batched image scheme as default. previously was using… Oct 6, 2015
matlab Update ilsvrc_2012_mean.mat to W x H x C, update demo and add comments May 30, 2015
models This is a test layer Aug 25, 2015
python Travis scripts for python3 and pytest for cmake. Also fixes CUDA CMak… Jul 21, 2015
scripts Travis scripts for python3 and pytest for cmake. Also fixes CUDA CMak… Jul 21, 2015
src update some uncomment Sep 20, 2015
tools Add the change in tools/ Sep 14, 2015
.Doxyfile update doxygen config to stop warnings Sep 3, 2014
.gitignore update gitignore Sep 16, 2015
.travis.yml Travis scripts for python3 and pytest for cmake. Also fixes CUDA CMak… Jul 21, 2015
CMakeLists.txt Travis scripts for python3 and pytest for cmake. Also fixes CUDA CMak… Jul 21, 2015 clarify the license and copyright terms of the project Aug 7, 2014 replace bundled install instructions with link to site Feb 10, 2014
LICENSE update Readme and License file Sep 9, 2015
Makefile remove all cuda related flags in Makefile Aug 27, 2015
Makefile.config Removed unused variable in base_conv_layer Sep 19, 2015
Makefile.config.example Add commented out helpers for homebrew users Apr 2, 2015 Update Jul 24, 2016
caffe.cloc [fix] stop cloc complaint about cu type Sep 4, 2014

#This was experimental branch of Caffe for OpenCL, we know recommend you use the now official OpenCL port of Caffe in BVLC GitHub Repo at

###OpenCL Caffe Experimental branch by AMD Reserach- No new development is happing on it.

This is an OpenCL implementation of Caffe, a mainstream DNN framework ( It includes a largely complete Caffe feature set as of August 2015. The project is under active development to improve performance and add new features. Contributions from the community are welcome.

OpenCL ( is an open standard parallel programming language for heterogeneous platforms. OpenCL is supported by a variety of commercial chip manufacturers.

####Branches We have three branches in this repo.

-stable, the stable branch for users

-dev, the developer branch, we encourage people to contribute on this branch

-master, the original Caffe's master branch against which our code is synchronized.

####Design features -All Caffe layers ported to OpenCL

-Performance improvement by batched implementation for conv layer based on clBLAS

-The user can choose the optimal batch number depending on H/W properties, image size and minibatch size

-Supports OpenCL 2.0, 1.2

-Implemented in C++ and OpenCL, maintaining the same interfaces as the original Caffe

-Users can directly run DNN models: AlexNet, VGG-16 and VGG-19

Note: More features are planned in the near future. Currently this implementation has been verified and tuned on AMD devices (CPUs/GPUs/APUs). Compatibility across different chip manufacturers will be considered for future addition.


We intend to keep updating the latest performance as we make optimizations. Fury results are preliminary and are actively being improved.

  • Training speed (Model: AlexNet, minibatch size 128)
Platform Speed (images per second)
AMD W9100 & A10-7850k 255
AMD R9 Fury & A10-7850k 261
AMD R290X @1000MHz & A10-7850k 268
AMD S9150 @900MHz & Xeon E5-2640 227
  • Recognition speed (Model: AlexNet, minibatch size 128)
Platform Speed (images per second)
AMD W9100 & A10-7850k 590
AMD R9 Fury & A10-7850k 699
AMD R290X @1000MHz & A10-7850k 606
AMD S9150 @900MHz & Xeon E5-2640 452

####Wiki For more information on how to install, use or contribute to this code base, please visit our wiki page:

#Contributors Junli Gu, Yibing Liu, Yuan Gao, Maohua Zhu

We thank Mauricio Breternitz, Hanjin Chu and Greg Stoner for their technical suggestions and support.

If you have any questions, please send an email to

###Support needed As an open source project, we hope to maintain an open dynamics and sharing culture. We encourage the contribution and support from the community to improve it together.

###License The original Caffe is provided in the BSD 2-Clause license open source license. The OpenCL ports written by AMD is covered by AMD license. We encourage the contribution and support from external, your contribution will be covered either by BSD 2-Clause license or whichever your preferred license.

Original Caffe information


Caffe is a deep learning framework made with expression, speed, and modularity in mind. It is developed by the Berkeley Vision and Learning Center (BVLC) and community contributors.

Check out the project site for all the details like

and step-by-step examples.

Join the chat at

Please join the caffe-users group or gitter chat to ask questions and talk about methods and models. Framework development discussions and thorough bug reports are collected on Issues.

Happy brewing!

License and Citation

Caffe is released under the BSD 2-Clause license. The BVLC reference models are released for unrestricted use.

Please cite Caffe in your publications if it helps your research:

  Author = {Jia, Yangqing and Shelhamer, Evan and Donahue, Jeff and Karayev, Sergey and Long, Jonathan and Girshick, Ross and Guadarrama, Sergio and Darrell, Trevor},
  Journal = {arXiv preprint arXiv:1408.5093},
  Title = {Caffe: Convolutional Architecture for Fast Feature Embedding},
  Year = {2014}