COMPMID-477 - Optimized Direct Convolution 3x3 and 5x5 (f32) for Bifrost.

Each work-item computes 4x3 output elements in case of 3x3 convolution and 4x2 in case of 5x5 convolution Change-Id: I6ebbaff8b7e971c1f90d5845c0b58d2a40f39df5 Reviewed-on: http://mpd-gerrit.cambridge.arm.com/84345 Reviewed-by: Anthony Barbier <anthony.barbier@arm.com> Tested-by: Kaizen <jeremy.johnson+kaizengerrit@arm.com>
author: Gian Marco Iodice <gianmarco.iodice@arm.com> 2017-08-16 18:38:32 +0100
committer: Anthony Barbier <anthony.barbier@arm.com> 2018-11-02 16:35:24 +0000
commit: 1246b63ca04cb067f26ae860688647224d6ba24e (patch)
tree: a805ff1faa1de1b06f3569926ec2f09b63ecdb5f /src/runtime/CL/functions/CLConvolutionLayer.cpp
parent: f583fb74ecaad9d57672e1422d68566a78bb503e (diff)
download: ComputeLibrary-1246b63ca04cb067f26ae860688647224d6ba24e.tar.gz
1 files changed, 3 insertions, 0 deletions
diff --git a/src/runtime/CL/functions/CLConvolutionLayer.cpp b/src/runtime/CL/functions/CLConvolutionLayer.cpp
index ff94e9d7a2..b1b83985d0 100644
--- a/src/runtime/CL/functions/CLConvolutionLayer.cpp
+++ b/src/runtime/CL/functions/CLConvolutionLayer.cpp
@@ -113,6 +113,9 @@ void CLConvolutionLayer::configure(const ICLTensor *input, const ICLTensor *weig
     const DataType dt                   = input->info()->data_type();
     const int      fixed_point_position = input->info()->fixed_point_position();
 
+    // Set the GPU target for matrix multiply
+    _mm_kernel.set_target(CLScheduler::get().target());
+
     _has_bias             = (biases != nullptr);
     _are_weights_reshaped = weights_info.are_reshaped();
author	Gian Marco Iodice <gianmarco.iodice@arm.com>	2017-08-16 18:38:32 +0100
committer	Anthony Barbier <anthony.barbier@arm.com>	2018-11-02 16:35:24 +0000
commit	1246b63ca04cb067f26ae860688647224d6ba24e (patch)
tree	a805ff1faa1de1b06f3569926ec2f09b63ecdb5f /src/runtime/CL/functions/CLConvolutionLayer.cpp
parent	f583fb74ecaad9d57672e1422d68566a78bb503e (diff)
download	ComputeLibrary-1246b63ca04cb067f26ae860688647224d6ba24e.tar.gz