Anndrey24 opened a new pull request, #16951:
URL: https://github.com/apache/tvm/pull/16951

   This patch partly reverts the unification of scalable and non-scalable 
scheduling of conv2d NHWC for `arm_cpu` targets introduced in #16899.
   
   The non-scalable schedule for float32 splits the N axis (corresponding to 
number of output channels) by 16 in both the unified and the nonunified 
schedule versions, and then additionally splits the inner partitions by 4 in 
only the nonunified version to which this patch is reverting (first added in 
#16106). The two versions' behaviour would be equivalent if none of the padding 
on the N axis was removed during lowering, however we allow for that to happen 
as it proved to increase performance for very small convolutions.
   
   As it stands, there seems to be a regression in cases where the datatype is 
float32 and the number of output channels is greater than 16, a multiple of 4, 
and not a multiple of 16, because even with the removed padding the nonunified 
schedule is able to vectorise over 4 elements, while the unified version cannot 
vectorise over 16 elements anymore.
   
   Since all of the conv2d NHWC hybrid topi test cases used numbers of output 
channels either less than 16 or divisible by 16, this patch also adds a new 
case which falls in the aforementioned regression area.
   
   cc @lhutton1 @ekalda 


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to