mozga-intel edited a comment on issue #19265:
URL: 
https://github.com/apache/incubator-mxnet/issues/19265#issuecomment-707744679


   Hi, 
   
   Well, 
   When running our pre-model (this is a simple imitation of the LSTM model). 
While a test, I want to create a large LRN tensor, for example: (20758,500). It 
could be seen that ~ 170GB of memory is allocated for scratchpad computations. 
We can see that global memory is always true. Well, as a result, for a 
different oneDNN version, I got the error messages: as following:
   1. For a given v.1.3 version of mkldnn: `Segmentation fault: 11`
   2. For a given v.1.6 version of mkldnn:  `mxnet.base.MXNetError: MXNetError: 
could not create a primitive`
   
   This error is only visible for a large LRN tensor. Step-by-step reproduction 
gives many lights to this issue. If we look into the code, a lot of future will 
be visible. The standard Vanilla-LSTM algorithm of MKLDNN leads to allocate of 
blocks of memory based on this equation: ```sizeof(float) * work_space```, 
where ```work_space``` is an offset (in bytes). For a given test (input 
:20758,500) we can see that ~170 GB od memory is allocated for scratchpad 
computation, where `workspace = 47952392192 * sizeof(float) =  191809568768 
bytes ~ 170 GB`.  If you don't have enough space, you will get both errors: see 
**1** & **2**. In Intel, MKLDNN primitives can use either individual memory or 
global buffer memory for an intermediate computation. The first one might lead 
to getting better performance result since memory most likely will be attached 
to any thread. The second one, might save a lot of memory.
   
   **For brevity:** 
   The input tensor is` T x N x C`, well, for a given example `(10758, 500)`, T 
is `10758`, C is `500`, That means that we need at least `4 * 10758 * 500 * 500 
* 4 bytes ~ 40 GB`, or maybe more. Basically the work-space would be comparable 
with the grid size `n_layers * mb * n_times_stamps * 4 (gates) * max(sic, slc, 
dhsc) ^ 2`.  For a given oneDNN version (1.3 and 1.6) the size of work-space 
(i.e LSTM space) is equal `book<float>(num_elems, ....) ~ 40 GB * sizeof(T) = 
40 GB * 4 ~160GB`.  The upper_bound (the size of input tensor) is unlimited and 
is bounded by available memory space.  Well, the size of buffer which is need 
to allocate LSTM tensor is equal: `4 * 10758 * 500 * 500 * 4 bytes ~ 40 GB`.  
Yet, this value is multiply by the constant value of its type (in this case: 
<T> = float). 
   Approximately: it should be defined, as following: 
   1. The size of work-space * <T>, where <T> is <uint8_t> ~ * 1byte 
[potentially]
   1. The workspace is only limited by the total number of elements of a given 
tensor. 
   
   
   The total size which is required for a given tensor (the upper-bound of LSTM)
   `n^2 * m = memory_space / (16 bits)`


----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

For queries about this service, please contact Infrastructure at:
[email protected]



---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to