# DMA stuck at wait()

**URL:** <https://discuss.pynq.io/t/dma-stuck-at-wait/8573>\
**Category:** Support\
**Created:** [June 25, 2025, 11:29am UTC](https://discuss.pynq.io/t/dma-stuck-at-wait/8573 "2025-06-25T11:29:16Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Karan\_Mali](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.pynq.io/karan_mali/32/6557_2.png) [@Karan\_Mali](https://discuss.pynq.io/u/Karan_Mali)\
**Post date:** [June 25, 2025, 11:29am UTC](https://discuss.pynq.io/t/dma-stuck-at-wait/8573/1 "2025-06-25T11:29:16Z")

</div>

Hi everyone! , I am using PYNQ image version 3.0.1 and Vivado 2022.2 and currently facing a DMA wait() error as I am trying to run a validation test and measure inference time for image processing application.  
Following are the schematics and all IP config settings. The dut\_0 IP you may be seeing is an HLS generated ip whose I/O interfaces are ap\_fifo through which I want to send an array of image pixel values via DMA using AXI streaming interface.  
[design\_\_bd.pdf](https://discuss.pynq.io/uploads/short-url/hNuESk8LhXBDFRwUPBRwuVkcpqG.pdf) (277.7 KB)

 ![Screenshot from 2025-06-25 16-11-18](https://us1.discourse-cdn.com/flex019/uploads/pynq1/original/2X/1/123c32b925e0ee50a0b084ecfcf86f5e027292fa.png)  
 ![Screenshot from 2025-06-25 16-16-29](https://us1.discourse-cdn.com/flex019/uploads/pynq1/original/2X/9/903f3374cbfee6515aaceea265ec3e5d16f6a49d.png)  
 ![Screenshot from 2025-06-25 16-16-41](https://us1.discourse-cdn.com/flex019/uploads/pynq1/original/2X/5/5de9bca0b7b0eb2d033350f98c58894b34120df8.png)  
 ![Screenshot from 2025-06-25 16-17-24](https://us1.discourse-cdn.com/flex019/uploads/pynq1/original/2X/4/47d72ebbbe8db71755a3af841f2dfaabc7684a05.png)  
 ![Screenshot from 2025-06-25 16-17-34](https://us1.discourse-cdn.com/flex019/uploads/pynq1/original/2X/3/382cf6d1c47aed1cec0ac2ebad03d8c44856f4b9.png)  
The error I am facing is when trying to run inference is as follows with the code

````auto
from datetime import datetime

import numpy as np
from pynq import Overlay, allocate

class HardwareOverlay(Overlay):
    def __init__ (
        self, bitfile_name, x_shape, y_shape, dtype=np.float32, dtbo=None, download=True, ignore_version=False, device=None
    ):
        super(). __init__ (bitfile_name, dtbo=None, download=True, ignore_version=False, device=None)
        self.sendchannel = self.axi_dma_0.sendchannel
        self.recvchannel = self.axi_dma_0.recvchannel
        self.input_buffer = allocate(shape=x_shape, dtype=dtype)
        self.output_buffer = allocate(shape=y_shape, dtype=dtype)

    def _print_dt(self, timea, timeb, N):
        dt = timeb - timea
        dts = dt.seconds + dt.microseconds * 10**-6
        rate = N / dts
        print(f"Classified {N} samples in {dts} seconds ({rate} inferences / s)")
        return dts, rate

    def inference(self, X, debug=True, profile=False, encode=None, decode=None):
        """
        Obtain the predictions of the NN implemented in the FPGA.
        Parameters:
        - X : the input vector. Should be numpy ndarray.
        - dtype : the data type of the elements of the input/output vectors.
                  Note: it should be set depending on the interface of the accelerator; if it uses 'float'
                  types for the 'data' AXI-Stream field, 'np.float32' dtype is the correct one to use.
                  Instead if it uses 'ap_fixed<A,B>', 'np.intA' is the correct one to use (note that A cannot
                  any integer value, but it can assume {..., 8, 16, 32, ...} values. Check `numpy`
                  doc for more info).
                  In this case the encoding/decoding has to be computed by the PS. For example for
                  'ap_fixed<16,6>' type the following 2 functions are the correct one to use for encode/decode
                  'float' -> 'ap_fixed<16,6>':
                  ```
                    def encode(xi):
                        return np.int16(round(xi * 2**10)) # note 2**10 = 2**(A-B)
                    def decode(yi):
                        return yi * 2**-10
                    encode_v = np.vectorize(encode) # to apply them element-wise
                    decode_v = np.vectorize(decode)
                  ```
        - profile : boolean. Set it to `True` to print the performance of the algorithm in term of `inference/s`.
        - encode/decode: function pointers. See `dtype` section for more information.
        - return: an output array based on `np.ndarray` with a shape equal to `y_shape` and a `dtype` equal to
                  the namesake parameter.
        """
        if profile:
            timea = datetime.now()
        if encode is not None:
            X = encode(X)
        self.input_buffer[:] = X
        self.sendchannel.transfer(self.input_buffer)
        self.recvchannel.transfer(self.output_buffer)
        if debug:
            print("Transfer OK")
        self.sendchannel.wait()
        if debug:
            print("Send OK")
        self.recvchannel.wait()
        if debug:
            print("Receive OK")
        # result = self.output_buffer.copy()
        if decode is not None:
            self.output_buffer = decode(self.output_buffer)

        if profile:
            timeb = datetime.now()
            dts, rate = self._print_dt(timea, timeb, len(X))
            return self.output_buffer, dts, rate
        else:
            return self.output_buffer

````

```auto
for data, targets in iter(test_loader):
    testdata = data.numpy()
    lables = targets.numpy()
    infer = HardwareOverlay('design_1.bit', testdata.shape, lables.shape)
    y_hw, latency , throughput = infer.inference(testdata, profile=True)
    break

```

here the test\_loader is a pytorch dataloader object  
error as follows (this is produced when I interrupt the kernel execution, if not interrupted then cell runs forever)  
**output stuck (no acknowledgement of completed transfer)**  
Transfer OK  
**error after interrupt**

```auto
KeyboardInterrupt Traceback (most recent call last)
Input In [4], in <cell line: 28>()
     30 lables = targets.numpy()
     31 nn = NeuralNetworkOverlay('snn.bit', testdata.shape, lables.shape)
---> 32 y_hw, latency , throughput = nn.predict(testdata, profile=True)
     33 break

Input In [1], in NeuralNetworkOverlay.predict(self, X, debug, profile, encode, decode)
     58 if debug:
     59 print("Transfer OK")
---> 60 self.sendchannel.wait()
     61 if debug:
     62 print("Send OK")

File /usr/local/share/pynq-venv/lib/python3.10/site-packages/pynq/lib/dma.py:181, in _SDMAChannel.wait(self)
    179 if error & 0x40:
    180 raise RuntimeError("DMA Decode Error (invalid address)")
--> 181 if self.idle:
    182 break
    183 if not self._flush_before:

File /usr/local/share/pynq-venv/lib/python3.10/site-packages/pynq/lib/dma.py:80, in _SDMAChannel.idle(self)
     73 @property
     74 def idle(self):
     75 """True if the DMA engine is idle
     76 
     77 `transfer` can only be called when the DMA is idle
     78 
     79 """
---> 80 return self._mmio.read(self._offset + 4) & 0x02 == 0x02

KeyboardInterrupt: 

```

As you can see its stuck at sendchannel.wait(). I know its because of DMA not being idle but i have configured the DMA for transfer already using the control register for continuous operation. I have also gone through the DMA debugging posts on community forum but still no luck. I cannot get ILA to work as I have completely exhausted my BRAM so any help or advice is welcomed how should I solve this issue.

Also please let me know that the above IPI block diagram is correct or not as at this point I am really not sure.

---

<div class="post-metadata">

**Author:** ![marioruiz](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.pynq.io/marioruiz/32/2177_2.png) [@marioruiz](https://discuss.pynq.io/u/marioruiz)\
**Post date:** [June 25, 2025, 2:37pm UTC](https://discuss.pynq.io/t/dma-stuck-at-wait/8573/2 "2025-06-25T14:37:15Z")

</div>

Hi @Karan_Mali,

Welcome to the PYNQ community.

Why are you using a FIFO interface? I will suggest you switch to an AXI4-Stream interface (it should be a pragma). You will need to handle TKEEP and TLAST correctly as well.

I wrote a blog showing how to debug issues with the DMA.

> [@Debugging Common DMA Issues \[Part 3\]](https://discuss.pynq.io/t/debugging-common-dma-issues-part-3/7157):
>
> Debugging Common DMA Issues If you frequent the PYNQ forum, one of the most common questions/issues we get is why DMA transfer do not work. In part 3 of this series, I will reproduce these issues and use both the register\_map and ILA to identify and address these issues. DMA channel not started Often, we see users reporting something like this. --------------------------------------------------------------------------- RuntimeError Traceback (most recent call last) Input In [4], in \<cell line:…

Mario

---

<div class="post-metadata">

**Author:** ![Karan\_Mali](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.pynq.io/karan_mali/32/6557_2.png) [@Karan\_Mali](https://discuss.pynq.io/u/Karan_Mali)\
**Post date:** [June 26, 2025, 6:26am UTC](https://discuss.pynq.io/t/dma-stuck-at-wait/8573/3 "2025-06-26T06:26:19Z")

</div>

Hi @marioruiz ,  
Well I am using a FIFO interface because the input is an array that is not a power of 2 and thus cannot use ap\_axiu or ap\_axis (which would have given me AXI port with side channels) nonetheless when I use axis interface pragma for AXI4-Stream interface the port synthesizes into AXI stream without side channels and also the input data width is converted into bytes so as here originally my array length was 784 now it becomes 784\*8 = 6272 data width port, as this is non power of two I have a data width mismatch error with the AXI ports on DMA.  
I already have been through the blog above (thank you for writing such a comprehensive blog) but things are not working. Also I couldn’t try the ILA debug cause I have exhausted my BRAM. Thus I request your kind help if possible in this regard. Also can you help me verify the IPI block diagram, as you said I will need to handle TKEEP and TLAST , do you think I have made any mistake in the block design?

---

<div class="post-metadata">

**Author:** ![marioruiz](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.pynq.io/marioruiz/32/2177_2.png) [@marioruiz](https://discuss.pynq.io/u/marioruiz)\
**Post date:** [June 26, 2025, 4:34pm UTC](https://discuss.pynq.io/t/dma-stuck-at-wait/8573/4 "2025-06-26T16:34:01Z")

</div>

Hi @Karan_Mali,

I do not think there is enough information in your post to understand what the IP is doing.  
I will recommend that you instantiate the memory inside of the HLS IP.

> as you said I will need to handle TKEEP and TLAST , do you think I have made any mistake in the block design?

From what I could see these signals are not being handled at all.

`ap_axiu` supports widths that are not power of two but need to be divisible by 8, so 784 is a valid value.
