This occured after training a bit. It's the first time I think that I see this. Node n23g0014 in RWTH cluster.
Thread 390647 (active): "MainThread"
copy_to_device (returnn/returnn/torch/frontend/_backend.py:169)
copy_to_device (returnn/returnn/frontend/device.py:40)
get_dyn_size_ext_for_device (returnn/returnn/tensor/_dim_extra.py:837)
get_size_tensor (returnn/returnn/tensor/_dim_extra.py:1941)
pad (returnn/returnn/torch/frontend/_backend.py:538)
pad (returnn/returnn/frontend/array_.py:519)
aed_training (denoising_lm_aed_2024/models/aed_v3.py:252)
_returnn_train_step (i6_experiments/users/zeyer/train_v4.py:248)
_run_step (returnn/returnn/torch/engine.py:932)
train_epoch (returnn/returnn/torch/engine.py:465)
train (returnn/returnn/torch/engine.py:269)
execute_main_task (returnn/returnn/__main__.py:543)
main (returnn/returnn/__main__.py:744)
<module> (returnn/rnn.py:11)
#0 0x0000148ff604fff7 in ?? () from /lib64/libcuda.so.1
#1 0x0000148ff5e1a6b3 in ?? () from /lib64/libcuda.so.1
#2 0x0000148ff5e6b87b in ?? () from /lib64/libcuda.so.1
#3 0x0000148ff6951388 in ?? () from /lib64/libcuda.so.1
#4 0x0000148ff69517a8 in ?? () from /lib64/libcuda.so.1
#5 0x0000148ff5e3faec in ?? () from /lib64/libcuda.so.1
#6 0x0000148ff5eecb42 in ?? () from /lib64/libcuda.so.1
#7 0x0000148ff6912375 in ?? () from /lib64/libcuda.so.1
#8 0x0000148ff5f80e67 in ?? () from /lib64/libcuda.so.1
#9 0x00001490c01cf3a5 in ?? ()
from /home/az668407/work/py-envs/py3.12-torch2.7/lib/python3.12/site-packages/torch/lib/../../nvidia/cuda_runtime/lib/libcudart.so.12
#10 0x00001490c02307d8 in cudaStreamSynchronize ()
from /home/az668407/work/py-envs/py3.12-torch2.7/lib/python3.12/site-packages/torch/lib/../../nvidia/cuda_runtime/lib/libcudart.so.12
#11 0x000014906b917fae in at::native::copy_kernel_cuda(at::TensorIterator&, bool) ()
from /home/az668407/work/py-envs/py3.12-torch2.7/lib/python3.12/site-packages/torch/lib/libtorch_cuda.so
#12 0x00001490a5857252 in at::native::copy_impl(at::Tensor&, at::Tensor const&, bool) [clone .isra.0] ()
from /home/az668407/work/py-envs/py3.12-torch2.7/lib/python3.12/site-packages/torch/lib/libtorch_cpu.so
#13 0x00001490a5858bac in at::native::copy_(at::Tensor&, at::Tensor const&, bool) ()
from /home/az668407/work/py-envs/py3.12-torch2.7/lib/python3.12/site-packages/torch/lib/libtorch_cpu.so
#14 0x00001490a651fd35 in at::_ops::copy_::call(at::Tensor&, at::Tensor const&, bool) ()
from /home/az668407/work/py-envs/py3.12-torch2.7/lib/python3.12/site-packages/torch/lib/libtorch_cpu.so
#15 0x00001490a5b303f5 in at::native::_to_copy(at::Tensor const&, std::optional<c10::ScalarType>, std::optional<c10::Layout>, std::optional<c10::Device>, std::optional<bool>, bool, std::optional<c10::MemoryFormat>) ()
from /home/az668407/work/py-envs/py3.12-torch2.7/lib/python3.12/site-packages/torch/lib/libtorch_cpu.so
#16 0x00001490a695488d in c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor (at::Tensor const&, std::optional<c10::ScalarType>, std::optional<c10::Layout>, std::optional<c10::Device>, std::optional<bool>, bool, std::optional<c10::MemoryFormat>), &at::(anonymous namespace)::(anonymous namespace)::wrapper_CompositeExplicitAutograd___to_copy>, at::Tensor, c10::guts::typelist::typelist<at::Tensor const&, std::optional<c10::ScalarType>, std::optional<c10::Layout>, std::optional<c10::Device>, std::optional<bool>, bool, std::optional<c10::MemoryFormat> > >, at::Tensor (at::Tensor const&, std::optional<c10::ScalarType>, std::optional<c10::Layout>, std::optional<c10::Device>, std::optional<bool>, bool, std::optional<c10::MemoryFormat>)>::call(c10::OperatorKernel*, c10::DispatchKeySet, at::Tensor const&, std::optional<c10::ScalarType>, std::optional<c10::Layout>, std::optional<c10::Device>, std::optional<bool>, bool, std::optional<c10::MemoryFormat>) ()
...
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 570.195.03 Driver Version: 570.195.03 CUDA Version: 12.8 |
|-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA H100 On | 00000000:1B:00.0 Off | 0 |
| N/A 38C P0 121W / 700W | 40161MiB / 95830MiB | 100% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
| 1 NVIDIA H100 On | 00000000:2C:00.0 Off | 0 |
| N/A 44C P0 253W / 700W | 32999MiB / 95830MiB | 58% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
| 2 NVIDIA H100 On | 00000000:9D:00.0 Off | 0 |
| N/A 37C P0 121W / 700W | 32507MiB / 95830MiB | 100% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
| 3 NVIDIA H100 On | 00000000:AD:00.0 Off | 0 |
| N/A 45C P0 263W / 700W | 32581MiB / 95830MiB | 38% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| 0 N/A N/A 390647 C ...vs/py3.12-torch2.7/bin/python 40134MiB |
| 1 N/A N/A 255143 C ...vs/py3.12-torch2.7/bin/python 32872MiB |
| 2 N/A N/A 255142 C ...vs/py3.12-torch2.7/bin/python 32478MiB |
| 3 N/A N/A 236563 C ...vs/py3.12-torch2.7/bin/python 32552MiB |
+-----------------------------------------------------------------------------------------+
Probably some hardware issue. Not sure we can do much about it.
This occured after training a bit. It's the first time I think that I see this. Node
n23g0014in RWTH cluster.py-spy:gdb:nvidia-smi:Probably some hardware issue. Not sure we can do much about it.