Skip to content

Instantly share code, notes, and snippets.

@saethlin
Created September 15, 2019 16:20
Show Gist options
  • Select an option

  • Save saethlin/da11f3c37f17f3744361af1fb4970fd4 to your computer and use it in GitHub Desktop.

Select an option

Save saethlin/da11f3c37f17f3744361af1fb4970fd4 to your computer and use it in GitHub Desktop.
MPI errors for Samoxive
[c24b-s28.ufhpc:24695] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24695] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24695] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
[c24b-s28.ufhpc:24695] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
[c24b-s28.ufhpc:24697] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24697] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24697] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
[c24b-s28.ufhpc:24697] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
[c24b-s28.ufhpc:24698] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24698] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24698] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
[c24b-s28.ufhpc:24698] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
[c24b-s28.ufhpc:24691] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24691] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24691] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
[c24b-s28.ufhpc:24691] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
[c24b-s28.ufhpc:24696] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24696] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24696] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
[c24b-s28.ufhpc:24693] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24693] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24693] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
[c24b-s28.ufhpc:24696] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
[c24b-s28.ufhpc:24693] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
[c24b-s28.ufhpc:24688] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24689] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24692] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24689] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24689] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
[c24b-s28.ufhpc:24692] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24688] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24692] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
[c24b-s28.ufhpc:24688] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
[c24b-s28.ufhpc:24690] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24694] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685
[c24b-s28.ufhpc:24690] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24694] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177
[c24b-s28.ufhpc:24690] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
[c24b-s28.ufhpc:24694] OPAL ERROR: Unreachable in file ext2x_client.c at line 109
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
[c24b-s28.ufhpc:24689] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
[c24b-s28.ufhpc:24688] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
[c24b-s28.ufhpc:24692] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
[c24b-s28.ufhpc:24694] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
--------------------------------------------------------------------------
The application appears to have been direct launched using "srun",
but OMPI was not built with SLURM's PMI support and therefore cannot
execute. There are several options for building PMI support under
SLURM, depending upon the SLURM version you are using:
version 16.05 or later: you can use SLURM's PMIx support. This
requires that you configure and build SLURM --with-pmix.
Versions earlier than 16.05: you must use either SLURM's PMI-1 or
PMI-2 support. SLURM builds PMI-1 by default, or you can manually
install PMI-2. You must then build Open MPI using --with-pmi pointing
to the SLURM PMI library location.
Please configure as appropriate and try again.
--------------------------------------------------------------------------
*** An error occurred in MPI_Init_thread
*** on a NULL communicator
*** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort,
*** and potentially your MPI job)
[c24b-s28.ufhpc:24690] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed!
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
[c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252
srun: error: c24b-s28: tasks 9,15,21,27,33,39,45,51,55,58,60: Exited with exit code 1
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b60f0029490: collective timeout seq=0
slurmstepd: error: c24a-s25 [0] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b60f0029490: COLL_FENCE_RING state seq=0
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 0:c24a-s25
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 1:c24a-s40, prev 5:c24b-s42
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b60f0029508, #0, in-use=0
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b60f0029540, #1, in-use=0
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b60f0029578, #2, in-use=1
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=1/prev=4/fwd=4
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]:
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s40,c24b-s[24,40,42]
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: c24b-s28
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/14990
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b1c94028220: collective timeout seq=0
slurmstepd: error: c24b-s42 [5] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b1c94028220: COLL_FENCE_RING state seq=0
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 5:c24b-s42
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 0:c24a-s25, prev 4:c24b-s40
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b1c94028298, #0, in-use=0
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b1c940282d0, #1, in-use=0
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b1c94028308, #2, in-use=1
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=1/prev=4/fwd=4
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]:
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s[25,40],c24b-s[24,40]
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: c24b-s28
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/14768
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b9c0c021ec0: collective timeout seq=0
slurmstepd: error: c24b-s40 [4] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b9c0c021ec0: COLL_FENCE_RING state seq=0
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 4:c24b-s40
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 5:c24b-s42, prev 3:c24b-s28
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b9c0c021f38, #0, in-use=0
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b9c0c021f70, #1, in-use=0
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b9c0c021fa8, #2, in-use=1
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=1/prev=4/fwd=4
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]:
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s[25,40],c24b-s[24,42]
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: c24b-s28
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/13946
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b5678006960: collective timeout seq=0
slurmstepd: error: c24a-s40 [1] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b5678006960: COLL_FENCE_RING state seq=0
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 1:c24a-s40
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 2:c24b-s24, prev 0:c24a-s25
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b56780069d8, #0, in-use=0
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b5678006a10, #1, in-use=0
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b5678006a48, #2, in-use=1
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=1/prev=4/fwd=4
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]:
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s25,c24b-s[24,40,42]
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: c24b-s28
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/14990
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b50f4006960: collective timeout seq=0
slurmstepd: error: c24b-s24 [2] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b50f4006960: COLL_FENCE_RING state seq=0
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 2:c24b-s24
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 3:c24b-s28, prev 1:c24a-s40
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b50f40069d8, #0, in-use=0
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b50f4006a10, #1, in-use=0
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b50f4006a48, #2, in-use=1
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=1/prev=4/fwd=5
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]:
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s[25,40],c24b-s[40,42]
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: c24b-s28
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/14990
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b3f20006960: collective timeout seq=0
slurmstepd: error: c24b-s28 [3] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b3f20006960: COLL_FENCE_RING state seq=0
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 3:c24b-s28
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 4:c24b-s40, prev 2:c24b-s24
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b3f200069d8, #0, in-use=0
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b3f20006a10, #1, in-use=0
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b3f20006a48, #2, in-use=1
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=0/prev=5/fwd=4
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]:
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s[25,40],c24b-s[24,40,42]
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: -
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS
slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/14990
slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1317 [pmixp_coll_tree_reset_if_to] mpi/pmix: ERROR: 0x2b60f0024560: collective timeout seq=0
slurmstepd: error: c24a-s25 [0] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1336 [pmixp_coll_tree_log] mpi/pmix: ERROR: 0x2b60f0024560: COLL_FENCE_TREE state seq=0 contribs: loc=1/prnt=0/child=4
slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1338 [pmixp_coll_tree_log] mpi/pmix: ERROR: my peerid: 0:c24a-s25
slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1341 [pmixp_coll_tree_log] mpi/pmix: ERROR: root host: 0:c24a-s25
slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1355 [pmixp_coll_tree_log] mpi/pmix: ERROR: child contribs [5]:
slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1382 [pmixp_coll_tree_log] mpi/pmix: ERROR: done contrib: c24a-s40,c24b-s[24,40,42]
slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1384 [pmixp_coll_tree_log] mpi/pmix: ERROR: wait contrib: c24b-s28
slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1391 [pmixp_coll_tree_log] mpi/pmix: ERROR: status: coll=COLL_COLLECT upfw=COLL_SND_NONE dfwd=COLL_SND_NONE
slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1393 [pmixp_coll_tree_log] mpi/pmix: ERROR: dfwd status: dfwd_cb_cnt=0, dfwd_cb_wait=0
slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1396 [pmixp_coll_tree_log] mpi/pmix: ERROR: bufs (offset/size): upfw 164/16415, dfwd 69/16415
Mon Sep 2 11:30:25 2019
_
| | _
___ ___ | |_| |_
/ __) _ \| (_ _)
( (_| |_| | | | |_
\___)___/ \_) \__)
Cosmic Lyman-alpha Transfer code
for 3D Octree Data
Aaron Smith (2017)
Working with 64 tasks, 2 threads
slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1317 [pmixp_coll_tree_reset_if_to] mpi/pmix: ERROR: 0x2b1c94025e90: collective timeout seq=0
slurmstepd: error: c24b-s42 [5] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1336 [pmixp_coll_tree_log] mpi/pmix: ERROR: 0x2b1c94025e90: COLL_FENCE_TREE state seq=0 contribs: loc=1/prnt=0/child=0
slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1338 [pmixp_coll_tree_log] mpi/pmix: ERROR: my peerid: 5:c24b-s42
slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1341 [pmixp_coll_tree_log] mpi/pmix: ERROR: root host: 0:c24a-s25
slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1345 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt host: 0:c24a-s25
slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1346 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt contrib:
slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1348 [pmixp_coll_tree_log] mpi/pmix: ERROR: [0:c24a-s25] false
slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1391 [pmixp_coll_tree_log] mpi/pmix: ERROR: status: coll=COLL_UPFWD upfw=COLL_SND_DONE dfwd=COLL_SND_NONE
slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1393 [pmixp_coll_tree_log] mpi/pmix: ERROR: dfwd status: dfwd_cb_cnt=0, dfwd_cb_wait=0
slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1396 [pmixp_coll_tree_log] mpi/pmix: ERROR: bufs (offset/size): upfw 88/16415, dfwd 69/16415
slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1317 [pmixp_coll_tree_reset_if_to] mpi/pmix: ERROR: 0x2b9c0c027850: collective timeout seq=0
slurmstepd: error: c24b-s40 [4] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1336 [pmixp_coll_tree_log] mpi/pmix: ERROR: 0x2b9c0c027850: COLL_FENCE_TREE state seq=0 contribs: loc=1/prnt=0/child=0
slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1338 [pmixp_coll_tree_log] mpi/pmix: ERROR: my peerid: 4:c24b-s40
slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1341 [pmixp_coll_tree_log] mpi/pmix: ERROR: root host: 0:c24a-s25
slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1345 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt host: 0:c24a-s25
slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1346 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt contrib:
slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1348 [pmixp_coll_tree_log] mpi/pmix: ERROR: [0:c24a-s25] false
slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1391 [pmixp_coll_tree_log] mpi/pmix: ERROR: status: coll=COLL_UPFWD upfw=COLL_SND_DONE dfwd=COLL_SND_NONE
slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1393 [pmixp_coll_tree_log] mpi/pmix: ERROR: dfwd status: dfwd_cb_cnt=0, dfwd_cb_wait=0
slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1396 [pmixp_coll_tree_log] mpi/pmix: ERROR: bufs (offset/size): upfw 88/16415, dfwd 69/16415
slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1317 [pmixp_coll_tree_reset_if_to] mpi/pmix: ERROR: 0x2b56740280d0: collective timeout seq=0
slurmstepd: error: c24a-s40 [1] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1336 [pmixp_coll_tree_log] mpi/pmix: ERROR: 0x2b56740280d0: COLL_FENCE_TREE state seq=0 contribs: loc=1/prnt=0/child=0
slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1338 [pmixp_coll_tree_log] mpi/pmix: ERROR: my peerid: 1:c24a-s40
slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1341 [pmixp_coll_tree_log] mpi/pmix: ERROR: root host: 0:c24a-s25
slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1345 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt host: 0:c24a-s25
slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1346 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt contrib:
slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1348 [pmixp_coll_tree_log] mpi/pmix: ERROR: [0:c24a-s25] false
slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1391 [pmixp_coll_tree_log] mpi/pmix: ERROR: status: coll=COLL_UPFWD upfw=COLL_SND_DONE dfwd=COLL_SND_NONE
slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1393 [pmixp_coll_tree_log] mpi/pmix: ERROR: dfwd status: dfwd_cb_cnt=0, dfwd_cb_wait=0
slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1396 [pmixp_coll_tree_log] mpi/pmix: ERROR: bufs (offset/size): upfw 88/16415, dfwd 69/16415
slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1317 [pmixp_coll_tree_reset_if_to] mpi/pmix: ERROR: 0x2b50f0024f60: collective timeout seq=0
slurmstepd: error: c24b-s24 [2] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state
slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1336 [pmixp_coll_tree_log] mpi/pmix: ERROR: 0x2b50f0024f60: COLL_FENCE_TREE state seq=0 contribs: loc=1/prnt=0/child=0
slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1338 [pmixp_coll_tree_log] mpi/pmix: ERROR: my peerid: 2:c24b-s24
slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1341 [pmixp_coll_tree_log] mpi/pmix: ERROR: root host: 0:c24a-s25
slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1345 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt host: 0:c24a-s25
slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1346 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt contrib:
slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1348 [pmixp_coll_tree_log] mpi/pmix: ERROR: [0:c24a-s25] false
slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1391 [pmixp_coll_tree_log] mpi/pmix: ERROR: status: coll=COLL_UPFWD upfw=COLL_SND_DONE dfwd=COLL_SND_NONE
slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1393 [pmixp_coll_tree_log] mpi/pmix: ERROR: dfwd status: dfwd_cb_cnt=0, dfwd_cb_wait=0
slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1396 [pmixp_coll_tree_log] mpi/pmix: ERROR: bufs (offset/size): upfw 88/16415, dfwd 69/16415
Octree grid data:
redshift = 1.90323
n_cells = 15976393
n_leafs = 13979344
r = [ -75 -75 -75 ]
[ -75 -75 -75 ]
[ ... ]
[ 73.2422 74.4141 72.6562 ]
[ 73.2422 74.4141 73.2422 ] kpc
w = [ 150000 75000 75000 ... 585.937 585.937 585.937 ] pc
V = [ 0 0 0 ... 2.01166e+08 2.01166e+08 2.01166e+08 ] pc^3
n_H = [ 0 0 0 ... 0.000277625 0.000245887 0.000238655 ] 1/cm^3
rho = [ 0 0 0 ... 6.11989e-28 5.42307e-28 5.26078e-28 ] g/cm^3
m = [ 0 0 0 ... 1819.02 1611.9 1563.67 ] Msun
T = [ 0 0 0 ... 1.69173e+06 1.78295e+06 1.7331e+06 ] K
v = [ 0 0 0 ]
[ 0 0 0 ]
[ ... ]
[ 82.0097 109.724 86.9507 ]
[ 80.6832 117.957 86.9683 ] km/s
Z = [ 0 0 0 ... 0.00107035 0.00158481 0.00105486 ]
x_HI = [ 0 0 0 ... 2.50207e-07 2.26406e-07 2.38765e-07 ]
x_HII = [ 0 0 0 ... 1 1 1 ]
min/max cell data:
ray_eps = 1.5e-07 pc
(x,y,z) = [(-75, -75, -75),
(75, 75, 75)] kpc
w = [9.15527, 2343.75] pc
V = [767.386, 1.28746e+10] pc^3
n_H = [2.42072e-06, 4908.58] 1/cm^3
rho = [5.41637e-30, 1.18512e-20] g/cm^3
m = [102.996, 4.94873e+06] Msun
T = [10.0004, 7.76786e+07] K
(vx,vy,vz) = [(-1775.27, -1726.74, -1206.34),
(1.79769e+303, 1.79769e+303, 1.79769e+303)] km/s
Z = [2e-06, 0.162382]
x_HI = [1e-10, 1]
x_HII = [0, 1]
SB Data Cube: 6.25 GB (1, 1024, 1024, 800)
Successfully set up initial conditions.
Simulation parameters:
Inital conditions file = /ufrc/narayanan/kimockb/FIRE2/h113_HR_sn1dy300ro100ss/snapdir_179/converted_snapshot_179.0.hdf5
Data output directory = /ufrc/narayanan/kimockb/experiments/median/663
Lyman-alpha constants:
f12 = 0.4162
nu0 = 2.466e+15 Hz
lambda0 = 1215.7 angstroms
DnuL = 9.936e+07 Hz
gDa = 0.539244
Photon parameters:
nbin = 800
n_photons = 1000000
Note: Frequencies are given in the rest frame.
freq_type = Delta_v
freq_lims = (-4000, 4000) km/s
resolution = 10 km/s
| x_res = 0.0350568
| nu_res = 8.22569e+10 Hz
| l_res = 0.0405515 ang
| Dv_res = 10 km/s
| R_res = 29979.2
T_exit = 10000 K
v_exit = (0, 0, 0) km/s (user specified)
r_insert = (0, 0, 0) pc
i_insert = 12891697
Check:
r[i_insert] = (73.2422, 73.2422, 73.2422) pc
v[i_insert] = (244.405, 104.288, -380.998) km/s
Dust properties: SMC dust model
sigma_dust = 1.17352e-19 cm^2
albedo = 0.32
f_ion = 0.01
Emission based on the cell recombination and collisional rates: (L_Lya_col/L_Lya = 0.966759)
i_mid = 7138689
j_cdf = [ 7.89562e-08 1.57912e-07 2.36869e-07 ... 1 1 1 ]
j_weights = [ 9.67412e-13 2.4547e-12 2.98404e-12 ... 7.95827e-12 5.66427e-12 5.6197e-12 ]
Strength of the Lyman-alpha source:
Luminosity = 2.43948e+10 Lsun
= 9.36518e+43 erg/s
Prod. Rate = 5.73148e+54 photons/s
LOS directions are healpix in RING ordering.
n_exp = -1
n_LOS = 1
k_LOS = [ 0 0 1 ]
Surface brightness image information:
pixels = 1024
radius = 75 kpc
pix dx = 146.484 pc
pix dl = 0.0169666 arcsec
pix dA = 0.000287866 arcsec^2
x-axis = [ 1 0 0 ]
y-axis = [ 0 1 0 ]
The core-skipping acceleration scheme is turned on.
[ Delta_nH0 < 0.5 nH0 ] [ Delta_v < 10 vth ]
Efficiency of non-local pre-calculation: (atau > 1) = 2.38459%
Lyman-alpha data:
kH0 = [ 0 0 0 ... 3.14999e-25 2.45908e-25 2.55298e-25 ] 1/cm
kdust = [ 0 0 0 ... 3.48729e-28 4.57314e-28 2.9544e-28 ] 1/cm
a = [ 0 0 0 ... 3.61496e-05 3.52127e-05 3.57156e-05 ]
atau = [ 0 0 0 ... 0 0 0 ]
min/max Lyman-alpha cell data:
(vx,vy,vz) = [(-1260.88, -1657.54, -1286.01),
(1.79769e+308, 1.79769e+308, 1.79769e+308)] vth
kH0 = [1.32774e-29, 1.92445e-09] 1/cm
kdust = [7.9063e-32, 5.26907e-17] 1/cm
T = [10.0004, 7.76786e+07] K
vth = sqrt(2*kB*T/mH)
= [0.406208, 1132.11] km/s
DnuD = vth*nu0/c
= [3.34134e+09, 9.31242e+12] Hz
a = 0.5*DnuL/DnuD
= [0.0148683, 5.33481e-06]
g = h*DnuD/2*kB*T = 0.539537*a
= [0.00801762, 2.87676e-06]
sigma0 = f12*sqrt(pi)*ee^2/(me*c*DnuD)
= [1.86513e-12, 6.69217e-16] cm^2
atau_NL = [0, 1.09126e+06]
tau0 (LOS) = [1.79769e+308, -1.79769e+308]
tau_dust = [1.79769e+308, -1.79769e+308]
+------------------------------------------------------+
| Initialization complete. Starting COLT calculations. |
+------------------------------------------------------+
00:24:46 remaining
00:23:39 remaining
00:22:12 remaining
00:20:57 remaining
00:19:39 remaining
00:18:20 remaining
00:17:02 remaining
00:15:44 remaining
00:14:24 remaining
00:13:06 remaining
00:11:48 remaining
00:10:30 remaining
00:09:13 remaining
00:07:54 remaining
00:06:35 remaining
slurmstepd: error: *** JOB 41034560 ON c24a-s25 CANCELLED AT 2019-09-02T12:00:45 DUE TO TIME LIMIT ***
slurmstepd: error: *** STEP 41034560.0 ON c24a-s25 CANCELLED AT 2019-09-02T12:00:45 DUE TO TIME LIMIT ***
srun: error: c24b-s28: task 3: Exited with exit code 1
@47oo

47oo commented Apr 30, 2021

Copy link
Copy Markdown

I have this problem too, on slurm 19.05.7 with pmix ,but ,when I run more than 200+ nodes and tasks-per-node > 50 ,I also find log on slurmd with wait contrib: hostname , do you solve this problem or know this problem happened

@saethlin

Copy link
Copy Markdown
Author

Unfortunately this was a long time ago, and I'm not an HPC user anymore. But here's what I can come up with.

I eventually settled on these MCA settings. I struggled mightily to figure out what they actually do, mostly I guessed around based on what the local admins told me:

export OMPI_MCA_pml="^ucx"
export OMPI_MCA_btl="self,vader,openib"
export OMPI_MCA_oob_tcp_listen_mode="listen_thread"

And I did eventually discover that I needed to provide a specific pmix version, which at one point changed based on what software was available on the cluster. What eventually worked for me:

srun --mpi=pmix_v2 /path/to/executable

@47oo

47oo commented May 1, 2021

Copy link
Copy Markdown

Thanks,I will try your method and change pmix or slurm version to test next week, and will tell you the result.

@47oo

47oo commented Jun 4, 2021

Copy link
Copy Markdown

We found the problem ,we modfiy some code from pmix like handle_timeout to enlarge, then UNREACHABLE not appeared, thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment