Created
September 15, 2019 16:20
-
-
Save saethlin/da11f3c37f17f3744361af1fb4970fd4 to your computer and use it in GitHub Desktop.
MPI errors for Samoxive
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| [c24b-s28.ufhpc:24695] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24695] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24695] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| [c24b-s28.ufhpc:24695] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| [c24b-s28.ufhpc:24697] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24697] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24697] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| [c24b-s28.ufhpc:24697] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| [c24b-s28.ufhpc:24698] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24698] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24698] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| [c24b-s28.ufhpc:24698] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| [c24b-s28.ufhpc:24691] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24691] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24691] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| [c24b-s28.ufhpc:24691] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| [c24b-s28.ufhpc:24696] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24696] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24696] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| [c24b-s28.ufhpc:24693] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24693] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24693] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| [c24b-s28.ufhpc:24696] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| [c24b-s28.ufhpc:24693] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| [c24b-s28.ufhpc:24688] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24689] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24692] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24689] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24689] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| [c24b-s28.ufhpc:24692] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24688] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24692] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| [c24b-s28.ufhpc:24688] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| [c24b-s28.ufhpc:24690] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24694] PMIX ERROR: PMIX TEMPORARILY UNAVAILABLE in file ptl_tcp.c at line 685 | |
| [c24b-s28.ufhpc:24690] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24694] PMIX ERROR: UNREACHABLE in file ptl_usock.c at line 177 | |
| [c24b-s28.ufhpc:24690] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| [c24b-s28.ufhpc:24694] OPAL ERROR: Unreachable in file ext2x_client.c at line 109 | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| [c24b-s28.ufhpc:24689] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| [c24b-s28.ufhpc:24688] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| [c24b-s28.ufhpc:24692] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| [c24b-s28.ufhpc:24694] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| -------------------------------------------------------------------------- | |
| The application appears to have been direct launched using "srun", | |
| but OMPI was not built with SLURM's PMI support and therefore cannot | |
| execute. There are several options for building PMI support under | |
| SLURM, depending upon the SLURM version you are using: | |
| version 16.05 or later: you can use SLURM's PMIx support. This | |
| requires that you configure and build SLURM --with-pmix. | |
| Versions earlier than 16.05: you must use either SLURM's PMI-1 or | |
| PMI-2 support. SLURM builds PMI-1 by default, or you can manually | |
| install PMI-2. You must then build Open MPI using --with-pmi pointing | |
| to the SLURM PMI library location. | |
| Please configure as appropriate and try again. | |
| -------------------------------------------------------------------------- | |
| *** An error occurred in MPI_Init_thread | |
| *** on a NULL communicator | |
| *** MPI_ERRORS_ARE_FATAL (processes in this communicator will now abort, | |
| *** and potentially your MPI job) | |
| [c24b-s28.ufhpc:24690] Local abort before MPI_INIT completed completed successfully, but am not able to aggregate error messages, and not able to guarantee that all other processes were killed! | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| [c24b-s28.ufhpc:24677] PMIX ERROR: UNREACHABLE in file ptl_tcp_component.c at line 1252 | |
| srun: error: c24b-s28: tasks 9,15,21,27,33,39,45,51,55,58,60: Exited with exit code 1 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b60f0029490: collective timeout seq=0 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b60f0029490: COLL_FENCE_RING state seq=0 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 0:c24a-s25 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 1:c24a-s40, prev 5:c24b-s42 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b60f0029508, #0, in-use=0 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b60f0029540, #1, in-use=0 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b60f0029578, #2, in-use=1 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=1/prev=4/fwd=4 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]: | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s40,c24b-s[24,40,42] | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: c24b-s28 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/14990 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b1c94028220: collective timeout seq=0 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b1c94028220: COLL_FENCE_RING state seq=0 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 5:c24b-s42 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 0:c24a-s25, prev 4:c24b-s40 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b1c94028298, #0, in-use=0 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b1c940282d0, #1, in-use=0 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b1c94028308, #2, in-use=1 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=1/prev=4/fwd=4 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]: | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s[25,40],c24b-s[24,40] | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: c24b-s28 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/14768 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b9c0c021ec0: collective timeout seq=0 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b9c0c021ec0: COLL_FENCE_RING state seq=0 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 4:c24b-s40 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 5:c24b-s42, prev 3:c24b-s28 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b9c0c021f38, #0, in-use=0 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b9c0c021f70, #1, in-use=0 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b9c0c021fa8, #2, in-use=1 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=1/prev=4/fwd=4 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]: | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s[25,40],c24b-s[24,42] | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: c24b-s28 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/13946 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b5678006960: collective timeout seq=0 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b5678006960: COLL_FENCE_RING state seq=0 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 1:c24a-s40 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 2:c24b-s24, prev 0:c24a-s25 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b56780069d8, #0, in-use=0 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b5678006a10, #1, in-use=0 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b5678006a48, #2, in-use=1 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=1/prev=4/fwd=4 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]: | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s25,c24b-s[24,40,42] | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: c24b-s28 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/14990 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b50f4006960: collective timeout seq=0 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b50f4006960: COLL_FENCE_RING state seq=0 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 2:c24b-s24 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 3:c24b-s28, prev 1:c24a-s40 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b50f40069d8, #0, in-use=0 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b50f4006a10, #1, in-use=0 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b50f4006a48, #2, in-use=1 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=1/prev=4/fwd=5 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]: | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s[25,40],c24b-s[40,42] | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: c24b-s28 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/14990 | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:738 [pmixp_coll_ring_reset_if_to] mpi/pmix: ERROR: 0x2b3f20006960: collective timeout seq=0 | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:756 [pmixp_coll_ring_log] mpi/pmix: ERROR: 0x2b3f20006960: COLL_FENCE_RING state seq=0 | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:758 [pmixp_coll_ring_log] mpi/pmix: ERROR: my peerid: 3:c24b-s28 | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:765 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor id: next 4:c24b-s40, prev 2:c24b-s24 | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b3f200069d8, #0, in-use=0 | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b3f20006a10, #1, in-use=0 | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:775 [pmixp_coll_ring_log] mpi/pmix: ERROR: Context ptr=0x2b3f20006a48, #2, in-use=1 | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:786 [pmixp_coll_ring_log] mpi/pmix: ERROR: seq=0 contribs: loc=0/prev=5/fwd=4 | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:788 [pmixp_coll_ring_log] mpi/pmix: ERROR: neighbor contribs [6]: | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:821 [pmixp_coll_ring_log] mpi/pmix: ERROR: done contrib: c24a-s[25,40],c24b-s[24,40,42] | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:823 [pmixp_coll_ring_log] mpi/pmix: ERROR: wait contrib: - | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:825 [pmixp_coll_ring_log] mpi/pmix: ERROR: status=PMIXP_COLL_RING_PROGRESS | |
| slurmstepd: error: c24b-s28 [3] pmixp_coll_ring.c:829 [pmixp_coll_ring_log] mpi/pmix: ERROR: buf (offset/size): 7256/14990 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1317 [pmixp_coll_tree_reset_if_to] mpi/pmix: ERROR: 0x2b60f0024560: collective timeout seq=0 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1336 [pmixp_coll_tree_log] mpi/pmix: ERROR: 0x2b60f0024560: COLL_FENCE_TREE state seq=0 contribs: loc=1/prnt=0/child=4 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1338 [pmixp_coll_tree_log] mpi/pmix: ERROR: my peerid: 0:c24a-s25 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1341 [pmixp_coll_tree_log] mpi/pmix: ERROR: root host: 0:c24a-s25 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1355 [pmixp_coll_tree_log] mpi/pmix: ERROR: child contribs [5]: | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1382 [pmixp_coll_tree_log] mpi/pmix: ERROR: done contrib: c24a-s40,c24b-s[24,40,42] | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1384 [pmixp_coll_tree_log] mpi/pmix: ERROR: wait contrib: c24b-s28 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1391 [pmixp_coll_tree_log] mpi/pmix: ERROR: status: coll=COLL_COLLECT upfw=COLL_SND_NONE dfwd=COLL_SND_NONE | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1393 [pmixp_coll_tree_log] mpi/pmix: ERROR: dfwd status: dfwd_cb_cnt=0, dfwd_cb_wait=0 | |
| slurmstepd: error: c24a-s25 [0] pmixp_coll_tree.c:1396 [pmixp_coll_tree_log] mpi/pmix: ERROR: bufs (offset/size): upfw 164/16415, dfwd 69/16415 | |
| Mon Sep 2 11:30:25 2019 | |
| _ | |
| | | _ | |
| ___ ___ | |_| |_ | |
| / __) _ \| (_ _) | |
| ( (_| |_| | | | |_ | |
| \___)___/ \_) \__) | |
| Cosmic Lyman-alpha Transfer code | |
| for 3D Octree Data | |
| Aaron Smith (2017) | |
| Working with 64 tasks, 2 threads | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1317 [pmixp_coll_tree_reset_if_to] mpi/pmix: ERROR: 0x2b1c94025e90: collective timeout seq=0 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1336 [pmixp_coll_tree_log] mpi/pmix: ERROR: 0x2b1c94025e90: COLL_FENCE_TREE state seq=0 contribs: loc=1/prnt=0/child=0 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1338 [pmixp_coll_tree_log] mpi/pmix: ERROR: my peerid: 5:c24b-s42 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1341 [pmixp_coll_tree_log] mpi/pmix: ERROR: root host: 0:c24a-s25 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1345 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt host: 0:c24a-s25 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1346 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt contrib: | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1348 [pmixp_coll_tree_log] mpi/pmix: ERROR: [0:c24a-s25] false | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1391 [pmixp_coll_tree_log] mpi/pmix: ERROR: status: coll=COLL_UPFWD upfw=COLL_SND_DONE dfwd=COLL_SND_NONE | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1393 [pmixp_coll_tree_log] mpi/pmix: ERROR: dfwd status: dfwd_cb_cnt=0, dfwd_cb_wait=0 | |
| slurmstepd: error: c24b-s42 [5] pmixp_coll_tree.c:1396 [pmixp_coll_tree_log] mpi/pmix: ERROR: bufs (offset/size): upfw 88/16415, dfwd 69/16415 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1317 [pmixp_coll_tree_reset_if_to] mpi/pmix: ERROR: 0x2b9c0c027850: collective timeout seq=0 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1336 [pmixp_coll_tree_log] mpi/pmix: ERROR: 0x2b9c0c027850: COLL_FENCE_TREE state seq=0 contribs: loc=1/prnt=0/child=0 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1338 [pmixp_coll_tree_log] mpi/pmix: ERROR: my peerid: 4:c24b-s40 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1341 [pmixp_coll_tree_log] mpi/pmix: ERROR: root host: 0:c24a-s25 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1345 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt host: 0:c24a-s25 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1346 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt contrib: | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1348 [pmixp_coll_tree_log] mpi/pmix: ERROR: [0:c24a-s25] false | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1391 [pmixp_coll_tree_log] mpi/pmix: ERROR: status: coll=COLL_UPFWD upfw=COLL_SND_DONE dfwd=COLL_SND_NONE | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1393 [pmixp_coll_tree_log] mpi/pmix: ERROR: dfwd status: dfwd_cb_cnt=0, dfwd_cb_wait=0 | |
| slurmstepd: error: c24b-s40 [4] pmixp_coll_tree.c:1396 [pmixp_coll_tree_log] mpi/pmix: ERROR: bufs (offset/size): upfw 88/16415, dfwd 69/16415 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1317 [pmixp_coll_tree_reset_if_to] mpi/pmix: ERROR: 0x2b56740280d0: collective timeout seq=0 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1336 [pmixp_coll_tree_log] mpi/pmix: ERROR: 0x2b56740280d0: COLL_FENCE_TREE state seq=0 contribs: loc=1/prnt=0/child=0 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1338 [pmixp_coll_tree_log] mpi/pmix: ERROR: my peerid: 1:c24a-s40 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1341 [pmixp_coll_tree_log] mpi/pmix: ERROR: root host: 0:c24a-s25 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1345 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt host: 0:c24a-s25 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1346 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt contrib: | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1348 [pmixp_coll_tree_log] mpi/pmix: ERROR: [0:c24a-s25] false | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1391 [pmixp_coll_tree_log] mpi/pmix: ERROR: status: coll=COLL_UPFWD upfw=COLL_SND_DONE dfwd=COLL_SND_NONE | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1393 [pmixp_coll_tree_log] mpi/pmix: ERROR: dfwd status: dfwd_cb_cnt=0, dfwd_cb_wait=0 | |
| slurmstepd: error: c24a-s40 [1] pmixp_coll_tree.c:1396 [pmixp_coll_tree_log] mpi/pmix: ERROR: bufs (offset/size): upfw 88/16415, dfwd 69/16415 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1317 [pmixp_coll_tree_reset_if_to] mpi/pmix: ERROR: 0x2b50f0024f60: collective timeout seq=0 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll.c:281 [pmixp_coll_log] mpi/pmix: ERROR: Dumping collective state | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1336 [pmixp_coll_tree_log] mpi/pmix: ERROR: 0x2b50f0024f60: COLL_FENCE_TREE state seq=0 contribs: loc=1/prnt=0/child=0 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1338 [pmixp_coll_tree_log] mpi/pmix: ERROR: my peerid: 2:c24b-s24 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1341 [pmixp_coll_tree_log] mpi/pmix: ERROR: root host: 0:c24a-s25 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1345 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt host: 0:c24a-s25 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1346 [pmixp_coll_tree_log] mpi/pmix: ERROR: prnt contrib: | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1348 [pmixp_coll_tree_log] mpi/pmix: ERROR: [0:c24a-s25] false | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1391 [pmixp_coll_tree_log] mpi/pmix: ERROR: status: coll=COLL_UPFWD upfw=COLL_SND_DONE dfwd=COLL_SND_NONE | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1393 [pmixp_coll_tree_log] mpi/pmix: ERROR: dfwd status: dfwd_cb_cnt=0, dfwd_cb_wait=0 | |
| slurmstepd: error: c24b-s24 [2] pmixp_coll_tree.c:1396 [pmixp_coll_tree_log] mpi/pmix: ERROR: bufs (offset/size): upfw 88/16415, dfwd 69/16415 | |
| Octree grid data: | |
| redshift = 1.90323 | |
| n_cells = 15976393 | |
| n_leafs = 13979344 | |
| r = [ -75 -75 -75 ] | |
| [ -75 -75 -75 ] | |
| [ ... ] | |
| [ 73.2422 74.4141 72.6562 ] | |
| [ 73.2422 74.4141 73.2422 ] kpc | |
| w = [ 150000 75000 75000 ... 585.937 585.937 585.937 ] pc | |
| V = [ 0 0 0 ... 2.01166e+08 2.01166e+08 2.01166e+08 ] pc^3 | |
| n_H = [ 0 0 0 ... 0.000277625 0.000245887 0.000238655 ] 1/cm^3 | |
| rho = [ 0 0 0 ... 6.11989e-28 5.42307e-28 5.26078e-28 ] g/cm^3 | |
| m = [ 0 0 0 ... 1819.02 1611.9 1563.67 ] Msun | |
| T = [ 0 0 0 ... 1.69173e+06 1.78295e+06 1.7331e+06 ] K | |
| v = [ 0 0 0 ] | |
| [ 0 0 0 ] | |
| [ ... ] | |
| [ 82.0097 109.724 86.9507 ] | |
| [ 80.6832 117.957 86.9683 ] km/s | |
| Z = [ 0 0 0 ... 0.00107035 0.00158481 0.00105486 ] | |
| x_HI = [ 0 0 0 ... 2.50207e-07 2.26406e-07 2.38765e-07 ] | |
| x_HII = [ 0 0 0 ... 1 1 1 ] | |
| min/max cell data: | |
| ray_eps = 1.5e-07 pc | |
| (x,y,z) = [(-75, -75, -75), | |
| (75, 75, 75)] kpc | |
| w = [9.15527, 2343.75] pc | |
| V = [767.386, 1.28746e+10] pc^3 | |
| n_H = [2.42072e-06, 4908.58] 1/cm^3 | |
| rho = [5.41637e-30, 1.18512e-20] g/cm^3 | |
| m = [102.996, 4.94873e+06] Msun | |
| T = [10.0004, 7.76786e+07] K | |
| (vx,vy,vz) = [(-1775.27, -1726.74, -1206.34), | |
| (1.79769e+303, 1.79769e+303, 1.79769e+303)] km/s | |
| Z = [2e-06, 0.162382] | |
| x_HI = [1e-10, 1] | |
| x_HII = [0, 1] | |
| SB Data Cube: 6.25 GB (1, 1024, 1024, 800) | |
| Successfully set up initial conditions. | |
| Simulation parameters: | |
| Inital conditions file = /ufrc/narayanan/kimockb/FIRE2/h113_HR_sn1dy300ro100ss/snapdir_179/converted_snapshot_179.0.hdf5 | |
| Data output directory = /ufrc/narayanan/kimockb/experiments/median/663 | |
| Lyman-alpha constants: | |
| f12 = 0.4162 | |
| nu0 = 2.466e+15 Hz | |
| lambda0 = 1215.7 angstroms | |
| DnuL = 9.936e+07 Hz | |
| gDa = 0.539244 | |
| Photon parameters: | |
| nbin = 800 | |
| n_photons = 1000000 | |
| Note: Frequencies are given in the rest frame. | |
| freq_type = Delta_v | |
| freq_lims = (-4000, 4000) km/s | |
| resolution = 10 km/s | |
| | x_res = 0.0350568 | |
| | nu_res = 8.22569e+10 Hz | |
| | l_res = 0.0405515 ang | |
| | Dv_res = 10 km/s | |
| | R_res = 29979.2 | |
| T_exit = 10000 K | |
| v_exit = (0, 0, 0) km/s (user specified) | |
| r_insert = (0, 0, 0) pc | |
| i_insert = 12891697 | |
| Check: | |
| r[i_insert] = (73.2422, 73.2422, 73.2422) pc | |
| v[i_insert] = (244.405, 104.288, -380.998) km/s | |
| Dust properties: SMC dust model | |
| sigma_dust = 1.17352e-19 cm^2 | |
| albedo = 0.32 | |
| f_ion = 0.01 | |
| Emission based on the cell recombination and collisional rates: (L_Lya_col/L_Lya = 0.966759) | |
| i_mid = 7138689 | |
| j_cdf = [ 7.89562e-08 1.57912e-07 2.36869e-07 ... 1 1 1 ] | |
| j_weights = [ 9.67412e-13 2.4547e-12 2.98404e-12 ... 7.95827e-12 5.66427e-12 5.6197e-12 ] | |
| Strength of the Lyman-alpha source: | |
| Luminosity = 2.43948e+10 Lsun | |
| = 9.36518e+43 erg/s | |
| Prod. Rate = 5.73148e+54 photons/s | |
| LOS directions are healpix in RING ordering. | |
| n_exp = -1 | |
| n_LOS = 1 | |
| k_LOS = [ 0 0 1 ] | |
| Surface brightness image information: | |
| pixels = 1024 | |
| radius = 75 kpc | |
| pix dx = 146.484 pc | |
| pix dl = 0.0169666 arcsec | |
| pix dA = 0.000287866 arcsec^2 | |
| x-axis = [ 1 0 0 ] | |
| y-axis = [ 0 1 0 ] | |
| The core-skipping acceleration scheme is turned on. | |
| [ Delta_nH0 < 0.5 nH0 ] [ Delta_v < 10 vth ] | |
| Efficiency of non-local pre-calculation: (atau > 1) = 2.38459% | |
| Lyman-alpha data: | |
| kH0 = [ 0 0 0 ... 3.14999e-25 2.45908e-25 2.55298e-25 ] 1/cm | |
| kdust = [ 0 0 0 ... 3.48729e-28 4.57314e-28 2.9544e-28 ] 1/cm | |
| a = [ 0 0 0 ... 3.61496e-05 3.52127e-05 3.57156e-05 ] | |
| atau = [ 0 0 0 ... 0 0 0 ] | |
| min/max Lyman-alpha cell data: | |
| (vx,vy,vz) = [(-1260.88, -1657.54, -1286.01), | |
| (1.79769e+308, 1.79769e+308, 1.79769e+308)] vth | |
| kH0 = [1.32774e-29, 1.92445e-09] 1/cm | |
| kdust = [7.9063e-32, 5.26907e-17] 1/cm | |
| T = [10.0004, 7.76786e+07] K | |
| vth = sqrt(2*kB*T/mH) | |
| = [0.406208, 1132.11] km/s | |
| DnuD = vth*nu0/c | |
| = [3.34134e+09, 9.31242e+12] Hz | |
| a = 0.5*DnuL/DnuD | |
| = [0.0148683, 5.33481e-06] | |
| g = h*DnuD/2*kB*T = 0.539537*a | |
| = [0.00801762, 2.87676e-06] | |
| sigma0 = f12*sqrt(pi)*ee^2/(me*c*DnuD) | |
| = [1.86513e-12, 6.69217e-16] cm^2 | |
| atau_NL = [0, 1.09126e+06] | |
| tau0 (LOS) = [1.79769e+308, -1.79769e+308] | |
| tau_dust = [1.79769e+308, -1.79769e+308] | |
| +------------------------------------------------------+ | |
| | Initialization complete. Starting COLT calculations. | | |
| +------------------------------------------------------+ | |
| 00:24:46 remaining | |
| 00:23:39 remaining | |
| 00:22:12 remaining | |
| 00:20:57 remaining | |
| 00:19:39 remaining | |
| 00:18:20 remaining | |
| 00:17:02 remaining | |
| 00:15:44 remaining | |
| 00:14:24 remaining | |
| 00:13:06 remaining | |
| 00:11:48 remaining | |
| 00:10:30 remaining | |
| 00:09:13 remaining | |
| 00:07:54 remaining | |
| 00:06:35 remaining | |
| slurmstepd: error: *** JOB 41034560 ON c24a-s25 CANCELLED AT 2019-09-02T12:00:45 DUE TO TIME LIMIT *** | |
| slurmstepd: error: *** STEP 41034560.0 ON c24a-s25 CANCELLED AT 2019-09-02T12:00:45 DUE TO TIME LIMIT *** | |
| srun: error: c24b-s28: task 3: Exited with exit code 1 |
Author
Unfortunately this was a long time ago, and I'm not an HPC user anymore. But here's what I can come up with.
I eventually settled on these MCA settings. I struggled mightily to figure out what they actually do, mostly I guessed around based on what the local admins told me:
export OMPI_MCA_pml="^ucx"
export OMPI_MCA_btl="self,vader,openib"
export OMPI_MCA_oob_tcp_listen_mode="listen_thread"
And I did eventually discover that I needed to provide a specific pmix version, which at one point changed based on what software was available on the cluster. What eventually worked for me:
srun --mpi=pmix_v2 /path/to/executable
Thanks,I will try your method and change pmix or slurm version to test next week, and will tell you the result.
We found the problem ,we modfiy some code from pmix like handle_timeout to enlarge, then UNREACHABLE not appeared, thanks!
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
I have this problem too, on slurm 19.05.7 with pmix ,but ,when I run more than 200+ nodes and tasks-per-node > 50 ,I also find log on slurmd with wait contrib: hostname , do you solve this problem or know this problem happened