Quick followup...

I recompiled with -g -O0 to get more debug info:

Program received signal SIGSEGV, Segmentation fault.
[Switching to Thread 800e0aac0 (LWP 100538/slurmctld)]
0x00007fffff6f9418 in ?? ()
(gdb) where
#0  0x00007fffff6f9418 in ?? ()
#1  0x0000000800bf616f in qsort () from /lib/libc.so.7
#2  0x000000080131b06e in _block_sync_core_bitmap (job_ptr=0x802033a08,
     cr_type=1) at dist_tasks.c:489
#3  0x000000080131c2cd in cr_dist (job_ptr=0x802033a08, cr_type=1)
     at dist_tasks.c:896
#4  0x00000008013223f3 in cr_job_test (job_ptr=0x802033a08,
     bitmap=0x8020413e0, min_nodes=1, max_nodes=500000, req_nodes=1, 
mode=0,
     cr_type=1, job_node_req=NODE_CR_ONE_ROW, cr_node_cnt=1,
     cr_part_ptr=0x800e0cd68, node_usage=0x800eac668, exc_core_bitmap=0x0)
     at job_test.c:2530
#5  0x0000000801316a42 in _run_now (job_ptr=0x802033a08, 
bitmap=0x8020413e0,
     min_nodes=1, max_nodes=500000, req_nodes=1, job_node_req=1,
     preemptee_candidates=0x0, preemptee_job_list=0x7fffff6f9a90,
     exc_core_bitmap=0x0) at select_cons_res.c:1498
#6  0x0000000801317e63 in select_p_job_test (job_ptr=0x802033a08,
     bitmap=0x8020413e0, min_nodes=1, max_nodes=500000, req_nodes=1, 
mode=0,
     preemptee_candidates=0x0, preemptee_job_list=0x7fffff6f9a90,
     exc_core_bitmap=0x0) at select_cons_res.c:2002
#7  0x00000000004cb6d3 in select_g_job_test ()
#8  0x000000000046160f in deallocate_nodes ()
#9  0x0000000000460b0d in deallocate_nodes ()
#10 0x00000000004621cd in select_nodes ()
---Type <return> to continue, or q <return> to quit---
#11 0x000000000043e1d8 in rehash_jobs ()
#12 0x000000000043e66e in job_allocate ()
#13 0x000000000046d439 in validate_operator ()
#14 0x000000000046b978 in slurmctld_req ()
#15 0x000000000042fec9 in main ()

On 03/28/13 20:33, Jason bacon wrote:
>
> I just ported slurm to FreeBSD, with very few issues.
>
> I've set up two simple test clusters, one on old P4 systems and one on
> AMD64 systems.
>
> I have not run into any issues on the P4 systems.
>
> On the AMD64 systems, everything seems to work fine except srun, which
> causes slurmctld to crash.
> Has anyone seen something like this before?
>
> Thanks,
>
>       Jason
>
> Controller or compute node:
>
> <<<ROOT@porpoise>>>  /home/bacon 850 # sinfo
> PARTITION         AVAIL  TIMELIMIT  NODES  STATE NODELIST
> default-partitio*    up   infinite      1   idle porpoise
> <<<ROOT@porpoise>>>  /home/bacon 851 # scontrol show partitions
> PartitionName=default-partition
>      AllocNodes=ALL AllowGroups=ALL Default=YES
>      DefaultTime=NONE DisableRootJobs=NO GraceTime=0 Hidden=NO
>      MaxNodes=UNLIMITED MaxTime=UNLIMITED MinNodes=1
>      Nodes=porpoise
>      Priority=1 RootOnly=NO ReqResv=NO Shared=NO PreemptMode=OFF
>      State=UP TotalCPUs=2 TotalNodes=1 DefMemPerNode=UNLIMITED
> MaxMemPerNode=UNLIMITED
>
> <<<ROOT@porpoise>>>  /home/bacon 852 # srun -n1 hostname
> srun: error: slurm_receive_msg: Zero Bytes were transmitted or received
> srun: error: Unable to allocate resources: Zero Bytes were transmitted
> or received
>
> Controller:
>
> GNU gdb 6.1.1 [FreeBSD]
> Copyright 2004 Free Software Foundation, Inc.
> GDB is free software, covered by the GNU General Public License, and you are
> welcome to change it and/or distribute copies of it under certain
> conditions.
> Type "show copying" to see the conditions.
> There is absolutely no warranty for GDB.  Type "show warranty" for details.
> This GDB was configured as "amd64-marcel-freebsd"...(no debugging
> symbols found)...
> (gdb) run -D
> Starting program: /usr/local/sbin/slurmctld -D
> [New LWP 100380]
> [New Thread 800e041c0 (LWP 100380/initial thread)]
> slurmctld: pidfile not locked, assuming no running daemon
> slurmctld: Accounting storage NOT INVOKED plugin loaded
> slurmctld: slurmctld version 2.5.4 started on cluster cluster
> slurmctld: Munge cryptographic signature plugin loaded
> slurmctld: Consumable Resources (CR) Node Selection plugin loaded with
> argument 1
> slurmctld: preempt/none loaded
> slurmctld: Checkpoint plugin loaded: checkpoint/none
> slurmctld: Job accounting gather NOT_INVOKED plugin loaded
> slurmctld: debug:  No backup controller to shutdown
> slurmctld: switch NONE plugin loaded
> slurmctld: debug:  Reading slurm.conf file: /usr/local/etc/slurm.conf
> slurmctld: topology NONE plugin loaded
> slurmctld: debug:  No DownNodes
> slurmctld: sched: Backfill scheduler plugin loaded
> [New Thread 800e77540 (LWP 100245/slurmctld)]
> slurmctld: Recovered state of 1 nodes
> slurmctld: Recovered information about 0 jobs
> slurmctld: cons_res: select_p_node_init
> slurmctld: cons_res: preparing for 1 partitions
> slurmctld: debug:  Updating partition uid access list
> slurmctld: Recovered state of 0 reservations
> slurmctld: State of 0 triggers recovered
> slurmctld: read_slurm_conf: backup_controller not specified.
> slurmctld: cons_res: select_p_reconfigure
> slurmctld: cons_res: select_p_node_init
> slurmctld: cons_res: preparing for 1 partitions
> slurmctld: Running as primary controller
> slurmctld: debug:  Priority BASIC plugin loaded
> [New Thread 800e77380 (LWP 100322/slurmctld)]
> [New Thread 800e771c0 (LWP 100361/slurmctld)]
> [New Thread 800e0ae40 (LWP 100397/slurmctld)]
> [New Thread 800e0ac80 (LWP 100414/slurmctld)]
> slurmctld: debug:  power_save module disabled, SuspendTime<  0
> [Thread 800e0ac80 (LWP 100414/slurmctld) exited]
> slurmctld: debug:  Spawning registration agent for porpoise 1 hosts
> [New Thread 800e0aac0 (LWP 100414/slurmctld)]
> [New Thread 80201fc80 (LWP 100509/slurmctld)]
> [New Thread 80201fac0 (LWP 100524/slurmctld)]
> [New Thread 800eb9380 (LWP 100534/slurmctld)]
> slurmctld: auth plugin for Munge (http://code.google.com/p/munge/) loaded
> [New Thread 801c54c80 (LWP 100535/slurmctld)]
> [Thread 800eb9380 (LWP 100534/slurmctld) exited]
> [Thread 80201fac0 (LWP 100524/slurmctld) exited]
> slurmctld: debug:  validate_node_specs: node porpoise registered with 0 jobs
> [Thread 801c54c80 (LWP 100535/slurmctld) exited]
> [Thread 80201fc80 (LWP 100509/slurmctld) exited]
> [Thread 800e0aac0 (LWP 100414/slurmctld) exited]
> [New Thread 800eb9380 (LWP 100414/slurmctld)]
>
> Program received signal SIGSEGV, Segmentation fault.
> [Switching to Thread 800eb9380 (LWP 100414/slurmctld)]
> 0x00007fffff3f5ee8 in ?? ()
> (gdb) where
> #0  0x00007fffff3f5ee8 in ?? ()
> #1  0x0000000800b9f16f in qsort () from /lib/libc.so.7
> #2  0x000000080131876b in _block_sync_core_bitmap ()
>      from /usr/local/lib/slurm/select_cons_res.so
> #3  0x00000008013198b7 in cr_dist ()
>      from /usr/local/lib/slurm/select_cons_res.so
> #4  0x000000080131e76b in cr_job_test ()
>      from /usr/local/lib/slurm/select_cons_res.so
> #5  0x0000000801316c6d in select_p_job_test ()
>      from /usr/local/lib/slurm/select_cons_res.so
> #6  0x00000000004536f8 in list_find_feature ()
> #7  0x00000000004540cf in list_find_feature ()
> #8  0x0000000000456cd5 in select_nodes ()
> #9  0x000000000043d80f in job_allocate ()
> #10 0x000000000045c3a0 in validate_super_user ()
> #11 0x000000000046475d in slurmctld_req ()
> #12 0x000000000042faa5 in main ()
>


-- 
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Jason W. Bacon
[email protected]
http://personalpages.tds.net/~jwbacon
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Reply via email to