Commits · d2e38a2b51057bf168b0cb9d2e5c22794b7aae93 · tud-zih-energy / Slurm

Mar 21, 2016

Change point where burst buffer env vars are set · 54f314e7

Morris Jette authored 9 years ago

burst_buffer/cray: Set environment variables just before starting job rather
    than at job submission time to reflect persistent buffers created or
    modified while the job is pending.
bug 2545

54f314e7

Fix deadlock issue with burst_buffer/cray when a newly created burst · dcfa6ec0

Danny Auble authored 9 years ago

buffer is found.

Bug 2576

What happened was a function was doing a double read lock which isn't
awesome to begin with, but not really horrible (if all you are doing is
read locks anyway).  The problem was after the first lock was locked a
different thread was going for a write lock and so when the second
read lock came in it created deadlocked.

dcfa6ec0

Mar 18, 2016

Fix for srun abort on SIGSTOP+SIGCONT · 1ed38f26

Morris Jette authored 9 years ago

Avoid possibly aborting srun that gets simultaneous SIGSTOP+SIGCONT while
    creating the job step. The result is that the signal hanlder gets a
    argument (the signal received) of zero.

Here's a log, window 1:
$ srun hostname
srun: Job step creation temporarily disabled, retrying
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 18
srun: I Got signal 0
srun: Cancelled pending job step

Window 2:
$  kill -STOP 18696 ; kill -CONT 18696
$  kill -STOP 18696 ; kill -CONT 18696
$  kill -STOP 18696 ; kill -CONT 18696
....

bug 2494

1ed38f26

Mar 17, 2016

Prevent uid update from corrupting assoc_hash table. · 60b58b70

Tim Wickberg authored 9 years ago

The uid is used as part of the hash function, must remove old reference
and recalculate if it may change, otherwise _delete_assoc_hash
will not find it again when the association is removed, causing
slurmctld to segfault.

Bug 2560.

60b58b70

Mar 16, 2016

Update gang scheduling data structures when job changes in size · 701917cc

Morris Jette authored 9 years ago

Previous gang scheduling logic maintained information about resources
  originally allocated to the job and made scheduling decisions on
  that basis.
bug 2494

701917cc

gang scheduling for with manually job suspend/resume · 344d2eab

Morris Jette authored 9 years ago

Update gang scheduling table when job manually suspended or resumed. Prior
    logic could mess up job suspend/resume sequencing.
bug 2494

344d2eab

Fix issue when adding a new TRES to AccountingStorageTRES for the first · 6c436e34

Danny Auble authored 9 years ago

time.

https://bugs.schedmd.com/show_bug.cgi?id=2547

The code just wasn't fully baked before and was probably written before
a lot of the other supporting code was done i.e
assoc_mgr_set_assoc|qos_tres_cnt were done specifically for this kind of
thing.  Many of the usage structures weren't realloced either as well as
the tres_cnt local to each qos and assoc wasn't updated.  So all in all
pretty bad code - bad Danny.  This makes sure all this sets up and no
memory corruption happens.

6c436e34

Send burst buffer teardown immediately · d85cdcc7

Morris Jette authored 9 years ago

Generate burst buffer use completion email immediately afer teardown
    completes rather than at job purge time (likely minutes later).
bug 2539

d85cdcc7

Modify burst buffer stage out message · fae4c3d3

Morris Jette authored 9 years ago

Change burst buffer use completion message from
"SLURM Job_id=1360353 Name=tmp Staged Out, StageOut time 00:01:47" to
"SLURM Job_id=1360353 Name=tmp StageOut/Teardown time 00:01:47"

fae4c3d3

Mar 15, 2016
- acct_gather_energy/ipmi - add threshold for message logging · 18608974
  Alejandro Sanchez authored 9 years ago
  
  18608974
- Check that bb_state.tres_pos is set correctly to avoid overwriting CPU TRES. · 5708037d
  Tim Wickberg authored 9 years ago
  
  Bug 2543.
  5708037d
Mar 14, 2016
- Add option for TopologyParam=NoInAddrAnyCtld to make the slurmctld listen · 775c46de
  Danny Auble authored 9 years ago
  
  on only one port like TopologyParam=NoInAddrAny does for everything else.
  775c46de
- FreeBSD - set_oom_adj is Linux-specific, stub out to avoid errors. · b3f2359f
  Tim Wickberg authored 9 years ago
  
  There's no /proc on *BSD, and BSD handles OOM in a completely different way.
  b3f2359f
Mar 11, 2016

Fix job array step function printout. · 03d29e24

Tim Wickberg authored 9 years ago

Return [0-100:2] formatting, rather than [0,2,4,6,8,...] when using
a step function.

Was inadvertantly broken in 14.11 with commit 5ffdca92.

Bug 2535.

03d29e24

Mar 10, 2016
- Add NEWS for commit 3bb2e602 · a0be0dc5
  Morris Jette authored 9 years ago
  
  a0be0dc5
Mar 09, 2016

cray job requeue bug · fec5e03b

Morris Jette authored 9 years ago

Fix Cray NHC spawning on job requeue. Previous logic would leave nodes
allocated to a requeued job as non-usable on job termination.

Specifically, each job has a "cleaning/cleaned" flag. Once a job
terminates, the cleaning flag is set, then after the job node health
check completes, the value gets set to cleaned. If the job is requeued,
on its second (or subsequent) termination, the select/cray plugin
is called to launch the NHC. The plugin sees the "cleaned" flag
already set, it then logs:
error: select_p_job_fini: Cleaned flag already set for job 1283858, this should never happen
and returns, never launching the NHC. Since the termination of the
job NHC triggers releasing job resources (CPUs, memory, and GRES),
those resources are never released for use by other jobs.

Bug 2384

fec5e03b

Correctly parse nids in slurmconfgen_smw.py · 88ccc111

David Gloe authored 9 years ago

An error in slurmconfgen_smw.py caused it to parse the nic as the nid.
On some systems those values differ, causing the generated slurm.conf file to
be incorrect.

Bug 2532.

88ccc111

Mar 08, 2016

Fix route/topology plugin to prevent segfault in sbcast. · 897c4b27

Bill Brophy authored 9 years ago

route_p_split_hostlist was not thread-safe, and would cause
one of several segfaults depending on where in the initialization
code each thread was.

Bug 2495.

897c4b27

Fix displayed value for RoutePlugin. · 14c51e65
Tim Wickberg authored 9 years ago
```
Was incorrectly displaying "(null)" even when loaded successfully.
```
14c51e65

Mar 05, 2016
- Fixed double read lock on getting job's gres/tres. · b23a57cf
  Danny Auble authored 9 years ago
  
  b23a57cf
Mar 04, 2016
- Fix issue where steps weren't always getting the gres/tres involved. · b294f81b
  Danny Auble authored 9 years ago
  
  b294f81b
Mar 03, 2016

Fix issue with sbcast not doing a correct fanout. · 72f13426
Danny Auble authored 9 years ago

72f13426
Fix getting reservations to database when database is down. · 5c43d754
Brian Christiansen authored 9 years ago
```
Bug 2507
```
5c43d754

Increase step GRES variable size · 7f0bdc84

Morris Jette authored 9 years ago

Step GRES value changed from type "int" to "int64_t" to support larger
values. Previous logic could fail in step allocation values over 32-bits.
Other GRES values are 64-bit.

7f0bdc84

Force close on exec on first 256 file descriptors when launching a · f502f1e5

Danny Auble authored 9 years ago

slurmstepd to close potential open ones.

It was pointed out the slurmd using acct_gather_energy/ipmi links to
freeipmi which could possibly open /dev/ipmi0 without the close on exec
flag set as root while launching a step leaving it open in the users app.

What this does is sets the flag on the first 256 to mitigate the concern.

Reported by Maksym Planeta.

Bug 2506

f502f1e5

Mar 02, 2016

Backfill scheduler to validate correct job partition · efd9d35e

Gary B Skouson authored 9 years ago

Previous logic tested whatever the job's partition pointer indicated
rather than the partition we are trying to run the job in. This bug
was introduced in Slurm version 15.08.5, Nov 16, 2015, commit
94f0e948
bug 2499

efd9d35e

Remove a duplicate xmalloc · 2d5066e7
Thomas Cadeau authored 9 years ago

2d5066e7

Mar 01, 2016

Update NEWS as well. · a058ff4a
Tim Wickberg authored 9 years ago

a058ff4a

Defer suspend until launch completes · 52fe3de1

Morris Jette authored 9 years ago

Insure that a job is completely launched before trying to suspend it.
Previous logic would start suspend logic early in the life of the
slurmstepd process, after it's listening socket was open but before
the tasks were launched. This defers the suspend logic until after
all prologs and setup completes and the tasks are launched. This is
important in the case of gang scheduling, in which newly launched
jobs can be immediately suspended.
bug 2494

52fe3de1

Feb 26, 2016
- Set correct reason when a QOS' MaxTresMins is violated. · 745568f2
  Danny Auble authored 9 years ago
  
  745568f2
- Add not to slurm.conf man page about SallocDefaultCommand and TaskPlugins. · b5b349b0
  Tim Wickberg authored 9 years ago
  
  Add note to slurm.conf man page about setting "--cpu_bind=no" as part of SallocDefaultCommand if a TaskPlugin is in use.
  b5b349b0
Feb 25, 2016
- Fix issue where SocketsPerBoard didn't translate to Sockets when CPUS= · fcae2193
  Danny Auble authored 9 years ago
  
  was also given.
  fcae2193
Feb 24, 2016
- Make it so scontrol update part qos= will take away a partition QOS from · 3a7470ae
  Danny Auble authored 9 years ago
  
  a partition.
  3a7470ae
- Make it possible to change CPUsPerTask with scontrol. · de28c13a
  Danny Auble authored 9 years ago
  
  This also reverts most of commit fa331e30 as well as commit bd9fa830 which would try to set the pn_min_cpus every time a job was updated. If a job didn't request node counts then they were hosed. This commit takes away the magic which was screwing things up. Now the person gets what they asked for without magic changing things. Bug 2302 Bug 2742 Bug 2478
  de28c13a
- Fix issue where when updating a job the pn_min_cpus was updated · bd9fa830
  Danny Auble authored 9 years ago
  
  erroneously.
  bd9fa830
- BGQ - Tighter locks around structures when nodes/cables change state. · c5925f41
  Danny Auble authored 9 years ago
  
  c5925f41
- BGQ - Remove redeclaration of job_read_lock. · fd3dedda
  Danny Auble authored 9 years ago
  
  fd3dedda
Feb 23, 2016

Fix issue with resizing jobs and limits not be kept track of correctly. · 92ac0dcd

Danny Auble authored 9 years ago

This whole process could probably be done better by keeping track of
old values and new values and only calling one function instead of a
pre and post function, but that can probably wait for future generations
of the code as it works now and is probably adequate for the time being.

Bug 2352

92ac0dcd

Feb 19, 2016

BurstBuffer/cray pre-run race condtition fix · e8959ae9

Morris Jette authored 9 years ago

BurstBuffer/cray - Defer job cancellation or time limit while "pre-run"
    operation in progress to avoid inconsistent state due to multiple calls
    to job termination functions.
bug 2454

e8959ae9

Update NEWS for start of v15.08.9 · 4aa2e3c2
Morris Jette authored 9 years ago

4aa2e3c2