如何解决sbatch error 内存规格不能满足
我想提交一个顺序作业,但我得到了:
sbatch: error: Memory specification can not be satisfied
sbatch: error: Batch job submission failed: Requested node configuration is not available
这是我的 .sh 文件:
#SBATCH --nodes=1
#SBATCH --time=01:00:00
#SBATCH --job-name=job-8-0
#SBATCH --mem=64000mb
#SBATCH --exclusive
module purge
module load gcc-8.3.0-gcc-4.8.5-tu6ftrf
echo "Starting job-8-0"
echo "Starting at `date`"
cd code
srun gcc -Wno-return-type file1.cpp file2.cpp file3.cpp file4.cpp file5.cpp main.cpp -o myExperiment -lstdc++ -lm
srun ./myExperiment 8 0
echo "Experiment 8-0 finished with exit code $? at: `date`"
节点login01信息为:
NodeName=login01 Arch=x86_64 CoresPerSocket=8
CPUAlloc=0 CPUErr=0 CPUTot=32 CPULoad=29.91
AvailableFeatures=(null)
ActiveFeatures=(null)
Gres=gpu:8
NodeAddr=10.0.50.0 NodeHostName=login01 Version=17.11
OS=Linux 3.10.0-957.el7.x86_64 #1 SMP Thu Nov 8 23:39:32 UTC 2018
RealMemory=1 AllocMem=0 FreeMem=65001 Sockets=2 Boards=1
State=IDLE+DRAIN ThreadsPerCore=2 TmpDisk=0 Weight=1 Owner=N/A MCS_label=N/A
BootTime=2021-05-25T13:13:10 SlurmdStartTime=2021-05-25T16:35:31
CfgTRES=cpu=32,mem=1M,billing=32
AllocTRES=
CapWatts=n/a
CurrentWatts=0 LowestJoules=0 ConsumedJoules=0
ExtSensorsJoules=n/s ExtSensorsWatts=0 ExtSensorsTemp=n/s
Reason=gres/gpu count too low (0 < 8) [slurm@2021-06-28T13:38:40]
还有其他节点的 FreeMem=122000 或 121000,...等超过 64000mb
这是超级计算机的规格:
• 操作系统:Linux CentOS 7
• 300 个计算节点
• 每个节点有:
o 2 个 CPU:至强 E5-2650 8 核 2.000GHz(共 16 核)
o 2 个双 AMD FirePro S10000 GPU
o 内存:128 GB 内存
• 调度程序:Slurm
当我打开 节点.conf
NodeName=login01 NodeAddr=10.0.50.0 CPUs=32 Procs=32 Sockets=2 CoresPerSocket=8 ThreadsPerCore=2 State=IDLE
对于所有节点都是一样的。
这是slurm.conf
ClusterName=sanam
ControlMachine=mgmt01
ControlAddr=10.0.1.254
SlurmUser=slurm
SlurmdUser=root
SlurmctldPort=6817
SlurmdPort=6818
AuthType=auth/munge
StateSaveLocation=/var/spool/slurm/ctld
SlurmdSpoolDir=/var/spool/slurmd
SwitchType=switch/none
MpiDefault=none
SlurmctldPidFile=/var/run/slurmctld.pid
SlurmdPidFile=/var/run/slurmd.pid
ProctrackType=proctrack/pgid
ReturnToService=2
GresTypes=gpu
# TIMERS
SlurmctldTimeout=300
SlurmdTimeout=300
InactiveLimit=0
MinJobAge=300
KillWait=30
Waittime=0
SelectType=select/cons_res
SelectTypeParameters=CR_Core
# SCHEDULING
SchedulerType=sched/backfill
FastSchedule=1
# LOGGING AND ACCOUNTING
AccountingStorageType=accounting_storage/slurmdbd
JobAcctGatherFrequency=30
JobAcctGatherType=jobacct_gather/linux
SlurmctldDebug=3
SlurmctldLogFile=/var/log/slurm/slurmctld.log
SlurmdDebug=3
SlurmdLogFile=/var/log/slurm/slurmd.log
JobCompType=jobcomp/none
AccountingStorageHost=10.0.1.254
include /etc/slurm/nodes.conf
include /etc/slurm/partitions.conf
#include /etc/slurm/gres.conf
“内存规格不能满足”是什么原因造成的?
我应该用 srun 指定 RAM 和 CPU 吗?特别是运行实验的第二个 srun?
版权声明:本文内容由互联网用户自发贡献,该文观点与技术仅代表作者本人。本站仅提供信息存储空间服务,不拥有所有权,不承担相关法律责任。如发现本站有涉嫌侵权/违法违规的内容, 请发送邮件至 dio@foxmail.com 举报,一经查实,本站将立刻删除。