[Mesa-users] System crashes when evolving beyond zams
Francis Timmes
fxt44 at mac.com
Wed Apr 21 04:26:25 UTC 2021
hi yash,
all these symptoms in this thread point toward an environment
that is too small for mesa to run; memory or disk or both.
if so, there is little we can do, and the only solution is a
run-time environment with more resources.
fxt
> On Apr 20, 2021, at 3:05 AM, YASH MEHUL MEHTA via Mesa-users <mesa-users at lists.mesastar.org> wrote:
>
>
>
> Yash Mehta
> Undergrad, IISc
> From: YASH MEHUL MEHTA <yashmehta at iisc.ac.in>
> Sent: Monday, April 19, 2021 10:57:01 PM
> To: Warrick Ball <W.H.Ball at bham.ac.uk>
> Subject: Re: [Mesa-users] System crashes when evolving beyond zams
>
> Hi
> I somehow managed so the system wouldn't shut down when the error came. I am attaching the files used to run the program (I used ./re x850, let me know if I should send the photos as well. Here is the error message:
>
> __________________________________________________________________________________________________________________________________________________
>
> step lg_Tmax Teff lg_LH lg_Lnuc Mass H_rich H_cntr N_cntr Y_surf eta_cntr zones retry
> lg_dt_yr lg_Tcntr lg_R lg_L3a lg_Lneu lg_Mdot He_core He_cntr O_cntr Z_surf gam_cntr iters
> age_yr lg_Dcntr lg_L lg_LZ lg_Lphoto lg_Dsurf C_core C_cntr Ne_cntr Z_cntr v_div_cs dt_limit
> __________________________________________________________________________________________________________________________________________________
>
> 980 7.665227 2.924E+04 4.027162 4.027162 10.000000 10.000000 0.243803 0.000064 0.240200 -5.153846 859 0
> 5.175202 7.665227 0.603939 -20.328974 2.864231 -99.000000 0.000000 0.756100 0.000002 0.000100 0.023529 5
> 1.7196E+07 1.375769 4.026439 -99.000000 -99.000000 -8.819006 0.000000 0.000001 0.000010 0.000097 0.000E+00 varcontrol
>
> 981 7.666598 2.917E+04 4.030124 4.030124 10.000000 10.000000 0.235612 0.000064 0.240200 -5.157652 861 0
> 5.164296 7.666598 0.607525 -20.252297 2.866960 -99.000000 0.000000 0.764291 0.000002 0.000100 0.023752 5
> 1.7342E+07 1.379027 4.029394 -99.000000 -99.000000 -8.823396 0.000000 0.000001 0.000010 0.000097 0.000E+00 varcontrol
>
> 982 7.668002 2.910E+04 4.033056 4.033056 10.000000 10.000000 0.227417 0.000064 0.240200 -5.161262 861 0
> 5.154129 7.668002 0.611115 -20.174214 2.869667 -99.000000 0.000000 0.772486 0.000001 0.000100 0.023983 5
> 1.7484E+07 1.382443 4.032316 -99.000000 -99.000000 -8.827795 0.000000 0.000001 0.000010 0.000097 0.000E+00 varcontrol
>
> 983 7.669419 2.903E+04 4.035955 4.035955 10.000000 10.000000 0.219358 0.000064 0.240200 -5.164685 864 0
> 5.143876 7.669419 0.614695 -20.095895 2.872353 -99.000000 0.000000 0.780545 0.000001 0.000100 0.024217 5
> 1.7624E+07 1.385932 4.035210 -99.000000 -99.000000 -8.832186 0.000000 0.000001 0.000010 0.000097 0.000E+00 varcontrol
>
> 984 7.670875 2.895E+04 4.038818 4.038818 10.000000 10.000000 0.211292 0.000064 0.240200 -5.167960 863 0
> 5.133771 7.670875 0.618271 -20.015968 2.875011 -99.000000 0.000000 0.788611 0.000001 0.000100 0.024459 5
> 1.7760E+07 1.389568 4.038064 -11.439140 -99.000000 -8.836579 0.000000 0.000001 0.000010 0.000097 0.000E+00 varcontrol
>
> 985 7.672349 2.888E+04 4.041652 4.041652 10.000000 10.000000 0.203351 0.000064 0.240200 -5.171046 863 0
> 5.124112 7.672349 0.621839 -19.935494 2.877651 -99.000000 0.000000 0.796552 0.000001 0.000100 0.024705 5
> 1.7893E+07 1.393291 4.040894 -11.439140 -99.000000 -8.840967 0.000000 0.000001 0.000010 0.000097 0.000E+00 varcontrol
>
> Load1_eosDT_Table ierr 5014 1 2 NaN NaN
> Load1_eosDT_Table ierr 5014 1 1 NaN NaN
> Load1_eosDT_Table ierr 5014 1 2 NaN NaN
> failed while reading /home/yash/Downloads/mesa-r15140/data/eosFreeEOS_data/mesa-FreeEOS_02z00x.data
>
> ... snip
>
> From: Warrick Ball <W.H.Ball at bham.ac.uk>
> Sent: 17 April 2021 02:54
> To: YASH MEHUL MEHTA <yashmehta at iisc.ac.in>
> Cc: mesa-users at lists.mesastar.org <mesa-users at lists.mesastar.org>
> Subject: Re: [Mesa-users] System crashes when evolving beyond zams
>
> External Email
>
>
> Hi again,
>
> > I have made absolutely no changes to the default install except setting z=0.0001 and setting terminating-at-zams to false (and later, introducing garbage collection).
>
> The reason we ask for inlists is so that it's completely clear what you have changed. For example, when you "changed z", did you just change `initial_z` or did you also change `Zbase` in the `&kap` namelist? I've run both anyway and it doesn't matter but this is the kind of detail that makes it easier to make sure that we're *exactly* reproducing what you've done.
>
> That aside, I don't think this is a RAM problem unless your VM is running something else that's using ~4.3 GB of RAM. I created a fresh work directory with r15140 and changed `initial_z` to 0.0001 (and `Zbase` to 0.0001, but that didn't matter) and deleted the ZAMS stopping condition
>
> ! stop when the star nears ZAMS (Lnuc/L > 0.99)
> Lnuc_div_L_zams_limit = 0.99d0
> stop_near_zams = .true.
>
> The run finished after 1066 steps because it hit the lower limit on central H. Based on the results from `top` and `time -v`, it peaked at ~2.7 GB of memory, which is much less than the 7 GB your VM has.
>
> I don't use VMs myself so probably can't help you any further. Sorry!
>
> Cheers,
> Warrick
>
>
> ___________
>
> Warrick Ball
> Postdoc, School of Physics and Astronomy
> University of Birmingham, Edgbaston, Birmingham B15 2TT
> W.H.Ball at bham.ac.uk
> +44 (0)121 414 4552
>
> On Fri, 16 Apr 2021, yashmehta at iisc.ac.in wrote:
>
> > So I did a bit of experimentation:
> > Running with ./rn and with ./re 1000 both results in the system crashing exactly at step 1011. This is true with and without garbage collection.
> > Interestingly, I was able to achieve step 1023 exactly once, after restarting my laptop, force closing all tasks, and maximizing RAM (this run was without garbage collection, and was done using ./re 1000). But I was not able to reproduce this despite
> > using the same setup multiple times.
> >
> > I have made absolutely no changes to the default install except setting z=0.0001 and setting terminating-at-zams to false (and later, introducing garbage collection).
> > There is no error message. Either the virtual machine simply crashes, or the entire laptop shuts down. I have never faced this problem in the past despite some heavy computation.
> >
> > __________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________
> > From: Warrick Ball <W.H.Ball at bham.ac.uk>
> > Sent: 16 April 2021 20:31
> > To: YASH MEHUL MEHTA <yashmehta at iisc.ac.in>
> > Cc: mesa-users at lists.mesastar.org <mesa-users at lists.mesastar.org>
> > Subject: Re: [Mesa-users] System crashes when evolving beyond zams
> > External Email
> >
> >
> > Hi,
> >
> > If it is a lack of RAM, you can start by restarting your run from a photo. If you haven't changed the options, the `photos/` directory will contain files that are either `x` followed by a 3-digit number or just a 4-digit number. e.g. `x050` or
> > `1000`. You can try restarting from one of these towards the end of your run, before the crash. You say you reached model 1011, so try
> >
> > ./re x950
> >
> > or perhaps
> >
> > ./re 1000
> >
> > If that gets further than the original run with `./rn`, it's probably because you don't have enough RAM. To keep your runs going, you can try using the "garbage collector" by adding e.g.
> >
> > num_steps_for_garbage_collection = 100 ! every 100 steps; you might need to adjust this
> >
> > to `&star_job` and perhaps
> >
> > report_garbage_collection = .true. ! instead of the default .false.
> >
> > to see that the garbage collection is working.
> >
> > The garbage collector works by de-allocating the EoS handles. Those that are still needed for the evolution are automatically reloaded, so in effect this drops those handles that aren't still required, thereby freeing some memory.
> >
> > This isn't guaranteed to work. There might be a situation where the model simply needs more EoS tables than you have RAM, but this might let you squeeze by.
> >
> > All that said, if the model still crashes at the same point after a restart, it might not be the RAM and you need to send more information: inlists to reproduce the error and the full error message.
> >
> > Cheers,
> > Warrick
> >
> > ___________
> >
> > Warrick Ball
> > Postdoc, School of Physics and Astronomy
> > University of Birmingham, Edgbaston, Birmingham B15 2TT
> > W.H.Ball at bham.ac.uk
> > +44 (0)121 414 4552
> >
> > On Fri, 16 Apr 2021, mesa-users at lists.mesastar.org wrote:
> >
> > > Hi
> > > My name is Yash Mehta, I am running mesa-r15140 on my virtual machine with 7000MB RAM. (I am running ubuntu16.04 on my windows10 home, 8gb Dell XPS 15)
> > > Evolving a star upto zams is causing no issues, but when trying to evolve a 15M_sun star with initial z=0.0001 beyond zams, the system crashes somewhere midway (typically after 1011th iteration; zams is reached at 938th iteration)
> > > Right before the system crashes, the terminal shows (after the data about the 1011th iteration):
> > > "write /home/yash/Downloads/mesa-r15140/data/eosDT_data/cache/mesa-FreeEOS_00z10x.bin"
> > > "write /home/yash/Downloads/mesa-r15140/data/eosDT_data/cache/mesa-FreeEOS_02z10x.bin"
> > > Is the issue simply lack of RAM or could it be something else? Also, is there a way out?
> > >
> > >
> >
> >
> <restart_photo><rn><mk><README.rst><re><inlist_pgstar><inlist_project><inlist><clean>_______________________________________________
> mesa-users at lists.mesastar.org
> https://lists.mesastar.org/mailman/listinfo/mesa-users
More information about the Mesa-users
mailing list