[Mesa-users] System crashes when evolving beyond zams

Francis Timmes fxt44 at mac.com
Wed Apr 21 04:26:25 UTC 2021


hi yash,

all these symptoms in this thread point toward an environment 
that is too small for mesa to run; memory or disk or both. 

if so, there is little we can do, and the only solution is a 
run-time environment with more resources.

fxt





> On Apr 20, 2021, at 3:05 AM, YASH MEHUL MEHTA via Mesa-users <mesa-users at lists.mesastar.org> wrote:
> 
> 
> 
> Yash Mehta
> Undergrad, IISc
> From: YASH MEHUL MEHTA <yashmehta at iisc.ac.in>
> Sent: Monday, April 19, 2021 10:57:01 PM
> To: Warrick Ball <W.H.Ball at bham.ac.uk>
> Subject: Re: [Mesa-users] System crashes when evolving beyond zams
>  
> Hi
> I somehow managed so the system wouldn't shut down when the error came. I am attaching the files used to run the program (I used ./re x850, let me know if I should send the photos as well. Here is the error message:
> 
> __________________________________________________________________________________________________________________________________________________
> 
>        step    lg_Tmax     Teff     lg_LH      lg_Lnuc     Mass       H_rich     H_cntr     N_cntr     Y_surf   eta_cntr   zones  retry
>    lg_dt_yr    lg_Tcntr    lg_R     lg_L3a     lg_Lneu     lg_Mdot    He_core    He_cntr    O_cntr     Z_surf   gam_cntr   iters  
>      age_yr    lg_Dcntr    lg_L     lg_LZ      lg_Lphoto   lg_Dsurf   C_core     C_cntr     Ne_cntr    Z_cntr   v_div_cs       dt_limit
> __________________________________________________________________________________________________________________________________________________
> 
>         980   7.665227  2.924E+04   4.027162   4.027162  10.000000  10.000000   0.243803   0.000064   0.240200  -5.153846    859      0
>    5.175202   7.665227   0.603939 -20.328974   2.864231 -99.000000   0.000000   0.756100   0.000002   0.000100   0.023529      5
>  1.7196E+07   1.375769   4.026439 -99.000000 -99.000000  -8.819006   0.000000   0.000001   0.000010   0.000097  0.000E+00    varcontrol
> 
>         981   7.666598  2.917E+04   4.030124   4.030124  10.000000  10.000000   0.235612   0.000064   0.240200  -5.157652    861      0
>    5.164296   7.666598   0.607525 -20.252297   2.866960 -99.000000   0.000000   0.764291   0.000002   0.000100   0.023752      5
>  1.7342E+07   1.379027   4.029394 -99.000000 -99.000000  -8.823396   0.000000   0.000001   0.000010   0.000097  0.000E+00    varcontrol
> 
>         982   7.668002  2.910E+04   4.033056   4.033056  10.000000  10.000000   0.227417   0.000064   0.240200  -5.161262    861      0
>    5.154129   7.668002   0.611115 -20.174214   2.869667 -99.000000   0.000000   0.772486   0.000001   0.000100   0.023983      5
>  1.7484E+07   1.382443   4.032316 -99.000000 -99.000000  -8.827795   0.000000   0.000001   0.000010   0.000097  0.000E+00    varcontrol
> 
>         983   7.669419  2.903E+04   4.035955   4.035955  10.000000  10.000000   0.219358   0.000064   0.240200  -5.164685    864      0
>    5.143876   7.669419   0.614695 -20.095895   2.872353 -99.000000   0.000000   0.780545   0.000001   0.000100   0.024217      5
>  1.7624E+07   1.385932   4.035210 -99.000000 -99.000000  -8.832186   0.000000   0.000001   0.000010   0.000097  0.000E+00    varcontrol
> 
>         984   7.670875  2.895E+04   4.038818   4.038818  10.000000  10.000000   0.211292   0.000064   0.240200  -5.167960    863      0
>    5.133771   7.670875   0.618271 -20.015968   2.875011 -99.000000   0.000000   0.788611   0.000001   0.000100   0.024459      5
>  1.7760E+07   1.389568   4.038064 -11.439140 -99.000000  -8.836579   0.000000   0.000001   0.000010   0.000097  0.000E+00    varcontrol
> 
>         985   7.672349  2.888E+04   4.041652   4.041652  10.000000  10.000000   0.203351   0.000064   0.240200  -5.171046    863      0
>    5.124112   7.672349   0.621839 -19.935494   2.877651 -99.000000   0.000000   0.796552   0.000001   0.000100   0.024705      5
>  1.7893E+07   1.393291   4.040894 -11.439140 -99.000000  -8.840967   0.000000   0.000001   0.000010   0.000097  0.000E+00    varcontrol
> 
>  Load1_eosDT_Table ierr        5014           1           2                       NaN                       NaN
>  Load1_eosDT_Table ierr        5014           1           1                       NaN                       NaN
>  Load1_eosDT_Table ierr        5014           1           2                       NaN                       NaN
> failed while reading /home/yash/Downloads/mesa-r15140/data/eosFreeEOS_data/mesa-FreeEOS_02z00x.data
> 
> ... snip
> 
> From: Warrick Ball <W.H.Ball at bham.ac.uk>
> Sent: 17 April 2021 02:54
> To: YASH MEHUL MEHTA <yashmehta at iisc.ac.in>
> Cc: mesa-users at lists.mesastar.org <mesa-users at lists.mesastar.org>
> Subject: Re: [Mesa-users] System crashes when evolving beyond zams
>  
> External Email
> 
> 
> Hi again,
> 
> > I have made absolutely no changes to the default install except setting z=0.0001 and setting terminating-at-zams to false (and later, introducing garbage collection).
> 
> The reason we ask for inlists is so that it's completely clear what you have changed.  For example, when you "changed z", did you just change `initial_z` or did you also change `Zbase` in the `&kap` namelist?  I've run both anyway and it doesn't matter but this is the kind of detail that makes it easier to make sure that we're *exactly* reproducing what you've done.
> 
> That aside, I don't think this is a RAM problem unless your VM is running something else that's using ~4.3 GB of RAM.  I created a fresh work directory with r15140 and changed `initial_z` to 0.0001 (and `Zbase` to 0.0001, but that didn't matter) and deleted the ZAMS stopping condition
> 
>    ! stop when the star nears ZAMS (Lnuc/L > 0.99)
>      Lnuc_div_L_zams_limit = 0.99d0
>      stop_near_zams = .true.
> 
> The run finished after 1066 steps because it hit the lower limit on central H.  Based on the results from `top` and `time -v`, it peaked at ~2.7 GB of memory, which is much less than the 7 GB your VM has.
> 
> I don't use VMs myself so probably can't help you any further.  Sorry!
> 
> Cheers,
> Warrick
> 
> 
> ___________
> 
> Warrick Ball
> Postdoc, School of Physics and Astronomy
> University of Birmingham, Edgbaston, Birmingham B15 2TT
> W.H.Ball at bham.ac.uk
> +44 (0)121 414 4552
> 
> On Fri, 16 Apr 2021, yashmehta at iisc.ac.in wrote:
> 
> > So I did a bit of experimentation:
> > Running with ./rn and with ./re 1000 both results in the system crashing exactly at step 1011. This is true with and without garbage collection.
> > Interestingly, I was able to achieve step 1023 exactly once, after restarting my laptop, force closing all tasks, and maximizing RAM (this run was without garbage collection, and was done using ./re 1000). But I was not able to reproduce this despite
> > using the same setup multiple times.
> >
> > I have made absolutely no changes to the default install except setting z=0.0001 and setting terminating-at-zams to false (and later, introducing garbage collection).
> > There is no error message. Either the virtual machine simply crashes, or the entire laptop shuts down. I have never faced this problem in the past despite some heavy computation.
> >
> > __________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________________
> > From: Warrick Ball <W.H.Ball at bham.ac.uk>
> > Sent: 16 April 2021 20:31
> > To: YASH MEHUL MEHTA <yashmehta at iisc.ac.in>
> > Cc: mesa-users at lists.mesastar.org <mesa-users at lists.mesastar.org>
> > Subject: Re: [Mesa-users] System crashes when evolving beyond zams
> > External Email
> >
> >
> > Hi,
> >
> > If it is a lack of RAM, you can start by restarting your run from a photo.  If you haven't changed the options, the `photos/` directory will contain files that are either `x` followed by a 3-digit number or just a 4-digit number. e.g. `x050` or
> > `1000`.  You can try restarting from one of these towards the end of your run, before the crash.  You say you reached model 1011, so try
> >
> >      ./re x950
> >
> > or perhaps
> >
> >      ./re 1000
> >
> > If that gets further than the original run with `./rn`, it's probably because you don't have enough RAM.  To keep your runs going, you can try using the "garbage collector" by adding e.g.
> >
> >      num_steps_for_garbage_collection = 100 ! every 100 steps; you might need to adjust this
> >
> > to `&star_job` and perhaps
> >
> >      report_garbage_collection = .true. ! instead of the default .false.
> >
> > to see that the garbage collection is working.
> >
> > The garbage collector works by de-allocating the EoS handles.  Those that are still needed for the evolution are automatically reloaded, so in effect this drops those handles that aren't still required, thereby freeing some memory.
> >
> > This isn't guaranteed to work.  There might be a situation where the model simply needs more EoS tables than you have RAM, but this might let you squeeze by.
> >
> > All that said, if the model still crashes at the same point after a restart, it might not be the RAM and you need to send more information: inlists to reproduce the error and the full error message.
> >
> > Cheers,
> > Warrick
> >
> > ___________
> >
> > Warrick Ball
> > Postdoc, School of Physics and Astronomy
> > University of Birmingham, Edgbaston, Birmingham B15 2TT
> > W.H.Ball at bham.ac.uk
> > +44 (0)121 414 4552
> >
> > On Fri, 16 Apr 2021, mesa-users at lists.mesastar.org wrote:
> >
> > > Hi
> > > My name is Yash Mehta, I am running mesa-r15140 on my virtual machine with 7000MB RAM. (I am running ubuntu16.04 on my windows10 home, 8gb Dell XPS 15)
> > > Evolving a star upto zams is causing no issues, but when trying to evolve a 15M_sun star with initial z=0.0001 beyond zams, the system crashes somewhere midway (typically after 1011th iteration; zams is reached at 938th iteration)
> > > Right before the system crashes, the terminal shows (after the data about the 1011th iteration):
> > > "write /home/yash/Downloads/mesa-r15140/data/eosDT_data/cache/mesa-FreeEOS_00z10x.bin"
> > > "write /home/yash/Downloads/mesa-r15140/data/eosDT_data/cache/mesa-FreeEOS_02z10x.bin"
> > > Is the issue simply lack of RAM or could it be something else? Also, is there a way out?
> > >
> > >
> >
> >
> <restart_photo><rn><mk><README.rst><re><inlist_pgstar><inlist_project><inlist><clean>_______________________________________________
> mesa-users at lists.mesastar.org
> https://lists.mesastar.org/mailman/listinfo/mesa-users



More information about the Mesa-users mailing list