-
Once you’ve logged into the head node of Cloud-Sandbox, go into the
/save/ec2-userdirectory and create a new directory based on your affiliated organization. (Ex.mkdir USER_AFFILIATION). Go into the newly made directory for the next step. -
Clone the latest Cloud-Sandbox code repository off the GitHub repository:
git clone https://github.com/ioos/Cloud-Sandbox.git- Create a job file for your model. The keywords in this job file will inform the rest of the workflow. Change directory into
Cloud-Sandbox/cloudflow/job/jobsCopy the cloud_sandbox_experiment.template file into a new file called your_model_name.exp. Inside that new file, you must specify at minimum the:
- MODEL (model name/class)
- JOBTYPE (model job type, be explicit on the workflow name for the given model class if it's seperate from the standard implementation class of 'experiment')
- APP (model application, if its a simple job run, then use 'basic')
- EXEC (executable)
- any executable dependencies as a new variable (such as an input file)
- MODEL_DIR (path to model directory on the Cloud-Sandbox)
- Create an AWS cluster configuration file for your model. Change directory into
Cloud-Sandbox/cloudflow/cluster/config Create a new directory based on your affiliated organization. (Ex. mkdir USER_AFFILIATION). Copy template.ioos into a new file called your_model_name.config to specify the AWS cloud resource configuration you would like your model to run on. Move the new file your_model_name.config in your user affiliation directory. In that file you can edit the following variables:
- nodeType (Eligible AWS node instances are listed within the
cloudflow/cluster/AWSCluster.pyPython script under variableawsTypeswith the associated CPU core count) - nodeCount (Number of nodes you want to utilize of the given AWS node instance. A word of caution as the Cloud-Sandbox AWS account does have caps on the number of nodes you can allocate for a given instance)
- tags (The Values for “Name” and “Project” should reflect your model name and affiliation).
- Build the workflow for your model -
workflow_main.pyChange directory into
Cloud-Sandbox/cloudflow/workflowsEdit the workflow_main.py script.
- In the function
main()under the for loopfor jobfile in joblist:- add an
elif jobtype ==block with your JOBTYPE that you specified in the job file from Step #3.- You can copy the block for
ucla-romsfor most models and replace the jobtype. Otherwise, choose or create a function inCloud-Sandbox/cloudflow/workflows/flows.py
- You can copy the block for
- If you need to execute a specific Python environment pathway, then you will also need to replace the
-S python3 -usyntax in the very first line with the Python environment pathway. For example, the first line will end up looking like#/usr/bin/env /pathway/to/python.
- add an
- Build the workflow for your model -
your_model_name.pyWe will create a file to read in the keywords from your job file by creating a new script and class for your model. Change directory into
Cloud-Sandbox/cloudflow/jobWithin that directory, you will want to copy the file called Model_Experiment_Template.py to your_model_name_Experiment.py. Within the your_model_name_Experiment.py file, you will see that the original configuration of this file is reflected strictly based on the cloud_sandbox_experiment.template job file you’ve copied and modified in Step #3.
- Rename the “Model_Experiment_Template” Python class name to “your_model_name_Experiment” so this can now reflect a unique Python class for your own specific model with the basic approach for model execution.
- If your model execution only needs to know essentially the location of the model run directory and then executable itself, then you don’t need to modify anything else in this file.
- If your model executable needs more information (e.g, model argument, specific model libraries to be linked) that you’ve included
your_model_name.expfile in Step #3, then you will need to include that information within theparseConfigfunction inside your new Python class.- e.g., if you added the variable
"IN_FILE" : "path/to/inputfile"in youryour_model_name.expjob file, then you will need to addself.IN_FILE = cfDict['IN_FILE']in under theparseConfigfunction. - If you want to add more optional arguments for your model experiement workflow besides its standard model execution that's defined in your
your_model_name.expjob file, then you will also need to addself.OPTIONAL_ARGUMENT_NAME = cfDict.get('OPTIONAL_ARGUMENT_NAME', "DEFAULT_VALUE")in under theparseConfigfunction.
- e.g., if you added the variable
- Build the workflow for your model -
JobFactory.pyEditJobFactory.pyfile and at the very top of the script, you will now add a new import statement to reflect the new “your_model_Experiment” Python class you’ve constructed from theyour_model_name_Experiment.pyfile you created in Step #6
- e.g.,
from cloudflow.job.your_model_name_Experiment import your_model_name_ExperimentInside class JobFactory, edit the job function
- Insert an
if/elifstatement that reflects your respectiveMODELclass andjobtypeclass variable defined in youryour_model_name.expfile constructed in Step #3 - Call the new Python class you created in Step #6 and imported here
newjob = your_model_name(configfile, NPROCS)
- Build the workflow for your model -
tasks.pyIf your model executable does not need more information than the template job file (just the pathway to the model run directory and the executable inyour_model_name.basicjob file in Step #3) then skip this step and move to the next one. If you added variables to your job file, read this step.
Change directory into
Cloud-Sandbox/cloudflow/workflowsEdit tasks.py
- Edit the
template_runfunction- Inside that function, add an
elifstatement within the code block to include the extra arguments for your model from your job fileyour_model_name.expcreated in Step #3 to include in the launcher script that you will modify in Step #9. - Add the new variable(s) you created in your job file to the
resultcommand.- e.g., if you added the variable
"IN_FILE" : "path/to/inputfile"in youryour_model_name.expjob file,- define the variable
IN_FILE = job.IN_FILEand addstr(IN_FILE)in thesubprocess.runcommand within the square brackets.
- define the variable
- Make sure you put the new variables at the end of the other strings in the square brackets.
- e.g., if you added the variable
- Copy and modify the code logic like in the
schism,dflowfm, orucla-romscode blocks for each model class, then add the jobtype of your specific model class.
- Inside that function, add an
- Build the workflow for your model -
basic_launcher.pyThis controls the launch of your model inside the Cloud-Sandbox. Change directory into
Cloud-Sandbox/cloudflow/workflowsModify basic_launcher.sh
- If you completed Step #8 due to model information required from the job file to kick start the executable,
- make an
ifshell script block to ingest the special input argument(s)- You can simply follow along with the code blocks for
schism,dflowfm, orucla-romsfor theexportstatements. - The
$7in theseexportstatements refer to the 7th string added in theresult = subprocess.runcommand intasks.py, which should be the new variable you added in Step #8. If you added more than one new variable, they would be$8,$9, and so on.
- You can simply follow along with the code blocks for
- make an
Next, go towards the bottom of the script where you see a set of if/elif blocks of code for specific model suites.
- Construct a code block for “your_model” that points to a shell launcher script and add the specific arguments required to run your model
- That launcher script for “your_model” will be constructed in the next step, where it takes specific arguments required to run your model.
- The default requirements for each of the launcher scripts are the model run directory and the pathway to the model executable. If more is required for your given model to launch, then include those as well similar to the code logic you see in the other if/elif code blocks.
-
Build the workflow for your model -
your_model_run.shCopy themodel_basic_run.shfile to a new file calledyour_model_run.sh. Inside that file, you will see the general workflow template that you will need to modify: (1) Load the required compilers and libraries used to compile your model (2) Export any required environmental variables needed for the AWS cloud environment to run your given model executable and (3) Call the method to run your model with the given executable (e.g., mpirun or mpiexec). -
Now, go back to
experiment_launcher.shand ensure that the if/elif code block you’ve constructed in Step #9 is matching the syntax of calling that specific shell script. Make sure to be explicit on your model code base workflow based on yourJOBTYPEandAPPoptions you've inserted within theyour_model_name.expin Step #3. Check that the code logic also includes the script arguments required to properly run the given model suite.
- You may need to add an
ifstatement forMPIOPTSfor your model depending on how your model is configured.
- Run your model Finally, we can now attempt to run the model! Change directory into
Cloud-Sandbox/cloudflowMake sure you are in the cloudflow directory before running any models. Follow the steps below to submit the Cloud-Sandbox job submission to the background of the head node and monitor your job progress:
./workflows/workflow_main.py ./cluster/configs/your_model_name.config ./job/jobs/your_model_name.exp &> your_model_test.out &To see the progress of the Cloud-Sandbox execution of your model
tail -f your_model_test.out