You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 64f778c
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: pages/flavours.md
+11-22Lines changed: 11 additions & 22 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,18 +1,14 @@
1
1
---
2
-
title: Choosing the right environment size
2
+
title: Choose the right environment
3
3
type: Using BioShell
4
4
description: How to choose the right number of CPUs, memory, and storage for your BioShell environment.
5
5
---
6
6
7
-
{% include callout.html type="note" content="A flavour is the combination of virtual CPUs and memory allocated to your BioShell environment, essentially the “spec” of your cloud computer. Different flavours suit different workloads, just as you might choose a lightweight laptop for email but a workstation for video editing. You pick a flavour when you request access and can request a change later if your needs grow." %}
7
+
{% include callout.html type="note" content="A flavour is the combination of virtual CPUs and memory allocated to your BioShell environment, essentially the “spec” of your virtual machine. Different flavours suit different workloads." %}
8
8
9
-
Cloud systems are shared research resources. As a general principle, you are encouraged to
10
-
request resources that closely match your actual needs. This supports fair access for all
11
-
users and preserves capacity for everyone.
9
+
As a general principle, we encourage you to request resources that closely match your actual needs.
12
10
13
-
Estimating requirements can be challenging, particularly at the start of a project you may
14
-
not yet know which software tools you will use or how demanding they will be. The guidance
15
-
below is designed to help you make a reasonable first choice and adjust from there.
11
+
Estimating requirements can be challenging, particularly at the start of a project you may not yet know which software tools you will use or how demanding they will be. The guidance below is designed to help you make a reasonable first choice and adjust from there.
16
12
17
13
18
14
## A familiar starting point {#familiar-starting-point}
@@ -27,8 +23,7 @@ background applications.
27
23
28
24
If you are new to BioShell, or unsure of your requirements, starting with a
29
25
**laptop-equivalent size** (4 CPUs / 8–16 GB RAM) is a reasonable default. You can always
30
-
request a larger environment if you find you need it.
31
-
26
+
request a larger VM flavour when you need it.
32
27
33
28
## Suggested sizes by workload {#workload-sizes}
34
29
@@ -49,14 +44,14 @@ datasets.
49
44
|**Memory**| 4–8 GB |
50
45
|**Storage**| Up to 100 GB |
51
46
52
-
**Example:**John is starting a research project analysing drought-resistant genes from 20
47
+
**Example:**Fred is starting a research project analysing drought-resistant genes from 20
53
48
crop samples (~140 GB raw data). His pipeline runs quality control (`FASTQC`), adapter
54
49
trimming (`cutadapt`), alignment and annotation (`blast`, `SPAdes`), and phylogenetic tree
55
50
construction (`MrBayes`). Of these, `blast` and `SPAdes` are the most CPU- and
56
-
memory-intensive tools in the pipeline, but because John is selecting out a small set of
51
+
memory-intensive tools in the pipeline, but because Fred is selecting out a small set of
57
52
drought-resistant genes rather than whole genomes, each run only needs 2–4 CPUs and
58
53
under 10 GB of RAM. A balanced environment at this size handles the pipeline
59
-
comfortably. If John later extends the analysis to many more genes, those steps become more
54
+
comfortably. If Fred later extends the analysis to many more genes, those steps become more
60
55
CPU-bound and he should move to the medium size below with more cores.
61
56
62
57
@@ -73,7 +68,7 @@ particularly those involving large in-memory data objects.
73
68
|**Storage**| Variable, depends on sample count |
74
69
75
70
**Example:** Michael is running the
76
-
[**SIH scRNAvigator notebooks**](https://github.com/Sydney-Informatics-Hub/scrna-analysis) in
71
+
[**scRNAvigator notebooks**](https://github.com/Sydney-Informatics-Hub/scrna-analysis) in
annotation, differential gene expression, and pathway enrichment analysis. Integration and
79
74
doublet detection steps load large data objects into memory simultaneously, making this
@@ -189,16 +184,10 @@ installed by default on almost every Linux system, useful if `htop` isn't availa
189
184
## Beyond a single environment: when to consider HPC {#beyond-single-environment}
190
185
191
186
BioShell environments are well suited to interactive work and moderate-scale pipelines, but
192
-
they have a ceiling. If your workload keeps growing, more samples in parallel, whole genomes
193
-
rather than subsets, cohorts scaling into the hundreds, a single environment may no longer
187
+
they have a ceiling. If your workload keeps growing, a single environment may no longer
194
188
be the most efficient option.
195
189
196
-
197
-
{% include callout.html type="tip" content="Before requesting a very large environment, test your pipeline end-to-end on a small subset of your data (a handful of samples, or a reduced reference) on a modest environment. This confirms the pipeline runs correctly and gives you a realistic estimate of per-sample time and resource use, information you’ll need whether you stay on BioShell or move to an HPC system." %}
198
-
199
-
200
190
Once your pipeline is validated, high-throughput or many-sample workloads are often better
201
191
suited to a national HPC facility than to a single cloud environment. The
0 commit comments