# How to select appropirate cluster specs for hail

**URL:** <https://discuss.hail.is/t/how-to-select-appropirate-cluster-specs-for-hail/2919>\
**Category:** Hail Query & hailctl\
**Created:** [October 24, 2022, 2:33pm UTC](https://discuss.hail.is/t/how-to-select-appropirate-cluster-specs-for-hail/2919 "2022-10-24T14:33:44Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![igorm](https://avatars.discourse-cdn.com/v4/letter/i/e47774/32.png) [@igorm](https://discuss.hail.is/u/igorm)\
**Post date:** [October 24, 2022, 2:33pm UTC](https://discuss.hail.is/t/how-to-select-appropirate-cluster-specs-for-hail/2919/1 "2022-10-24T14:33:44Z")

</div>

Hi,

Can you please provide some guidance how to think about / design the appropriate cluster for hail?

Example: 10k single sample vcf files (~100M variants) imported to MatrixTable.

In oder to efficiently conduct aggregation queries (speed, avoiding running out of memory etc…) on such dataset what kind of cluster would be a good choice in terms of:

- specs for master node
- specs for slave node
- number of slave nodes

Thanks!

---

<div class="post-metadata">

**Author:** ![danking](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/danking/32/43_2.png) [@danking](https://discuss.hail.is/u/danking)\
**Post date:** [October 24, 2022, 8:02pm UTC](https://discuss.hail.is/t/how-to-select-appropirate-cluster-specs-for-hail/2919/2 "2022-10-24T20:02:59Z")

</div>

We generally recommend using whatever autoscaling is provided by your cloud of choice. Import your data into a matrix table and save it in that format (never `import_vcf` and then immediately do analysis). Use spot or preemptible workers unless your pipeline has a “shuffle” (basically: `key_by` and `key_rows_by`). Use a leader/master node with ~16 cores and ~60 GB of RAM. Worker nodes can generally be whatever the standard instance type is. Some operations take a `block_size` parameter which you can set to smaller values if you run into RAM problems.
