# Setting number of preemptible workers in \`hailctl dataproc start\`

**URL:** <https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105>\
**Category:** Hail Query & hailctl\
**Created:** [October 2, 2019, 3:37pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105 "2019-10-02T15:37:39Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Danish436](https://avatars.discourse-cdn.com/v4/letter/d/9d8465/32.png) [@Danish436](https://discuss.hail.is/u/Danish436)\
**Post date:** [October 2, 2019, 3:37pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/1 "2019-10-02T15:37:39Z")

</div>

I am writing hailctl dataproc start (cluster name) —p12 and —vep GRCh37

It is not recognozing —p12 to increase the size of the cores

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [October 2, 2019, 3:40pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/2 "2019-10-02T15:40:02Z")

</div>

the `-p` should have a single dash, not a double dash.

The full name is `--num-preemptible-workers`

---

<div class="post-metadata">

**Author:** ![Danish436](https://avatars.discourse-cdn.com/v4/letter/d/9d8465/32.png) [@Danish436](https://discuss.hail.is/u/Danish436)\
**Post date:** [October 2, 2019, 5:30pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/3 "2019-10-02T17:30:53Z")

</div>

Just to import a VCF of 40,000 into a mt, should we specify different memory or different other requirements for cluster or do you suggest just increasing the number of nodes?

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [October 2, 2019, 5:32pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/4 "2019-10-02T17:32:30Z")

</div>

No, just increasing the number of nodes should be fine.

---

<div class="post-metadata">

**Author:** ![Danish436](https://avatars.discourse-cdn.com/v4/letter/d/9d8465/32.png) [@Danish436](https://discuss.hail.is/u/Danish436)\
**Post date:** [October 2, 2019, 8:11pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/5 "2019-10-02T20:11:59Z")

</div>

I have used the following to launch my cluster:

hailctl dataproc start development --vep GRCh37 --num-preemptible-workers 40

I am running the following now on 350 cores

import hail as hl  
import hail.expr.aggregators as agg  
import hail.methods  
import pandas as pd  
from typing import \*  
import random  
hl.init()

hl.import\_vcf(‘gs://wes\_development/complete.pheno.n100.vcf.gz’,force\_bgz=True,force=True).write(‘gs://wes\_development/pipeline\_file.mt’, overwrite=True)  
an\_g = hl.read\_matrix\_table(‘gs://wes\_development/pipeline\_file.mt’)  
an\_g = hl.vep(an\_g, ‘gs://hail-common/vep/vep/vep85-loftee-gcloud.json’)  
an\_g.describe()  
print(‘writing’)  
an\_g.write(‘gs://wes\_development/genotype\_annotations\_new.mt’, overwrite=True)

The write command at the end has already taken an hour; but it is still running in GCP. The progress bar shows me the following:  
[Stage 1:\> (0 + 7) / 7]

I don’t think the progress bar has moved yet. so not sure what’s happening

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [October 2, 2019, 8:14pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/6 "2019-10-02T20:14:52Z")

</div>

try loading with `min_partitions=256` on `import_vcf` – this will divide the data into more than 7 chunks, which will both use more than 7 cores, and make progress more evident.

While the default partitioning is generally fine, with VEP it’s good to ensure a higher amount of parallelism.

---

<div class="post-metadata">

**Author:** ![Danish436](https://avatars.discourse-cdn.com/v4/letter/d/9d8465/32.png) [@Danish436](https://discuss.hail.is/u/Danish436)\
**Post date:** [October 2, 2019, 8:16pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/7 "2019-10-02T20:16:30Z")

</div>

Thanks.

So should I use the following command on import:

hl.import\_vcf(‘gs://wes\_development/complete.pheno.n100.vcf.gz’,force\_bgz=True,force=True, min\_partitions=256).write(‘gs://wes\_development/pipeline\_file.mt’, overwrite=True)

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [October 2, 2019, 8:17pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/8 "2019-10-02T20:17:59Z")

</div>

yes, try that. However, remove the `force` option – this luckily isn’t getting used since you use `force_bgz`, but it’s quite dangerous.

---

<div class="post-metadata">

**Author:** ![alanwilter](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/alanwilter/32/390_2.png) [@alanwilter](https://discuss.hail.is/u/alanwilter)\
**Post date:** [May 7, 2020, 8:57am UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/9 "2020-05-07T08:57:18Z")

</div>

And what about `mt.repartition` once you had done `import_vcf ` rather than building again another cluster, if short of only 7 cores?

BTW, did last reply from @tpoterba work for you @Danish436?

I’m preparing a similar project to run vep on 80Gb vcf.gz file, so I’ve very interested in this topic and any similar you may point to me.

Thanks, Alan

---

<div class="post-metadata">

**Author:** ![danking](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/danking/32/43_2.png) [@danking](https://discuss.hail.is/u/danking)\
**Post date:** [May 7, 2020, 3:23pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/10 "2020-05-07T15:23:01Z")

</div>

Hi @alanwilter,

You should avoid using repartition. Repartition performs a “shuffle” which requires non-preemptible workers and is failure prone. I think you might find my [overview of efficiently using Hail](https://github.com/danking/hail-cloud-docs/blob/master/how-to-cloud-with-hail.md#efficiently-using-hail) helpful. You do not need to stop and start your cluster just to change the number of cores. You can [dynamically change the cluster size](https://github.com/danking/hail-cloud-docs/blob/master/how-to-cloud-with-hail.md#dynamic-cluster-size).

You called your file a “80Gb vcf.gz file”. Is your file really gzipped and not bgzipped? gzipped files cannot be read in parallel, so Hail will be extremely slow and not use your worker nodes. If it your VCF is gzipped, tou should decompress that file and re-compress it as a bgzipped file.

If your file is actually bgzipped but uses the file extension “gz,” then `force_bgz=True` will import the file in parallel.

---

<div class="post-metadata">

**Author:** ![alanwilter](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/alanwilter/32/390_2.png) [@alanwilter](https://discuss.hail.is/u/alanwilter)\
**Post date:** [May 7, 2020, 5:34pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/11 "2020-05-07T17:34:08Z")

</div>

> [@danking](#):
>
> If your file is actually bgzipped but uses the file extension “gz,” then `force_bgz=True` will import the file in parallel.

Many thanks @danking. Indeed, I am doing as above since my colleague who created the file guaranteed me it is bgzipped.

---

<div class="post-metadata">

**Author:** ![alanwilter](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/alanwilter/32/390_2.png) [@alanwilter](https://discuss.hail.is/u/alanwilter)\
**Post date:** [May 7, 2020, 5:35pm UTC](https://discuss.hail.is/t/setting-number-of-preemptible-workers-in-hailctl-dataproc-start/1105/12 "2020-05-07T17:35:06Z")

</div>

And yes, I’ve been reading your docs a lot lately 🙂
