# "lost node" failures when running hl.experimental.run\_combiner()

**URL:** <https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529>\
**Category:** Hail Query & hailctl\
**Created:** [July 13, 2020, 11:48pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529 "2020-07-13T23:48:39Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![jinasong](https://avatars.discourse-cdn.com/v4/letter/j/9fc348/32.png) [@jinasong](https://discuss.hail.is/u/jinasong)\
**Post date:** [July 13, 2020, 11:48pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/1 "2020-07-13T23:48:39Z")

</div>

Hello,

I ran ‘hl.experimental.run\_combiner()’ with ‘branch factor value=100’ and ‘batch factor = 100’ for 1000 WGS gvcfs in GCP. In my successful run, there are some failed tasks causing multiple attempts and long runtime. I used the n1-highmem-32 machine type, spark.executor.mem = 72g and spark.executor.core = 8. Based on these spark conf parameters, each node launched up to 2 executors. Please give me some advice on how to remove these failure tasks. Thank you.

1. Round 1 (1000 gvcfs --\> 10 MT) : 234 failed / 250 tasks  
Error 1 : ExecutorLostFailure (executor 30 exited unrelated to the running tasks) Reason: Container marked as failed: container\_1594240183345\_0002\_01\_000045 on host: \*\*\*\*. Exit status: -100. Diagnostics: Container released on a _lost_ node.

2. Round 2 (10 MT --\> 1 MT) : 4567 failed / 95623 tasks  
Error 2 : ExecutorLostFailure (executor 78 exited caused by one of the running tasks) Reason: Container from a bad node: container\_1594240183345\_0002\_01\_000124 on host: \*\*\*\*. Exit status: 134. Diagnostics: [2020-07-11 08:44:42.110]Exception from container-launch.

Best,  
Jina

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [July 14, 2020, 12:19am UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/2 "2020-07-14T00:19:41Z")

</div>

This is quite weird. Can you paste the full stack trace? You really shouldn’t need to be using these high memory settings, there’s almost certainly something wrong in the Hail runtime that we can fix.

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [July 14, 2020, 12:27pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/3 "2020-07-14T12:27:00Z")

</div>

Also, we’ve fixed a few memory leaks in the last couple versions, so make sure you’re on the latest (0.2.49 at time of posting)

---

<div class="post-metadata">

**Author:** ![jinasong](https://avatars.discourse-cdn.com/v4/letter/j/9fc348/32.png) [@jinasong](https://discuss.hail.is/u/jinasong)\
**Post date:** [July 14, 2020, 9:19pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/4 "2020-07-14T21:19:55Z")

</div>

Hello Tim,

1. Unfortunately, the log file for 1000 gvcfs was overwritten by 200 gvcfs case ( my another run). Instead, I attached the lost node error parts in the 200 gvcfs case log file.

[_lost_ node error parts In the 200 gvcfs combiner.txt|attachment](https://discuss.hail.is/uploads/short-url/mEhopAsAGhUkxDASLPVQqWXzGJ3.txt) (6.2 KB)

Please review it. And let me know if you need more information for solving this issue. The total log file size is about 200M. By the way, I am curious if these errors are related to memory size.

1. For your information, with executor.core = 8, executor.mem = 38g, executor.memOverhead = 15g, when I ran run\_combiner for 1000 gvcfs, I got a lot of error of this type :

ExecutorLostFailure (executor 27 exited caused by one of the running tasks) Reason: Container killed by YARN for exceeding memory limits. 53.0 GB of 53 GB physical memory used. Consider boosting spark.yarn.executor.memoryOverhead or disabling yarn.nodemanager.vmem-check-enabled because of YARN-4714.

But after increasing the memory size of executor, this type error was removed.

1. In the addition, I found out that the output file sizes of run\_combiner() with the same input and different spark configuration is different. How can I find out if the output MTs are identical or not?

Thank you for your supports.

Best,  
Jina

---

<div class="post-metadata">

**Author:** ![kumarveerapen](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/kumarveerapen/32/355_2.png) [@kumarveerapen](https://discuss.hail.is/u/kumarveerapen)\
**Post date:** [July 30, 2020, 3:17pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/5 "2020-07-30T15:17:55Z")

</div>

Dear Jina

Apologies for our late reply.  
1 and 3) I’ll tag @tpoterba

We’re glad that your memory size issue was solved with the executor memory size tweaking.

---

<div class="post-metadata">

**Author:** ![jinasong](https://avatars.discourse-cdn.com/v4/letter/j/9fc348/32.png) [@jinasong](https://discuss.hail.is/u/jinasong)\
**Post date:** [July 30, 2020, 5:52pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/6 "2020-07-30T17:52:47Z")

</div>

Hello Kumar,

Thank you for the reply. Unfortunately, I could not resolve “lost node” error so far. If I can get any insight from the Hail team, I will really appreciate it.

Best,  
Jina

---

<div class="post-metadata">

**Author:** ![danking](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/danking/32/43_2.png) [@danking](https://discuss.hail.is/u/danking)\
**Post date:** [July 30, 2020, 6:22pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/7 "2020-07-30T18:22:44Z")

</div>

@jinasong,

Sorry for the recent instability in the Hail library. We’re investigating more thorough scale testing practices that will discover these problems before release. We believe your issue might be fixed in Hail 0.2.52.

---

<div class="post-metadata">

**Author:** ![jinasong](https://avatars.discourse-cdn.com/v4/letter/j/9fc348/32.png) [@jinasong](https://discuss.hail.is/u/jinasong)\
**Post date:** [July 30, 2020, 8:41pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/8 "2020-07-30T20:41:58Z")

</div>

Hi @danking,

I will retry it with the latest version of Hail and keep posting. Thank you for your support.

Best,  
Jina

---

<div class="post-metadata">

**Author:** ![jinasong](https://avatars.discourse-cdn.com/v4/letter/j/9fc348/32.png) [@jinasong](https://discuss.hail.is/u/jinasong)\
**Post date:** [August 17, 2020, 7:19pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/9 "2020-08-17T19:19:40Z")

</div>

Hi @danking,

I ran it again with Hail 0.2.54. In this new version, I found out the run\_combiner() function requires the information of ‘use\_genome\_default\_intervals’ and gave it ‘True’. After that, I can find a dramatically increased number of tasks. But, I encountered another type of error as below.

Error message : error reading tabix-indexed file gs://my-project/my-bucket/my-sample.g.vcf.gz: i=0, curOff=386139765735504, expected=386139765735424  
at is.hail.io.tabix.TabixLineIterator.next(TabixReader.scala:417)…

I have never seen this type of error in other previous versions.

Could you give me an idea to solve this issue? Thank you.

Best,  
Jina

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [August 17, 2020, 8:39pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/10 "2020-08-17T20:39:17Z")

</div>

We’ve replicated this issue with public data, and are working on a fix.

---

<div class="post-metadata">

**Author:** ![jinasong](https://avatars.discourse-cdn.com/v4/letter/j/9fc348/32.png) [@jinasong](https://discuss.hail.is/u/jinasong)\
**Post date:** [August 17, 2020, 11:22pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/11 "2020-08-17T23:22:41Z")

</div>

Sounds great. Looking forward to the good news.

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [August 18, 2020, 3:03pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/12 "2020-08-18T15:03:51Z")

</div>

OK, we’ve characterized the problem (it’s a bug in control flow when a VCF line ends on the last byte of a compressed BGZ block). It’s not a trivial one-line fix, so stay tuned.

---

<div class="post-metadata">

**Author:** ![jinasong](https://avatars.discourse-cdn.com/v4/letter/j/9fc348/32.png) [@jinasong](https://discuss.hail.is/u/jinasong)\
**Post date:** [August 26, 2020, 9:35pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/13 "2020-08-26T21:35:38Z")

</div>

Hi Tim,

I just saw that the new Hail version 0.2.55 released. I wonder if the issue in the run\_combiner() function was resolved. Thank you.

Best,  
Jina

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [August 27, 2020, 1:12pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/14 "2020-08-27T13:12:48Z")

</div>

This is fixed but the fix went in after 0.2.55. We can make a new release today.

---

<div class="post-metadata">

**Author:** ![jinasong](https://avatars.discourse-cdn.com/v4/letter/j/9fc348/32.png) [@jinasong](https://discuss.hail.is/u/jinasong)\
**Post date:** [September 4, 2020, 7:02am UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/15 "2020-09-04T07:02:29Z")

</div>

Hi Tim,

I updated Hail, as of version 0.2.56 and tested run\_combiner() function with 100, 1k, and 10k gvcf files (average size : 6G) each. Run\_combiner() runs for 100 gvcfs and 1k gvcfs completed successfully through multiple attempts of failed subtasks showing similar messages as before. Run time in a new version was faster than in the previous Hail version. Thanks much for your and your team’s work.

Q1. By the way, the sizes of output MTs are different from the outputs from previous Hail version 0.2.52.  
: A sparse MT size for 100 gvcfs - 317G (in v0.2.56), 600G (in v0.2.52)  
: A sparse MT size for 1k gvcfs - 2.8T (in v0.2.56), 6T (in v0.2.52)  
Please let me know how I should interpret this.

Q2. In addition, unfortunately, the job for 10k gvcfs has been failed. The first round of 100 batches, merging 100 gvcfs to a sparse MT, completed successfully, but it failed when starting the second round with the error message as below. If you let me know how to resolve this issue, I will really appreciate it.

– Caused by: java.io.IOException: All datanodes [DatanodeInfoWithStorage[\*\*\*\*\*\*\*,DISK]] are bad

I found out 100 sparse MTs generated by the first round in a run\_combiner() run in my temp storage.  
Is it possible to combine 100 MTs to 1 MT with any other Hail function?

Thank you.  
-Jina

---

<div class="post-metadata">

**Author:** ![jinasong](https://avatars.discourse-cdn.com/v4/letter/j/9fc348/32.png) [@jinasong](https://discuss.hail.is/u/jinasong)\
**Post date:** [September 8, 2020, 11:53pm UTC](https://discuss.hail.is/t/lost-node-failures-when-running-hl-experimental-run-combiner/1529/16 "2020-09-08T23:53:45Z")

</div>

Please ignore my first question. The output size is the same as before. Sorry about that. I am looking forward to your advice for combining 10k gvcfs successfully I mentioned in my second question. Thank you

-Jina
