# A problem with KeyTable.from\_pandas in hail v0.1

**URL:** <https://discuss.hail.is/t/a-problem-with-keytable-from-pandas-in-hail-v0-1/461>\
**Category:** Help \[0.1\]\
**Created:** [April 12, 2018, 2:51pm UTC](https://discuss.hail.is/t/a-problem-with-keytable-from-pandas-in-hail-v0-1/461 "2018-04-12T14:51:04Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Wei\_Zhang](https://avatars.discourse-cdn.com/v4/letter/w/3d9bf3/32.png) [@Wei\_Zhang](https://discuss.hail.is/u/Wei_Zhang)\
**Post date:** [April 12, 2018, 2:51pm UTC](https://discuss.hail.is/t/a-problem-with-keytable-from-pandas-in-hail-v0-1/461/1 "2018-04-12T14:51:04Z")

</div>

Hi everyone,

I’m new to hail. I have a problem annotating vds using the KeyTable generated from a pandas dataframe. The code looks like this:

> # load gene expression data
> 
> residuals = pd.read\_table(‘eQTL/data/residuals\_5.txt’, index\_col=0)  
> residuals = residuals.T
> 
> # Create a permutation matrix
> 
> permutations = pd.concat([residuals[gene]]\*n, axis=1)  
> permutations = permutations.apply(np.random.permutation)  
> permutations = pd.concat([residuals[gene], permutations], axis=1)  
> permutations.columns = [‘y’] + [‘p{}’.format(i+1) for i in range(n)]  
> permutations.reset\_index(inplace=True)
> 
> # Create a KeyTable from pandas dataframe
> 
> kt = KeyTable.from\_pandas(permutations).key\_by(‘index’)
> 
> # Annotate vds by the KeyTable
> 
> vds2\_cis = vds2\_cis.annotate\_samples\_table(kt, root=‘sa.pheno’)

Then I got the following error message:

Name: org.apache.toree.interpreter.broker.BrokerException  
Message: Traceback (most recent call last):  
File “/tmp/kernel-PySpark-34094fbd-6f73-467d-b19e-a06c898f987f/pyspark\_runner.py”, line 194, in   
eval(compiled\_code)  
File “”, line 1, in   
File “”, line 2, in annotate\_samples\_table  
File “/mnt/tmp/spark-aba52f07-bfd3-4ce2-9096-34956bfd6316/userFiles-83cc5db7-9eb6-47a3-a292-96e96b5c6b3c/hail-python.zip/hail/java.py”, line 121, in handle\_py4j  
‘Error summary: %s’ % (deepest, full, Env.hc().version, deepest))  
FatalError: SparkException:  
Error from python worker:  
/usr/bin/python: No module named pyspark  
PYTHONPATH was:  
/mnt/yarn/usercache/hadoop/filecache/260/\_\_spark\_libs\_\_8515568189444678508.zip/spark-core\_2.11-2.1.0.jar  
java.io.EOFException

Alternatively, I tried to save the pandas dataframe to a file, and then load it using hc.import\_table. This time it works fine:

> # Save the pandas dataframe to a file:
> 
> permutations.to\_csv(‘permutations.txt’, sep=‘\t’, index=False)
> 
> # Load this file as a KeyTable
> 
> kt2 = hc.import\_table(‘permutations.txt’, impute=True, missing=‘’).key\_by(‘index’)
> 
> # Annotate vds by the KeyTable
> 
> vds2\_cis = vds2\_cis.annotate\_samples\_table(kt2, root=‘sa.pheno’)

This doesn’t make sense since kt and kt2 look almost identical. Loading it every time from a file leads to a huge IO overhead since I’m looping through 20k genes.

I’d highly appreciate it if anyone can help me troubleshooting the above error message I got by running “KeyTable.from\_pandas” followed by “VariantDataset.annotate\_samples\_table”

Thank you very much!  
Wei

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [April 12, 2018, 2:53pm UTC](https://discuss.hail.is/t/a-problem-with-keytable-from-pandas-in-hail-v0-1/461/2 "2018-04-12T14:53:34Z")

</div>

This is a Spark setup problem. Your Spark worker nodes don’t have PySpark properly installed, and KeyTable.from\_pandas just calls pandas to Spark conversion inside Spark.

---

<div class="post-metadata">

**Author:** ![Wei\_Zhang](https://avatars.discourse-cdn.com/v4/letter/w/3d9bf3/32.png) [@Wei\_Zhang](https://discuss.hail.is/u/Wei_Zhang)\
**Post date:** [April 12, 2018, 3:05pm UTC](https://discuss.hail.is/t/a-problem-with-keytable-from-pandas-in-hail-v0-1/461/3 "2018-04-12T15:05:50Z")

</div>

Hi @tpoterba, Thanks for the quick reply! I’m curious why “KeyTable.from\_pandas” itself doesn’t trigger the problem but “vds.annotate\_samples\_table” does? If the problem is in “vds.annotate\_samples\_table”, why the KeyTable imported from a file works?

Anyway, can you point me to any online tutorial regarding how to install PySpark properly? Thank you!

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [April 12, 2018, 3:35pm UTC](https://discuss.hail.is/t/a-problem-with-keytable-from-pandas-in-hail-v0-1/461/4 "2018-04-12T15:35:21Z")

</div>

Spark (and Hail) compute lazily – Spark jobs won’t be run until they need to produce a result. This means that errors don’t always happen at the line of Python you expect them too, because multiple steps are fused together. The problem here was definitely the from\_pandas – any Spark “action” on this table would have crashed.

---

<div class="post-metadata">

**Author:** ![Wei\_Zhang](https://avatars.discourse-cdn.com/v4/letter/w/3d9bf3/32.png) [@Wei\_Zhang](https://discuss.hail.is/u/Wei_Zhang)\
**Post date:** [April 12, 2018, 4:20pm UTC](https://discuss.hail.is/t/a-problem-with-keytable-from-pandas-in-hail-v0-1/461/5 "2018-04-12T16:20:41Z")

</div>

@tpoterba That makes a lot of sense. Thanks for your explanation!
