# Spark memory error trying to write matrixtable

**URL:** <https://discuss.hail.is/t/spark-memory-error-trying-to-write-matrixtable/3007>\
**Category:** Hail Query & hailctl\
**Created:** [December 22, 2022, 9:56am UTC](https://discuss.hail.is/t/spark-memory-error-trying-to-write-matrixtable/3007 "2022-12-22T09:56:13Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![jbchang](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/jbchang/32/818_2.png) [@jbchang](https://discuss.hail.is/u/jbchang)\
**Post date:** [January 11, 2023, 6:01am UTC](https://discuss.hail.is/t/spark-memory-error-trying-to-write-matrixtable/3007/4 "2023-01-11T06:01:21Z")

</div>

Hi Tim et al—so I think I have tried to do what you’re suggesting, which is to annotate the columns of a matrixtable with a chunk of the phenotypes at a time, and then run a set of regressions, and then re-load the matrixtable and annotate with another chunk of phenotypes, etc. The main issue now is that it’s super slow. Any suggestions?

One thing that had come up previously is that there’s no way to quickly run logistic regression on phenotypes that all have different missingness patterns (see [this post](https://discuss.hail.is/t/phewas-on-dnanexus-ukb-rap/2967/13)). So, that’s why I’m currently just running them one-by-one.

Here’s the (somewhat simplified) code:

This chunk (thanks to Dan from [this post](https://discuss.hail.is/t/problem-loading-multiple-csv-files-for-annotation-on-dnanexus/3002/3)) runs every ~40 phenotypes (I use `hl.import_table()` to import ~40 phenotypes at a time)

```python
phecodes = ht.aggregate(hl.agg.collect_as_set(ht.phecode))
ht = ht.group_by(
    ht.sample_id
).aggregate(
    phenos = hl.dict(hl.agg.collect((ht.phecode, ht.row)))
)
ht = ht.annotate(**{
    phecode: ht.phenos.get(phecode)
    for phecode in phecodes
})
mt = mt.annotate_cols(**ht[mt.col_key])

```

Then, for each of the 40 `phenotype`s, I run the following code:

```python
regression_results = hl.logistic_regression_rows( 
    test = 'firth',
    y=mt_burden[f'{phenotype}'].has_phenotype,
    x=mt_burden.n_variants,
    covariates=[1.0,
                mt_burden.sex_float, 
                mt_burden.p21003_i0_squared,
                mt_burden.p21003_i0_float,
                mt_burden.p22009_a1,
                mt_burden.p22009_a2,
                mt_burden.p22009_a3,
                mt_burden.p22009_a4,
                mt_burden.p22009_a5,
                mt_burden.p22009_a6,
                mt_burden.p22009_a7,
                mt_burden.p22009_a8,
                mt_burden.p22009_a9,
                mt_burden.p22009_a10])

```

Best,  
Jeremy

---

_[View the full topic](https://discuss.hail.is/t/spark-memory-error-trying-to-write-matrixtable/3007)._
