# Individual GT call output handling issue

**URL:** <https://discuss.hail.is/t/individual-gt-call-output-handling-issue/3975>\
**Category:** Hail Query & hailctl\
**Created:** [October 28, 2024, 6:02pm UTC](https://discuss.hail.is/t/individual-gt-call-output-handling-issue/3975 "2024-10-28T18:02:30Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![cchunju8286](https://avatars.discourse-cdn.com/v4/letter/c/3bc359/32.png) [@cchunju8286](https://discuss.hail.is/u/cchunju8286)\
**Post date:** [October 28, 2024, 6:02pm UTC](https://discuss.hail.is/t/individual-gt-call-output-handling-issue/3975/1 "2024-10-28T18:02:30Z")

</div>

Hello,

I’m working on outputting genotype (`GT`) calls for each sample, but I’m hitting some issues with handling the matrix table. To avoid conflicts, I’ve ended up using `.collect()` to bring the entire table into memory, which isn’t ideal due to high memory consumption.

##### My script now:

sample = col.s  
(Filter the MatrixTable for the specific sample)  
sample\_mt = self.mt.filter\_cols(self.mt.s == sample)  
(Extract the GT column,Table handle this way to avoid structure conflict)  
sample\_entries\_table = sample\_mt.entries()  
sample\_entries\_table = sample\_entries\_table.select(‘GT’)  
(Convert the entries (GT field) to a simple array without keys)  
gt\_array = sample\_mt.entries().select(‘GT’).collect()

##### As I previously used with export():

(This outputs extra columns (locus, alleles, and s))  
sample\_entries\_table = sample\_mt.entries().select(‘GT’)  
sample\_entries\_table.export(output\_file)

I am not aware if this is the key column issue or function handling problem, but I feel that there is a more efficient way to extract the individual GT calls without using collect(). Any guide would be appreciated.

---

<div class="post-metadata">

**Author:** ![ehigham](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/ehigham/32/943_2.png) [@ehigham](https://discuss.hail.is/u/ehigham)\
**Post date:** [October 28, 2024, 6:16pm UTC](https://discuss.hail.is/t/individual-gt-call-output-handling-issue/3975/2 "2024-10-28T18:16:40Z")

</div>

Hi @cchunju8286,

Would you mind sharing what you want to do with the result? Maybe we can give you a better answer with more details.

You can omit the extra fields and globals in `Table.export` by dropping the key:

```python
>>> mt = sample_mt.entries().key_by().select_globals().select('GT')
>>> mt.describe()
----------------------------------------
Global fields:
    None
----------------------------------------
Row fields:
    'GT': call 
----------------------------------------
Key: []
----------------------------------------

```

Hope this helps,

---

<div class="post-metadata">

**Author:** ![cchunju8286](https://avatars.discourse-cdn.com/v4/letter/c/3bc359/32.png) [@cchunju8286](https://discuss.hail.is/u/cchunju8286)\
**Post date:** [November 18, 2024, 8:27am UTC](https://discuss.hail.is/t/individual-gt-call-output-handling-issue/3975/3 "2024-11-18T08:27:49Z")

</div>

Hi @ehigham ,  
sorry for the late response, I was focusing on other projects. The method you told me did work, and it solved the problem. The output I wanted is simply for each individual (like below):  
GT  
0/1  
0/0  
0/0

If you don’t mind I wanted to ask another question, cause I tried to parallelize the task, and it returns “joblib.externals.loky.process\_executor.BrokenProcessPool: A task has failed to un-serialize. Please ensure that the arguments of the function are all picklable.”  
I think what I understand is that this particular Hail matrix can not pass on to multiple tasks at the same time? I am also trying to avoid reading the matrix multiple times, so I would not overload the memory during the processing. Have you ever encounter this problem, or have an idea of this issue?

My code in the main script:

def process\_sample(sample):  
# Filter MatrixTable for a specific sample  
sample\_mt = mt.filter\_cols(mt.s == sample)  
output\_sample\_gt(sample\_mt, sample, chr, output\_dir)

```
# Parallelize the processing
Parallel(n_jobs=n_jobs)(
    delayed(process_sample)(sample) for sample in sample_ids
)

```

Let me know if the problem is clear.

Thanks again
