# Split\_multi\_hts AD field

**URL:** <https://discuss.hail.is/t/split-multi-hts-ad-field/4015>\
**Category:** Hail Query & hailctl\
**Created:** [December 18, 2024, 1:34pm UTC](https://discuss.hail.is/t/split-multi-hts-ad-field/4015 "2024-12-18T13:34:33Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![DBScan](https://avatars.discourse-cdn.com/v4/letter/d/bc79bd/32.png) [@DBScan](https://discuss.hail.is/u/DBScan)\
**Post date:** [December 18, 2024, 1:34pm UTC](https://discuss.hail.is/t/split-multi-hts-ad-field/4015/1 "2024-12-18T13:34:33Z")

</div>

Hi, I’m trying to split my MT generated by DRAGEN into a diallelic MT. I’m facing the issue that the AD field contains only a single entry for hom-ref samples, but non-hom-ref samples contain N entries where N is the number of alleles per site. Example:

| Column 1 | Alleles | Sample1 | Sample2 | Sample3 |
| --- | --- | --- | --- | --- |
| chr1:10000 | [“A”, “G”, “C”] | 0/0: **30** | 0/1:0,30,0 | 1/2:0,15,15 |
| | | | | |

So what I would need is the following:

| Column 1 | Alleles | Sample1 | Sample2 | Sample3 |
| --- | --- | --- | --- | --- |
| chr1:10000 | [“A”, “G”, “C”] | 0/0: **30,0,0** | 0/1:0,30,0 | 1/2:0,15,15 |
| | | | | |

I’ve tried to create a new AD field like this where each entry contains exactly N elements in the AD list:

```python
mt_annot = mt.annotate_entries(AD = hl.if_else(mt.GT.is_hom_ref(), (mt.AD.append(0) for entry in range(mt.GT.n_alt_alleles())), mt.AD))

```

But I get the following error:

```python
TypeError: 'Int32Expression' object cannot be interpreted as an integer

```

I’m struggling with the part where I need to add “0” as many times as there are n\_alt\_alleles, how could I achieve that?

---

<div class="post-metadata">

**Author:** ![patrick-schultz](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/patrick-schultz/32/265_2.png) [@patrick-schultz](https://discuss.hail.is/u/patrick-schultz)\
**Post date:** [December 18, 2024, 1:46pm UTC](https://discuss.hail.is/t/split-multi-hts-ad-field/4015/2 "2024-12-18T13:46:43Z")

</div>

As a general rule of thumb, python loops and comprehensions almost never mix the way you want with hail code. The way I would append `n_alt_alleles` zeroes is

```python
mt.AD.extend(hl.range(mt.GT.n_alt_alleles().map(lambda x: 0)))

```

@chrisvittal Are you familiar with this issue with DRAGEN generated datasets?

---

<div class="post-metadata">

**Author:** ![DBScan](https://avatars.discourse-cdn.com/v4/letter/d/bc79bd/32.png) [@DBScan](https://discuss.hail.is/u/DBScan)\
**Post date:** [December 18, 2024, 1:50pm UTC](https://discuss.hail.is/t/split-multi-hts-ad-field/4015/3 "2024-12-18T13:50:19Z")

</div>

Thanks for the super speedy reply!  
I’m not very proficient with Python, so I just tried out a bunch of options and neither of them worked.

I now get the following error:

```python
AttributeError: 'Int32Expression' object has no attribute 'map'

```

Edit:  
I think I actually need `mt.alleles` instead of `mt.GT.n_alt_alleles`.

---

<div class="post-metadata">

**Author:** ![patrick-schultz](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/patrick-schultz/32/265_2.png) [@patrick-schultz](https://discuss.hail.is/u/patrick-schultz)\
**Post date:** [December 18, 2024, 2:01pm UTC](https://discuss.hail.is/t/split-multi-hts-ad-field/4015/4 "2024-12-18T14:01:50Z")

</div>

Whoops, sorry, parentheses typo! That should be

```python
mt.AD.extend(hl.range(mt.GT.n_alt_alleles()).map(lambda x: 0))

```

Edit:  
You’re right, should really be

```python
mt.AD.extend(hl.range(mt.alleles.length() - 1).map(lambda x: 0))

```

---

<div class="post-metadata">

**Author:** ![DBScan](https://avatars.discourse-cdn.com/v4/letter/d/bc79bd/32.png) [@DBScan](https://discuss.hail.is/u/DBScan)\
**Post date:** [December 18, 2024, 2:06pm UTC](https://discuss.hail.is/t/split-multi-hts-ad-field/4015/5 "2024-12-18T14:06:38Z")

</div>

Perfect, this works! Thanks a lot. I can never figure out when to use `hl.len(mt.alleles)` vs `mt.alleles.length()` 😑

---

<div class="post-metadata">

**Author:** ![patrick-schultz](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/patrick-schultz/32/265_2.png) [@patrick-schultz](https://discuss.hail.is/u/patrick-schultz)\
**Post date:** [December 18, 2024, 2:07pm UTC](https://discuss.hail.is/t/split-multi-hts-ad-field/4015/6 "2024-12-18T14:07:52Z")

</div>

Either one, they’re the same!

---

<div class="post-metadata">

**Author:** ![chrisvittal](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/chrisvittal/32/327_2.png) [@chrisvittal](https://discuss.hail.is/u/chrisvittal)\
**Post date:** [December 18, 2024, 3:59pm UTC](https://discuss.hail.is/t/split-multi-hts-ad-field/4015/7 "2024-12-18T15:59:03Z")

</div>

I haven’t seen this before, no.

@DBScan how was this dataset generated?

---

<div class="post-metadata">

**Author:** ![DBScan](https://avatars.discourse-cdn.com/v4/letter/d/bc79bd/32.png) [@DBScan](https://discuss.hail.is/u/DBScan)\
**Post date:** [December 19, 2024, 8:18am UTC](https://discuss.hail.is/t/split-multi-hts-ad-field/4015/8 "2024-12-19T08:18:25Z")

</div>

Hi @chrisvittal , I’ve used DRAGEN iterative gVCF Genotyper 4.2 with the following options:

```python
--gg-discard-ac-zero true 
--gg-remove-nonref true

```
