# Filter variants based on other files

**URL:** <https://discuss.hail.is/t/filter-variants-based-on-other-files/2470>\
**Category:** Hail Query & hailctl\
**Created:** [February 9, 2022, 3:18pm UTC](https://discuss.hail.is/t/filter-variants-based-on-other-files/2470 "2022-02-09T15:18:30Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![yc000000](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/yc000000/32/640_2.png) [@yc000000](https://discuss.hail.is/u/yc000000)\
**Post date:** [February 9, 2022, 3:18pm UTC](https://discuss.hail.is/t/filter-variants-based-on-other-files/2470/1 "2022-02-09T15:18:30Z")

</div>

I have a Hail matrix table with variants and samples (h1) and a txt file from clinvar vcf. I would like to filter out the variants (row) that are not in the clinvar txt file but I am not sure how.

Hail matrix table is keyed on chr:pos and the array of alleles. I successfully imported txt file as hail table (no key) and tried to annotate h1 with this table then filter with the new field.

However, I don’t know how to key the txt file to be the same as h1 in order to annotate it then filter. Or, am I thinking it completely wrong?

Sorry for the basic question, I started Hail only recently, would greatly appreciate if someone has some idea for the situation.

Thank you!

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [February 9, 2022, 3:28pm UTC](https://discuss.hail.is/t/filter-variants-based-on-other-files/2470/2 "2022-02-09T15:28:28Z")

</div>

How are the variants formatted in the clinvar table?

It will look something like:

```auto
# annotate the clinvar table to create a locus and alleles
clinvar = clinvar.key_by('locus', 'alleles')

# semi_join_rows keeps the rows whose keys overlap with the table's keys
mt = mt.semi_join_rows(clinvar) 

```

---

<div class="post-metadata">

**Author:** ![yc000000](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/yc000000/32/640_2.png) [@yc000000](https://discuss.hail.is/u/yc000000)\
**Post date:** [February 9, 2022, 3:53pm UTC](https://discuss.hail.is/t/filter-variants-based-on-other-files/2470/3 "2022-02-09T15:53:26Z")

</div>

![image](https://canada1.discourse-cdn.com/flex036/uploads/hail/original/1X/d2fd4dcdad96f042d67335a3511f6bcfb46aaeb7.png)

these four are the columns that should be keyed… and h1 key is like  
chr1:17018956 [“A”,“T”]

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [February 9, 2022, 6:06pm UTC](https://discuss.hail.is/t/filter-variants-based-on-other-files/2470/4 "2022-02-09T18:06:40Z")

</div>

The missing bit can be:

```auto

clinvar = clinvar.annotate(
    locus = hl.locus(clinvar.Chromosome, clinvar.PositionVCF, reference_genome='GRCh37'),
    alleles = [clinvar.Reference, clinvar.AlternateAlleleVCF]

```

I should say, if you had the VCF just importing that would make this easier!
