# Inconsistent sample qc results

**URL:** <https://discuss.hail.is/t/inconsistent-sample-qc-results/1378>\
**Category:** Hail Query & hailctl\
**Created:** [April 21, 2020, 3:12pm UTC](https://discuss.hail.is/t/inconsistent-sample-qc-results/1378 "2020-04-21T15:12:08Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![wonu](https://avatars.discourse-cdn.com/v4/letter/w/f17d59/32.png) [@wonu](https://discuss.hail.is/u/wonu)\
**Post date:** [April 21, 2020, 3:12pm UTC](https://discuss.hail.is/t/inconsistent-sample-qc-results/1378/1 "2020-04-21T15:12:09Z")

</div>

Hi,

I’m running the following bit of code on exome sequenced data and finding that i get different results when I plot my sample call rate after. I was wondering if the order of commands makes a difference, or if there is something else I am missing. Thanks!

```auto
scz = hl.read_matrix_table('Analyses/scz_main.mt')                     

scz = hl.sample_qc(scz, name = 'sample_qc') 
         
scz.count_rows()  
                                                     
scz = scz.filter_rows(scz.alleles.length() <= 6)

scz = hl.split_multi_hts(scz, permit_shuffle=True)                     

scz = scz.filter_entries(
    hl.is_defined(scz.GT) &
    (
        (scz.GT.is_hom_ref() & 
            (
                 ((scz.AD[0] / scz.DP) < 0.8) | 
                (scz.GQ < 20) |
                (scz.DP < 20)
        	)
        ) |
        (scz.GT.is_het() & 
        	( 
                (((scz.AD[0] + scz.AD[1]) / scz.DP) < 0.8) | 
                ((scz.AD[1] / scz.DP) < 0.2) | 
                (scz.PL[0] < 20) |
                (scz.DP < 20)
        	)
        ) |
        (scz.GT.is_hom_var() & 
        	(
                ((scz.AD[1] / scz.DP) < 0.8) |
                (scz.PL[0] < 20) |
                (scz.DP < 20)
        	)
        )
    ),
    keep = False
)                                                                   

scz.count_rows() 

```

I’m using hail version 0.2.36-ed011219dd93

---

<div class="post-metadata">

**Author:** ![danking](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/danking/32/43_2.png) [@danking](https://discuss.hail.is/u/danking)\
**Post date:** [April 21, 2020, 5:41pm UTC](https://discuss.hail.is/t/inconsistent-sample-qc-results/1378/2 "2020-04-21T17:41:32Z")

</div>

Could you try explaining the issue in a different way, perhaps with the full, exact code and output for each scenario?

It makes sense that the call rate changes after you filter some entries, as described in the code you shared. Why do you expect sample call rate would not change?

---

<div class="post-metadata">

**Author:** ![wonu](https://avatars.discourse-cdn.com/v4/letter/w/f17d59/32.png) [@wonu](https://discuss.hail.is/u/wonu)\
**Post date:** [April 21, 2020, 6:08pm UTC](https://discuss.hail.is/t/inconsistent-sample-qc-results/1378/3 "2020-04-21T18:08:59Z")

</div>

Hi, that’s the exact code, but I’ve had to run it multiple times and seem to get one of two different outputs when I plot the call rate after filtering (see attached photos)

 ![callrate_geno](https://canada1.discourse-cdn.com/flex036/uploads/hail/original/1X/d0f1e06aa7ecce3a49f5db6f8362839bb592f323.png) ![callrate_geno1](https://canada1.discourse-cdn.com/flex036/uploads/hail/original/1X/70f8fdca0490a01669f11e24e2ba5e6ee057412c.png) .

Code for call rate:  
callrate\_geno = hl.plot.histogram(scz.sample\_qc.call\_rate, range=(0,1), legend=‘Call Rate’)  
export\_png(callrate\_geno, filename=‘Output/callrate\_geno.png’)

---

<div class="post-metadata">

**Author:** ![danking](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/danking/32/43_2.png) [@danking](https://discuss.hail.is/u/danking)\
**Post date:** [April 21, 2020, 7:06pm UTC](https://discuss.hail.is/t/inconsistent-sample-qc-results/1378/4 "2020-04-21T19:06:05Z")

</div>

Hmm. I’m sorry that’s happening!

Do you run this on a private cluster, in the cloud, or on a single machine? If you run it on the cloud, what command do you use to start the cluster?

When you run it multiple times, is each time on the same cluster/machine or on different clusters/machines?

Does this happen reliably? For example, if you run the script twice in a row like: `python3 myscript.py; mv Output/callrate_geno.png Output/callrate_geno-1.png; python3 myscript.py`, does it produce different results?

Is it possible for you provide us with a script and an example dataset that produces the two different plots?

The code you’ve shared is deterministic, it should not produce different results unless `Analyses/scz_main.mt` changes.

Did you change Hail versions between runs? It’s unlikely but possible that there was a bug in one of those versions. If you have an example that differs between two versions of Hail, we could investigate that further.

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [April 22, 2020, 11:42am UTC](https://discuss.hail.is/t/inconsistent-sample-qc-results/1378/5 "2020-04-22T11:42:06Z")

</div>

> [@wonu](#):
>
> Hi, that’s the exact code, but I’ve had to run it multiple times and seem to get one of two different outputs when I plot the call rate after filtering (see attached photos)

Are these from running sample\_qc before and after the pipeline in your original post above? I’d expect the call rate to go down after your `filter_entries`, since sample\_qc call rate is defined as the number of defined calls per sample divided by the total number of variants.
