# QQ plot downsampling

**URL:** <https://discuss.hail.is/t/qq-plot-downsampling/1278>\
**Category:** Hail Query & hailctl\
**Created:** [February 7, 2020, 3:00pm UTC](https://discuss.hail.is/t/qq-plot-downsampling/1278 "2020-02-07T15:00:55Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![pavlos](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/pavlos/32/303_2.png) [@pavlos](https://discuss.hail.is/u/pavlos)\
**Post date:** [February 7, 2020, 3:00pm UTC](https://discuss.hail.is/t/qq-plot-downsampling/1278/1 "2020-02-07T15:00:55Z")

</div>

Hi hail team,  
is it possible to downsample the QQ plot for specific p-values?  
I would like to show sparse datapoints for p-values \<0.01 and dense datapoints for p-values \>0.01.  
Is this something I can do using the hl.plot.qq function?

q = hl.plot.qq(gwas.p\_value, collect\_all=False, n\_divisions=100, title=f"{pheno\_name} QQ plot")

Thanks,  
Pavlos

---

<div class="post-metadata">

**Author:** ![johnc1231](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/johnc1231/32/286_2.png) [@johnc1231](https://discuss.hail.is/u/johnc1231)\
**Post date:** [February 7, 2020, 3:04pm UTC](https://discuss.hail.is/t/qq-plot-downsampling/1278/2 "2020-02-07T15:04:32Z")

</div>

Others can feel free to chime in if there’s some argument I don’t know about, but my intuition is that the way to do this is to filter or downsample your `gwas` table in advance.

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [February 7, 2020, 3:05pm UTC](https://discuss.hail.is/t/qq-plot-downsampling/1278/3 "2020-02-07T15:05:38Z")

</div>

Is the standard downsampling approach causing problems? If so, what are those problems?

---

<div class="post-metadata">

**Author:** ![pavlos](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/pavlos/32/303_2.png) [@pavlos](https://discuss.hail.is/u/pavlos)\
**Post date:** [February 7, 2020, 3:08pm UTC](https://discuss.hail.is/t/qq-plot-downsampling/1278/4 "2020-02-07T15:08:57Z")

</div>

It’s a request from faculty members here to have more datapoints for higher p-values as it’s a more correct qq-plot, specifically they would like me to set a cutoff of p-value 0.01. Perhaps I can do that by downsampling the gwas table in advance as @johnc1231 suggests. I was just wondering if there was any other way you have done this before.

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [February 7, 2020, 3:11pm UTC](https://discuss.hail.is/t/qq-plot-downsampling/1278/5 "2020-02-07T15:11:10Z")

</div>

> It’s a request from faculty members here to have more datapoints for higher p-values as it’s a more correct qq-plot

The `hl.agg.downsample` aggregator called by `qq` already does this – it does “visual” downsampling, merging points in adjacent pixels, essentially. Unless you have two points with exactly the same p-value, the right tail of the qq plot is going to be at full resolution.

---

<div class="post-metadata">

**Author:** ![pavlos](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/pavlos/32/303_2.png) [@pavlos](https://discuss.hail.is/u/pavlos)\
**Post date:** [February 7, 2020, 3:39pm UTC](https://discuss.hail.is/t/qq-plot-downsampling/1278/6 "2020-02-07T15:39:06Z")

</div>

that’s great news @tpoterba! pleased I won’t have to manually downsample the dataset. Thank you again for all your help.
