# Reducing serialization size on high-partition jobs

**URL:** <https://discuss.hail.is/t/reducing-serialization-size-on-high-partition-jobs/4254>\
**Category:** Development\
**Created:** [July 24, 2026, 2:07am UTC](https://discuss.hail.is/t/reducing-serialization-size-on-high-partition-jobs/4254 "2026-07-24T02:07:10Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Thouis](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/thouis/32/1359_2.png) [@Thouis](https://discuss.hail.is/u/Thouis)\
**Post date:** [July 24, 2026, 2:07am UTC](https://discuss.hail.is/t/reducing-serialization-size-on-high-partition-jobs/4254/1 "2026-07-24T02:07:10Z")

</div>

Hi all,

I’ve been working on some code for computing PRSes in AllOfUs, and was hitting warnings like:  
`WARN TaskSetManager: Stage 2 contains a task of very large size (1528 KiB). The maximum recommended task size is 1000 KiB.`

After some time debugging with claude, I tracked this down to jobs being launched across MTs with very many partitions (a single chromosome’s worth of AoU SNP calls, via hl.filter\_intervals). The size of the task grew with the size of the chromosome being processed.

A fairly minor change (moving RDDPartition our of RDD) seems to have fixed this:

> <https://github.com/thouis/hail/commit/ef4b2398a9b6a9dd7954a66eb5c73ea526612163>
>
> RDDPartition was defined as a case class nested inside the anonymous
> RDD in mapC…ollectPartitions, so it carried an implicit $outer field
> pointing back to the enclosing RDD instance. Every task's serialized
> Partition therefore dragged along the whole RDD's captured state
> (the full contexts array, compiled PartitionFn, reference genomes,
> etc.), not just its own data slice -- size scaling with partition
> count and triggering Spark's "task of very large size" warnings.
> 
> Moving RDDPartition to a top-level, package-private case class (same
> shape already used by TableStageToRDDPartition in RVDToTableStage.scala)
> removes the $outer capture; compute() already accesses f/fsBc/globalsBc
> directly as RDD instance state, so no functional change.

I ran `make -C hail jvm-test` - there were some failures but no new ones compared to before I changed this.

---

<div class="post-metadata">

**Author:** ![patrick-schultz](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/patrick-schultz/32/265_2.png) [@patrick-schultz](https://discuss.hail.is/u/patrick-schultz)\
**Post date:** [September 11, 2026, 11:02am UTC](https://discuss.hail.is/t/reducing-serialization-size-on-high-partition-jobs/4254/2 "2026-09-11T11:02:26Z")

</div>

Hi @Thouis,

Apologies for the delayed response. This looks like a good fix to me. I’ve gone ahead and created a PR from your branch.

Thanks for chasing down the issue!

---

<div class="post-metadata">

**Author:** ![Thouis](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/thouis/32/1359_2.png) [@Thouis](https://discuss.hail.is/u/Thouis)\
**Post date:** [September 11, 2026, 11:28am UTC](https://discuss.hail.is/t/reducing-serialization-size-on-high-partition-jobs/4254/3 "2026-09-11T11:28:46Z")

</div>

Awesome, thanks for taking it on!
