# IOException: No FileSystem for scheme: gs

**URL:** <https://discuss.hail.is/t/ioexception-no-filesystem-for-scheme-gs/2321>\
**Category:** Hail Batch & General Cloud\
**Created:** [November 1, 2021, 2:57pm UTC](https://discuss.hail.is/t/ioexception-no-filesystem-for-scheme-gs/2321 "2021-11-01T14:57:24Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![rahulch](https://avatars.discourse-cdn.com/v4/letter/r/7cd45c/32.png) [@rahulch](https://discuss.hail.is/u/rahulch)\
**Post date:** [November 1, 2021, 2:57pm UTC](https://discuss.hail.is/t/ioexception-no-filesystem-for-scheme-gs/2321/1 "2021-11-01T14:57:24Z")

</div>

while trying to run load\_dataset on AWS EMR I see below error. I am using pyspark and initiating hail.

`[hadoop@ip-172-31-101-148 ~]$ sudo pyspark Python 3.6.12 (default, May 18 2021, 22:47:55) [GCC 4.8.5 20150623 (Red Hat 4.8.5-28)] on linux Type "help", "copyright", "credits" or "license" for more information. Setting default log level to "WARN". To adjust logging level use sc.setLogLevel(newLevel). For SparkR, use setLogLevel(newLevel). 21/11/01 02:38:39 WARN HiveConf: HiveConf of name hive.server2.thrift.url does not exist 21/11/01 02:38:41 WARN Client: Neither spark.yarn.jars nor spark.yarn.archive is set, falling back to uploading libraries under SPARK_HOME. Welcome to ______ / __/__  ________ / /__ _\ \/ _ \/ _ `/ **/ '\_/  
/** / .\_\_/\_,_/_/ /_/\_\ version 2.4.4  
/_/

Using Python version 3.6.12 (default, May 18 2021 22:47:55)  
SparkSession available as ‘spark’.

> > > import hail as hl  
> > > hl.init(sc)  
> > > /usr/local/lib/python3.6/site-packages/hail/backend/backend.py:130: UserWarning: pip-installed Hail requires additional configuration options in Spark referring  
> > > to the path to the Hail Python module directory HAIL\_DIR,  
> > > e.g. /path/to/python/site-packages/hail:  
> > > spark.jars=HAIL\_DIR/hail-all-spark.jar  
> > > spark.driver.extraClassPath=HAIL\_DIR/hail-all-spark.jar  
> > > spark.executor.extraClassPath=./hail-all-spark.jar  
> > > ‘pip-installed Hail requires additional configuration options in Spark referring\n’  
> > > Running on Apache Spark version 2.4.4  
> > > SparkUI available at [http://ip-172-31-101-148.ec2.internal:4040](http://ip-172-31-101-148.ec2.internal:4040)  
> > > Welcome to  
> > > \_\_ \_\_ \<\>\_\_  
> > > / /\_/ /\_\_ \_\_/ /  
> > > / \_\_ / \_ `/ / / /_/ /_/\_,_/_/_/ version 0.2.37-7952b436bd70 LOGGING: writing to /home/hadoop/hail-20211101-0240-0.2.37-7952b436bd70.log mt = hl.experimental.load_dataset(name='dbSNP') Traceback (most recent call last): File "<stdin>", line 1, in <module> TypeError: load_dataset() missing 2 required positional arguments: 'version' and 'reference_genome' mt = hl.experimental.load_dataset(name='dbSNP',version='154',reference_genome='GRCh38') Traceback (most recent call last): File "<stdin>", line 1, in <module> File "/usr/local/lib/python3.6/site-packages/hail/experimental/datasets.py", line 33, in load_dataset with hl.hadoop_open(config_file, 'r') as f: File "<decorator-gen-6>", line 2, in hadoop_open File "/usr/local/lib/python3.6/site-packages/hail/typecheck/check.py", line 585, in wrapper return __original_func(*args_, **kwargs_) File "/usr/local/lib/python3.6/site-packages/hail/utils/hadoop_utils.py", line 79, in hadoop_open return Env.fs().open(path, mode, buffer_size) File "/usr/local/lib/python3.6/site-packages/hail/fs/hadoop_fs.py", line 12, in open handle = io.BufferedReader(HadoopReader(path, buffer_size), buffer_size=buffer_size) File "/usr/local/lib/python3.6/site-packages/hail/fs/hadoop_fs.py", line 45, in__init__self._jfile = Env.jutils().readFile(path, Env.backend()._jhc, buffer_size) File "/usr/lib/spark/python/lib/py4j-0.10.7-src.zip/py4j/java_gateway.py", line 1257, in__call__File "/usr/local/lib/python3.6/site-packages/hail/utils/java.py", line 211, in deco 'Error summary: %s' % (deepest, full, hail.__version__, deepest)) from None hail.utils.java.FatalError: IOException: No FileSystem for scheme: gs`

Any help is much appreciated

---

<div class="post-metadata">

**Author:** ![rahulch](https://avatars.discourse-cdn.com/v4/letter/r/7cd45c/32.png) [@rahulch](https://discuss.hail.is/u/rahulch)\
**Post date:** [November 1, 2021, 5:23pm UTC](https://discuss.hail.is/t/ioexception-no-filesystem-for-scheme-gs/2321/2 "2021-11-01T17:23:03Z")

</div>

> [@rahulch](#):
>
> hail.utils.java.FatalError: IOException: No FileSystem for scheme: gs`

wanted to add that I also tried below command since the default is gcp:

> > > mt = hl.experimental.load\_dataset(name=‘dbSNP’,version=‘154’,reference\_genome=‘GRCh38’,region=‘us’,cloud=‘aws’)

Traceback (most recent call last):

File “”, line 1, in

TypeError: load\_dataset() got an unexpected keyword argument ‘region’

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [November 1, 2021, 5:25pm UTC](https://discuss.hail.is/t/ioexception-no-filesystem-for-scheme-gs/2321/3 "2021-11-01T17:25:01Z")

</div>

> [@rahulch](#):
>
> version 0.2.37-7952b436bd70

You’re using a _super_ old version of Hail. if you update to latest, I think this should work.

---

<div class="post-metadata">

**Author:** ![rahulch](https://avatars.discourse-cdn.com/v4/letter/r/7cd45c/32.png) [@rahulch](https://discuss.hail.is/u/rahulch)\
**Post date:** [November 1, 2021, 6:16pm UTC](https://discuss.hail.is/t/ioexception-no-filesystem-for-scheme-gs/2321/4 "2021-11-01T18:16:32Z")

</div>

> [@rahulch](#):
>
> 7952b436bd70

tried that but still the same. I’m thinking of doing a fresh installation unless you can catch my mistake

python3.6 -m pip install hail --upgrade

> > > hl.init(sc)  
> > > /usr/local/lib/python3.6/site-packages/hail/backend/backend.py:130: UserWarning: pip-installed Hail requires additional configuration options in Spark referring  
> > > to the path to the Hail Python module directory HAIL\_DIR,  
> > > e.g. /path/to/python/site-packages/hail:  
> > > spark.jars=HAIL\_DIR/hail-all-spark.jar  
> > > spark.driver.extraClassPath=HAIL\_DIR/hail-all-spark.jar  
> > > spark.executor.extraClassPath=./hail-all-spark.jar  
> > > ‘pip-installed Hail requires additional configuration options in Spark referring\n’  
> > > Running on Apache Spark version 2.4.4  
> > > SparkUI available at [http://ip-172-31-101-148.ec2.internal:4040](http://ip-172-31-101-148.ec2.internal:4040)  
> > > Welcome to  
> > > \_\_ \_\_ \<\>\_\_  
> > > / /_/ /\_\_ \_\_/ /  
> > > / \_\_ / \_ `/ / /  
> > > /_/ /_/\_,_/_/_/ version 0.2.37-7952b436bd70  
> > > LOGGING: writing to /home/hadoop/hail-20211101-1814-0.2.37-7952b436bd70.log

---

<div class="post-metadata">

**Author:** ![tpoterba](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/tpoterba/32/61_2.png) [@tpoterba](https://discuss.hail.is/u/tpoterba)\
**Post date:** [November 1, 2021, 6:46pm UTC](https://discuss.hail.is/t/ioexception-no-filesystem-for-scheme-gs/2321/5 "2021-11-01T18:46:45Z")

</div>

Can you do:

```auto
pip show hail

```

and

```auto
python3.6 -c "import hail as hl; print(hl. __file__ )"

```

---

<div class="post-metadata">

**Author:** ![rahulch](https://avatars.discourse-cdn.com/v4/letter/r/7cd45c/32.png) [@rahulch](https://discuss.hail.is/u/rahulch)\
**Post date:** [November 1, 2021, 7:03pm UTC](https://discuss.hail.is/t/ioexception-no-filesystem-for-scheme-gs/2321/6 "2021-11-01T19:03:52Z")

</div>

Thanks @tpoterba

[hadoop@ip-172-31-101-148 ~]$ pip show hail

Name: hail  
Version: 0.2.74  
Summary: Scalable library for exploring and analyzing genomic data.  
Home-page: [https://hail.is](https://hail.is)  
Author: Hail Team  
Author-email: [hail@broadinstitute.org](mailto:hail@broadinstitute.org)  
License: UNKNOWN  
Location: /home/hadoop/.local/lib/python3.6/site-packages  
Requires: aiohttp, aiohttp-session, asyncinit, bokeh, boto3, botocore, decorator, Deprecated, dill, fsspec, gcsfs, google-cloud-storage, humanize, hurry.filesize, janus, nest-asyncio, numpy, pandas, parsimonious, PyJWT, pyspark, python-json-logger, requests, scipy, tabulate, tqdm  
Required-by:  
[hadoop@ip-172-31-101-148 ~] [hadoop@ip-172-31-101-148 ~] python3.6 -c “import hail as hl; print(hl. **file** )”  
Traceback (most recent call last):  
File “”, line 1, in   
File “/home/hadoop/.local/lib/python3.6/site-packages/hail/ **init**.py”, line 44, in   
from .table import Table, GroupedTable, asc, desc # noqa: E402  
File “/home/hadoop/.local/lib/python3.6/site-packages/hail/table.py”, line 3, in   
import pandas  
File “/home/hadoop/.local/lib/python3.6/site-packages/pandas/ **init**.py”, line 22, in   
from pandas.compat.numpy import (  
File “/home/hadoop/.local/lib/python3.6/site-packages/pandas/compat/numpy/ **init**.py”, line 21, in   
“this version of pandas is incompatible with numpy \< 1.15.4\n”  
ImportError: this version of pandas is incompatible with numpy \< 1.15.4  
your numpy version is 1.14.5.  
Please upgrade numpy to \>= 1.15.4 to use this pandas version  
[hadoop@ip-172-31-101-148 ~]$

---

<div class="post-metadata">

**Author:** ![danking](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.hail.is/danking/32/43_2.png) [@danking](https://discuss.hail.is/u/danking)\
**Post date:** [November 2, 2021, 2:39pm UTC](https://discuss.hail.is/t/ioexception-no-filesystem-for-scheme-gs/2321/7 "2021-11-02T14:39:08Z")

</div>

Hey @rahulch ,

How are you installing Hail? I don’t expect Hail-from-pip to work properly on EMR. I believe you have. three options:

- Some folks at harvard med [have some scripts](https://github.com/hms-dbmi/hail-on-EMR) for running Hail on Amazon.
- You could also [install Hail from source](https://hail.is/docs/0.2/install/other-cluster.html) on the master node of the spark cluster.
- We maintain a tool, `hailctl` (which is included in the Hail python package) for running Hail on Google Dataproc, if you’re able to use Google Cloud instead.
