Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 3 additions & 5 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -2,12 +2,12 @@ files: \.py$

repos:
- repo: https://github.com/PyCQA/autoflake
rev: v2.3.3
rev: v2.4.0
hooks:
- id: autoflake
args: [--in-place]
- repo: https://github.com/PyCQA/isort
rev: 9.0.0b1
rev: 9.0.1
hooks:
- id: isort
args: [--thirdparty, neo4j]
Expand All @@ -16,9 +16,7 @@ repos:
hooks:
- id: autopep8
- repo: https://github.com/PyCQA/docformatter
# Pin to v1.7.7 until this is resolved: https://github.com/PyCQA/docformatter/issues/344
# Messes up some of our query strings.
rev: v1.7.7
rev: v1.7.8
hooks:
- id: docformatter
args: [--in-place, --wrap-summaries, '88', --wrap-descriptions, '88']
Expand Down
11 changes: 11 additions & 0 deletions ACKNOWLEDGMENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,17 @@ We use the top 1M websites per country from the Google [Chrome User Experience R
This data is licensed under [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/).
No changes were made to the data.

## Hagezi

We use the [DNS blocklists](https://github.com/hagezi/dns-blocklists) maintained by
Hagezi. The names of these lists are resolved by IHR, which publishes the results
[here](https://github.com/InternetHealthReport/hagezi-blocklists-forward-dns). List
membership is attributed to Hagezi in IYP, the DNS data to IHR.

The blocklists are licensed under [GPL-3.0](https://www.gnu.org/licenses/gpl-3.0.html)
and the resolution results under [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0).
No changes were made to the data.

## Internet Health Report

We use three datasets from the [Internet Health Report](https://ihr.iijlab.net/) (that's
Expand Down
1 change: 1 addition & 0 deletions config.json.example
Original file line number Diff line number Diff line change
Expand Up @@ -118,6 +118,7 @@
"iyp.crawlers.google.crux_top1m_country",
"iyp.crawlers.iana.address_space",
"iyp.crawlers.openintel.dnsgraph",
"iyp.crawlers.hagezi.blocklists_forward_dns",
"iyp.crawlers.cloudflare.dns_top_locations",
"iyp.crawlers.cloudflare.dns_top_ases",
"iyp.crawlers.ooni.facebookmessenger",
Expand Down
1 change: 1 addition & 0 deletions documentation/data-sources.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@
| | Cloudflare Radar API endpoint radar/ranking/top (top 100 domain names)| https://radar.cloudflare.com | [README](https://github.com/InternetHealthReport/internet-yellow-pages/tree/main/iyp/crawlers/cloudflare#readme)| cloudflare.top100 |
| Emile Aben | AS names| https://github.com/emileaben/asnames | [README](https://github.com/InternetHealthReport/internet-yellow-pages/tree/main/iyp/crawlers/emileaben#readme)| emileaben.as_names |
| Google | CrUX top 1M websites per country| https://developer.chrome.com/docs/crux | [README](https://github.com/InternetHealthReport/internet-yellow-pages/tree/main/iyp/crawlers/google#readme) | google.crux_top1m_country |
| Hagezi | DNS blocklists, DNS resolution by IHR | https://github.com/hagezi/dns-blocklists | [README](https://github.com/InternetHealthReport/internet-yellow-pages/tree/main/iyp/crawlers/hagezi#readme) | hagezi.blocklists_forward_dns |
| IHR | AS Hegemony| https://www.ihr.live/en/documentation#AS-dependency | [README](https://github.com/InternetHealthReport/internet-yellow-pages/tree/main/iyp/crawlers/ihr#readme) | ihr.local_hegemony_v4, ihr.local_hegemony_v6 |
| | Country Dependency| https://www.ihr.live/en/documentation#Country-s-network-dependency | [README](https://github.com/InternetHealthReport/internet-yellow-pages/tree/main/iyp/crawlers/ihr#readme) | ihr.country_dependency |
| | ROV| https://www.ihr.live/en/documentation#Route-Origin-Validation | [README](https://github.com/InternetHealthReport/internet-yellow-pages/tree/main/iyp/crawlers/ihr#readme) | ihr.rov |
Expand Down
60 changes: 60 additions & 0 deletions iyp/crawlers/hagezi/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
# Hagezi DNS blocklists -- https://github.com/hagezi/dns-blocklists

[Hagezi's DNS blocklists](https://github.com/hagezi/dns-blocklists) are curated lists of
domain names grouped by the kind of content or behavior they are associated with (fake
shops, pop-up ads, threats, newly registered domains, DoH resolvers, dynamic DNS, URL
shorteners, piracy, gambling, social networks, NSFW).

IHR resolves hostnames in these lists (https://github.com/InternetHealthReport/hagezi-blocklists-forward-dns)
and publishes the A/AAAA records, the zone of each domain name, and the
authoritative name servers.

## Graph representation

**Blocklist membership:**

```Cypher
(:HostName {name: 'bit.ly'})-[:CATEGORIZED]->(:Tag {label: 'url-shortener'})
```

The Tag label is the name of the Hagezi list the name was taken from.
A name can be part of several lists.

List membership is Hagezi's data, so these relationships carry
`reference_org: 'Hagezi'` with a `reference_url_info` pointing to the [blocklist
repository](https://github.com/hagezi/dns-blocklists). All other relationships describe
the DNS measurement performed by IHR and carry `reference_org: 'IHR'`.

For CATEGORIZED relationships `reference_time_modification` is the date of the Hagezi
commit the names were fetched from (the `source.commit_date` field of the result file).

**IP resolution for hostnames:**

```Cypher
(:HostName {name: 'bit.ly'})-[:RESOLVES_TO]->(:IP {ip: '67.199.248.10'})
```

**Zone of a hostname:**

```Cypher
(:HostName {name: 'www.example.com'})-[:PART_OF]->(:DomainName {name: 'example.com'})
```

**Authoritative name servers managing zones:**

```Cypher
(:DomainName {name: 'example.com'})-[:MANAGED_BY]->(:HostName:AuthoritativeNameServer {name: 'a.iana-servers.net'})
```

Name servers are HostName nodes with the additional AuthoritativeNameServer label, and
their IPs are attached with RESOLVES_TO relationships as above.

## Dependence

This crawler is not depending on other crawlers.

## Notes

The crawler uses the GitHub API to find the most recent scan directory, since scans are
not published on a fixed schedule. It fails if the most recent scan is older than 30
days, which indicates that the measurement pipeline stopped working.
Loading
Loading