# Internet Measurement

> Infrastructure surveys, topology measurement, RPKI coverage, root-server and TLD ecosystems: reproducible cross-layer research queries on WhisperGraph.

*Source: https://www.whisper.security/docs/recipes/research*
*Published: 2026-05-04*
*Last updated: 2026-10-06*

---
You study the internet itself: topology, deployment trends, ecosystem structure, not one incident at a time but in aggregate, across billions of edges. The hard part isn't the analysis; it's getting clean, joined, planet-scale data to analyze. These recipes take you to the joined data: DNS, BGP peering and observed paths, RPKI, the root-server and registry ecosystems, and the physical internet, all in one Cypher surface. The recipes below are bulk and aggregate queries, written with the bounding that keeps them fast on a graph of 7.5B nodes and 39.8B edges.

A few rules that keep research queries honest at this scale:

- **Anchor or aggregate, never bare-scan a large label.** A query that touches all of `HOSTNAME`, `IPV4` or `NAMESERVER_FOR` without an anchored start will not finish. Anchor on a `{name:"..."}` node, or aggregate behind a `CALL db.*` histogram or a `whisper.*` ranking procedure.
- **Bound high-fan-out hops** with `WITH ... LIMIT` before you expand again.
- **Small reference labels are safe to scan.** `FEED_SOURCE`, `CATEGORY`, `VENDOR`, `CDN_POP`, `DNS_ROOT_INSTANCE` and `THREAT_SIGNAL_TYPE` are catalogues, and listing them is instant. Reach the large labels through an edge.
- **Treat `LINKS_TO` as a sampled crawl layer, not a web-scale link graph.** Anchor on a host, read its degree as a floor, and do not build a link-structure study on it alone.

See [Getting Started](/docs/whisper-graph/getting-started) for keys, the [Graph Schema](/docs/whisper-graph/schema) for the full label/edge model, and the [Procedures](/docs/whisper-graph/procedures) reference for the `CALL` surface.

> **Run it live.** Several of these measurement jobs have a guided, browser-runnable version that opens with a result on your own indicator:
> - [Digital Infrastructure Mapping](/products/intelligence/use-cases/infrastructure-supply-chain/infrastructure-mapping) — one indicator mapped to its owner and full footprint across every layer.
> - [Supply-Chain Dependency Mapping](/products/intelligence/use-cases/infrastructure-supply-chain/supply-chain) — every external provider a domain depends on, grouped by function, with dependency chains and single-vendor (SPOF) signals across facilities, cable landings and subsea cables.
> - [Investigate an Indicator](/products/intelligence/use-cases/threat-investigation/indicator) — verdict, hosting, routing, and the shared infrastructure around any domain, IP, ASN, or prefix.
>
> More live flows are on the [Research & OSINT use cases](/products/intelligence/use-cases/research-osint) page.

**Key concepts:** [BGP routing](/glossary/bgp-routing) · [RPKI ROA](/glossary/rpki-roa) · [MOAS conflict](/glossary/moas-conflict) · [Internet exchange point](/glossary/internet-exchange-point) · [Submarine cable](/glossary/submarine-cable) · [Tor exit node](/glossary/tor-exit-node) · [MITRE ATT&CK](/glossary/mitre-attack).

---

## Schema exploration

### What's actually in the graph

Before you write a traversal, confirm the label and edge exist. The most common cause of an empty result set is anchoring on a label that doesn't exist (there is no `DOMAIN` or `FQDN` label; every name is a `HOSTNAME`). The `db.*` procedures return precomputed histograms, instant even at this scale.

```cypher expect=rows>0 verified=2026-09-02
// Every node label with its live count
CALL db.labels()
```
**Returns:** `label, count`

**Sample output** (the five largest of 43 labels):
```json
[
  {"label": "HOSTNAME", "count": 2752403048},
  {"label": "IPV4", "count": 621441120},
  {"label": "EMAIL", "count": 237065663},
  {"label": "ORGANIZATION", "count": 119189847},
  {"label": "PHONE", "count": 60194142}
]
```

```cypher expect=rows>0 verified=2026-09-02
// Every edge type with its live count, source and target labels
CALL db.relationshipTypes()
```

**Sample output** (the five largest of 55 edge types, `sourceLabels`/`targetLabels` elided):
```json
[
  {"type": "NAMESERVER_FOR", "count": 9173662411},
  {"type": "ANNOUNCED_BY", "count": 4331089630},
  {"type": "RESOLVES_TO", "count": 3125689316},
  {"type": "CHILD_OF", "count": 2451196569},
  {"type": "REGISTERED_BY", "count": 916255242}
]
```

```cypher expect=rows>0 verified=2026-09-02
// Every property name in the graph
CALL db.propertyKeys() YIELD propertyKey RETURN propertyKey ORDER BY propertyKey LIMIT 200
```

**Costs:** milliseconds; histogram reads, no traversal; the column for edges is `type`, not `relationshipType`.

> **Why this matters:** the edge histogram *is* a research dataset, the shape of the global internet one `CALL` away. Use these counts to plan which traversals are cheap (anchored) and which need aggregation, and read `sourceLabels`/`targetLabels` off `db.relationshipTypes()` to learn an edge's direction before you write it.

**From here, →** [Confirm a property before you filter on it](#confirm-a-property-before-you-filter-on-it).

### Confirm a property before you filter on it

Filtering on a property that doesn't exist returns empty, not an error. A quick `keys()` read (or `db.propertyKeys()`) saves you from `WHERE h.fqdn = ...` when the property is `name`.

```cypher expect=rows>0 seed=185.220.101.1 verified=2026-09-02
// What properties does a threat-listed IP carry?
MATCH (ip:IPV4 {name: "185.220.101.1"})
RETURN keys(ip) AS properties
LIMIT 1
```
**Returns:** `properties`

**Sample output** (trimmed):
```json
[{"properties": ["id", "label", "name", "threatScore", "threatSources", "isThreat", "isTor", "threatLevel", "verdictLevel", "verdictScore", "verdictCoverage", "verdictBlocking"]}]
```

**Costs:** milliseconds; one indexed anchor and a property read; the key set differs by label, so check the label you are about to filter.

**From here, →** [The threat feed catalog](#the-threat-feed-catalog).

### The threat feed catalog

Feed coverage is a study in itself. The catalogue is small enough to list directly, and any indicator's `LISTED_IN` edges tell you which feeds and categories cover it.

```cypher expect=rows>0 verified=2026-09-02
// All threat-feed sources, by display name
MATCH (f:FEED_SOURCE)
RETURN f.displayName AS feed
ORDER BY f.displayName
LIMIT 15
```
**Returns:** `feed`

**Sample output**:
```json
[{"feed": "1Hosts Xtra"}, {"feed": "AlienVault Reputation"}, {"feed": "Bad Hosting ASN"}]
```

```cypher expect=rows>0 seed=185.220.101.1 verified=2026-09-02
// Which feeds and categories cover a known Tor exit?
MATCH (ip:IPV4 {name: "185.220.101.1"})-[:LISTED_IN]->(f:FEED_SOURCE)
MATCH (f)-[:BELONGS_TO]->(c:CATEGORY)
RETURN c.displayName AS category, collect(f.displayName) AS feeds
ORDER BY category
LIMIT 25
```

**Sample output**:
```json
[
  {"category": "General Blacklists", "feeds": ["GreenSnow Blacklist", "IPsum", "FireHOL Level 2", "duggytuxy-datashield-critical"]},
  {"category": "Spam", "feeds": ["StopForumSpam Listed IPs (7 day)"]},
  {"category": "TOR Network", "feeds": ["Tor Exit Nodes"]}
]
```

**Costs:** milliseconds; a scan of a small reference label, and two anchored hops for the per-indicator view; swap `FEED_SOURCE` for `CATEGORY` to list the categories.

> The graph indexes **134 feeds across 32 categories** with 14.5M `LISTED_IN` edges. The catalog spans block lists *and* trust lists, so the same query surface answers "known-bad?" and "known-good?". `.name` is the slug (`firehol-level2`, `tor`) and `.displayName` the readable label; the full list is in [Threat Feeds & Categories](/docs/whisper-graph/threat-feeds).

**From here, →** [BGP peering-degree, network by network](#bgp-peering-degree-network-by-network).

---

## Internet topology

### BGP peering-degree, network by network

Peering data lives in PeeringDB and route-collector dumps you have to download, parse and join yourself. In the graph, `BGP_NEIGHBOR` is the canonical ASN↔ASN adjacency edge, already materialized. Anchor on a set of ASNs and count peers in one round-trip.

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
// Peering degree for a sample of well-known networks
UNWIND ["AS13335", "AS3356", "AS15169", "AS2914"] AS asn_name
MATCH (a:ASN {name: asn_name})-[:BGP_NEIGHBOR]->(peer:ASN)
RETURN asn_name, count(peer) AS peer_count
ORDER BY peer_count DESC
```
**Returns:** `asn_name, peer_count`

**Sample output**:
```json
[
  {"asn_name": "AS3356", "peer_count": 6196},
  {"asn_name": "AS2914", "peer_count": 1461},
  {"asn_name": "AS13335", "peer_count": 1284},
  {"asn_name": "AS15169", "peer_count": 139}
]
```

**Costs:** milliseconds; one anchored hop per element, aggregated; `PEERS_WITH` still resolves as an alias, but write `BGP_NEIGHBOR`.

> **Reading it:** transit-heavy carriers (AS3356 Lumen, AS2914 NTT) sit at the top of the degree distribution; content networks (AS15169 Google) peer selectively. The degree gap is the structural difference between transit and content ASNs, visible in one query. For the whole distribution rather than a sample, `CALL whisper.bgpDegreeDistribution()` returns one row per `(inDegree, outDegree)` pair; see [BGP & RPKI](/docs/recipes/bgp-routing).

**From here, →** [Rank the densest networks by prefix count](#rank-the-densest-networks-by-prefix-count).

### Rank the densest networks by prefix count

For a top-of-distribution view without enumerating every ASN, use the ranking procedure. It reads a precomputed ranking, so it's instant where the equivalent per-ASN sweep would be expensive.

```cypher expect=rows>0 verified=2026-09-02
// The ASNs announcing the most prefixes
CALL whisper.topAsnsByPrefixCount(15)
YIELD asn, prefixCount
RETURN asn, prefixCount
LIMIT 15
```
**Returns:** `asn, prefixCount`

**Sample output**:
```json
[
  {"asn": "AS16509", "prefixCount": 22511},
  {"asn": "AS9808", "prefixCount": 21519},
  {"asn": "AS577", "prefixCount": 16258}
]
```

To study a single network's footprint, anchor and bound the fan-out:

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
// How many prefixes does Cloudflare announce?
MATCH (a:ASN {name: "AS13335"})-[:ROUTES]->(ap:ANNOUNCED_PREFIX)
RETURN count(ap) AS announced_prefixes
LIMIT 1
```

**Costs:** milliseconds; a precomputed ranking, then one anchored aggregated hop; the argument is a number of networks, not an ASN.

> The ordering shifts as the routing table does, so cite it with a date.

**From here, →** [Second-degree peering reach](#second-degree-peering-reach).

### Second-degree peering reach

The peering graph's structure shows up in the two-hop neighborhood: how many distinct networks are within two BGP hops. Bound the first hop hard before expanding; a large carrier's neighbor set is in the thousands.

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
// Distinct networks reachable within two BGP hops of Cloudflare
MATCH (a:ASN {name: "AS13335"})-[:BGP_NEIGHBOR]->(n1:ASN)
WITH DISTINCT n1 LIMIT 500
MATCH (n1)-[:BGP_NEIGHBOR]->(n2:ASN)
RETURN count(DISTINCT n2) AS two_hop_reach
LIMIT 1
```
**Returns:** `two_hop_reach`

**Sample output**:
```json
[{"two_hop_reach": 33555}]
```

**Costs:** milliseconds; two explicit hops with a `WITH ... LIMIT 500` between them, which is load-bearing.

> **Tip.** `BGP_NEIGHBOR` also works inside a bounded variable-length pattern: `MATCH p = (a:ASN {name: "AS13335"})-[:BGP_NEIGHBOR*2..2]->(n:ASN) WHERE n <> a RETURN n.name, length(p) LIMIT 10` samples the second ring directly. Keep `WHERE n <> a`: a peering mesh is undirected in practice, so a two-hop walk routinely lands back on the origin.

**From here, →** [ASN home jurisdiction distribution](#asn-home-jurisdiction-distribution).

### ASN home jurisdiction distribution

`HAS_COUNTRY` runs from `ASN` straight to `COUNTRY`, no city hop needed for the AS's registered jurisdiction. Aggregate across a sample to study where networks are domiciled.

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
// Home country for a sample of networks
UNWIND ["AS13335", "AS15169", "AS3356", "AS2914", "AS4837"] AS asn_name
MATCH (a:ASN {name: asn_name})-[:HAS_COUNTRY]->(c:COUNTRY)
RETURN c.name AS country, count(*) AS networks
ORDER BY networks DESC
```
**Returns:** `country, networks`

**Sample output**:
```json
[{"country": "US", "networks": 4}, {"country": "CN", "networks": 1}]
```

**Costs:** milliseconds; one anchored hop per element, aggregated.

**From here, →** [Which countries hold the most autonomous systems?](#which-countries-hold-the-most-autonomous-systems).

### Which countries hold the most autonomous systems?

Characterising the shape of the routed internet, you want the distribution of autonomous systems by registered country across the whole population, not a sample. The ranking procedure reads it precomputed.

```cypher expect=rows>0 verified=2026-09-02
// Which countries hold the most autonomous systems
CALL whisper.asnCountries(10) YIELD country, asns
RETURN country, asns
ORDER BY asns DESC
LIMIT 10
```
**Returns:** `country, asns`

**Sample output**:
```json
[
  {"country": "US", "asns": 31320},
  {"country": "BR", "asns": 8933},
  {"country": "IN", "asns": 6113},
  {"country": "RU", "asns": 5620}
]
```

**Costs:** milliseconds; a precomputed ranking; the argument is the number of countries to return, not a country code.

> These are *registered* countries from the routing registries, a legal-entity fact rather than a physical one: a network registered in one country routinely announces prefixes that terminate somewhere else. For where the traffic actually lands, join through the facility and exchange layer below.

**From here, →** [Outbound link degree from a domain](#outbound-link-degree-from-a-domain).

---

## The hyperlink layer (LINKS_TO)

`LINKS_TO` is a crawl-derived hyperlink edge between hostnames, in the same query surface as DNS, WHOIS and BGP, which means you can join link structure against routing without exporting a CSV. It is a sample of the crawled web, not a census of it. **The one rule: always anchor**, and read every degree as a floor.

### Outbound link degree from a domain

How many distinct hosts a site links out to is the simplest link-structure measurement, and anchored on the host it is an instant read.

```cypher expect=rows>0 seed=github.com verified=2026-09-02
// How many distinct hosts does github.com link out to?
MATCH (h:HOSTNAME {name: "github.com"})-[:LINKS_TO]->(target:HOSTNAME)
RETURN count(DISTINCT target) AS outbound_hosts
LIMIT 1
```
**Returns:** `outbound_hosts`

**Sample output**:
```json
[{"outbound_hosts": 189369}]
```

**Costs:** milliseconds; one anchored hop, aggregated; degree reflects what the crawl sample captured for that host.

**From here, →** [Join the link graph to routing — what networks does a site link out to?](#join-the-link-graph-to-routing-what-networks-does-a-site-link-out-to).

### Join the link graph to routing — what networks does a site link out to?

The join this makes: follow each outbound link to where it actually resolves and who routes it. Cap the link fan-out first, then traverse DNS→BGP for each.

```cypher expect=rows>0 seed=github.com verified=2026-09-02
// Outbound links → resolve each target → which networks host them
MATCH (h:HOSTNAME {name: "github.com"})-[:LINKS_TO]->(target:HOSTNAME)
WITH DISTINCT target LIMIT 200
MATCH (target)-[:RESOLVES_TO]->(ip:IPV4)-[:ANNOUNCED_BY]->(ap:ANNOUNCED_PREFIX)-[:ROUTES]->(a:ASN)
RETURN a.name AS asn, count(DISTINCT target) AS linked_hosts
ORDER BY linked_hosts DESC
LIMIT 15
```
**Returns:** `asn, linked_hosts`

**Sample output**:
```json
[
  {"asn": "AS13335", "linked_hosts": 43},
  {"asn": "AS16509", "linked_hosts": 38},
  {"asn": "AS396982", "linked_hosts": 8}
]
```

**Costs:** milliseconds; a four-layer join (link → DNS → BGP announcement → ASN) kept bounded by `WITH DISTINCT target LIMIT 200`; sign in to run the three-hop leg.

> **Why it's hard otherwise:** this crosses three datasets that normally live in three different tools. Here it's one statement. Widen the sample deliberately, and re-anchor on a linked host to walk further rather than writing a variable-length pattern over `LINKS_TO`.

**From here, →** [Is a network's announced space ROA-covered?](#is-a-network-s-announced-space-roa-covered).

---

## RPKI coverage

RPKI Route Origin Authorizations (`ROA`) are first-class nodes (3M of them), linked to the prefixes and origin ASNs they authorize. You can study deployment and validity without pulling and parsing the RIR trust-anchor dumps yourself.

### Is a network's announced space ROA-covered?

Cross-referencing announced prefixes against the RPKI repository means reconciling two separate feeds. In the graph both are nodes; `ROA_AUTHORIZES_ORIGIN` joins them, and the prefix each ROA covers is a property on the ROA itself.

```cypher expect=static seed=AS13335 verified=2026-10-06 reason="RPKI authorization data is not always present on a read, so this block shows a captured result rather than a live run"
// RPKI ROAs that authorize Cloudflare as an origin AS, with the max-length they permit
MATCH (roa:ROA)-[:ROA_AUTHORIZES_ORIGIN]->(a:ASN {name: "AS13335"})
WHERE roa.authSource = "rpki-roa"
RETURN roa.prefix AS authorized_prefix, roa.maxLength AS max_length,
       roa.trustAnchor AS trust_anchor
LIMIT 25
```
**Returns:** `authorized_prefix, max_length, trust_anchor`

**Sample output** (captured 2026-10-06):
```json
[
  {"authorized_prefix": "102.219.82.0/24", "max_length": 24, "trust_anchor": "afrinic"},
  {"authorized_prefix": "154.193.133.0/24", "max_length": 24, "trust_anchor": "afrinic"},
  {"authorized_prefix": "154.193.184.0/24", "max_length": 24, "trust_anchor": "afrinic"}
]
```

**Costs:** milliseconds; one inbound hop from an indexed ASN plus property reads; `count(roa)` sizes the set before you list it.

> **Empty result:** no rows means no ROA names this ASN as an origin, which is the unsigned state, not missing data. Zero rows is never a verdict.

> A `ROA` node has no `name`; its identity is the `(prefix, asn)` pair it authorizes. Two kinds of record share the label, told apart by `authSource`. An RPKI ROA (`rpki-roa`) carries `id, label, authSource, asn, prefix, maxLength, trustAnchor, validUntil`. A route object from an internet routing registry (`irr-route`) carries `id, label, authSource, asn, prefix, mnt_by, source_rir` and no `maxLength`, `trustAnchor` or `validUntil`, and it is the larger share of the label. `keys(roa)` shows which kind a sample is, and `WHERE roa.authSource = "rpki-roa"` keeps a study to RPKI.

**From here, →** [Cross-check an announcement against its authorization](#cross-check-an-announcement-against-its-authorization).

### Cross-check an announcement against its authorization

Walk from an IP to its announced prefix, then read the precomputed validation state and count the ROAs covering that exact prefix: the building block of route-origin validation, written as explicit single hops.

```cypher expect=rows>0,no-null-columns seed=1.1.1.1 verified=2026-09-30
// Does the prefix covering 1.1.1.1 validate, and which RPKI ROAs cover it?
MATCH (ip:IPV4 {name: "1.1.1.1"})-[:ANNOUNCED_BY]->(ap:ANNOUNCED_PREFIX)
OPTIONAL MATCH (roa:ROA)-[:ROA_AUTHORIZES_PREFIX]->(ap)
WHERE roa.authSource = "rpki-roa"
RETURN ap.name AS announced_prefix, ap.rpkiStatus AS rpki_status,
       ap.roaAsn AS roa_asn, ap.roaMaxLength AS roa_max_length,
       count(roa) AS covering_roas, collect(DISTINCT roa.asn) AS roa_origins
LIMIT 5
```
**Returns:** `announced_prefix, rpki_status, roa_asn, roa_max_length, covering_roas, roa_origins`

**Sample output:**
```json
[{"announced_prefix": "1.1.1.0/24", "rpki_status": "valid", "roa_asn": 13335, "roa_max_length": 24, "covering_roas": 1, "roa_origins": [13335]}]
```

**Costs:** milliseconds; one anchored hop plus one optional ROA hop; `rpkiStatus` takes `valid`, `invalid` or `not-found`, and `not-found` is the unsigned population.

> **Empty result:** `covering_roas: 0` with `rpki_status: "not-found"` is unsigned space, not a missing edge. Zero rows is never a verdict.

**From here, →** [MOAS conflicts — the early hijack signal](#moas-conflicts-the-early-hijack-signal).

### MOAS conflicts — the early hijack signal

A prefix announced by more than one origin AS is the leading early signal of a BGP hijack or route leak. For a study you want the population, not one network, so anchor on the conflict edge itself, bound it, and rank by how many origins are competing.

```cypher expect=static verified=2026-09-02 reason="multi-origin conflict state changes as prefixes are withdrawn and re-announced, so this block shows a captured result rather than a live run"
// The most heavily contested prefixes, and how many origins announce them
MATCH (p:ANNOUNCED_PREFIX)-[:CONFLICTS_WITH]->(other:ASN)
WITH DISTINCT p LIMIT 3000
MATCH (p)-[:CONFLICTS_WITH]->(o:ASN)
WITH p, collect(DISTINCT o.name) AS conflicting_origins
WHERE size(conflicting_origins) > 1
RETURN p.name AS prefix, size(conflicting_origins) AS origins,
       conflicting_origins[0..6] AS sample_origins
ORDER BY origins DESC
LIMIT 25
```
**Returns:** `prefix, origins, sample_origins`

**Sample output:**
```json
[
  {"prefix": "192.58.128.0/24", "origins": 21, "sample_origins": ["AS396549", "AS396738", "AS396739", "AS396707", "AS396576", "AS396686"]},
  {"prefix": "192.30.45.0/24", "origins": 12, "sample_origins": ["AS396549", "AS396578", "AS20362", "AS211369", "AS396555", "AS396566"]}
]
```

**Costs:** milliseconds with the `WITH DISTINCT p LIMIT 3000` bound; the same aggregation without the bound runs for many seconds.

> **Read the tail before you read the head.** The graph holds 11,049 `CONFLICTS_WITH` edges, and the ranking is dominated by prefixes that are *supposed* to have many origins: `192.58.128.0/24` (J-root) is anycast working correctly, not a stack of hijacks. A MOAS study's real work is separating anycast and legitimate multi-homing from the two- or three-origin cases that are anomalies; `p.moasIsLegitimate` is the graph's own read on that. Pair it with the per-announcement RPKI state (`rpkiStatus` / `roaAsn` / `roaMaxLength`): a MOAS conflict where one origin is RPKI-invalid is a far stronger signal than the conflict alone. See [BGP & RPKI](/docs/recipes/bgp-routing).
>
> Starting from a named network instead (`MATCH (a:ASN {name: "…"})-[:ROUTES]->(p) WHERE p.isMoas`) is a valid question with a usually-empty answer, because most networks are not in conflict. That empty result means *no conflict*, not *no data*.

**From here, →** [Where is a network physically present?](#where-is-a-network-physically-present).

---

## Physical infrastructure distributions

The layer DNS-only datasets don't have: data centers, internet exchanges, submarine cables and root-server instances as queryable nodes, joined to the networks that sit in them. This is where you study the *physical* topology of the internet.

### Where is a network physically present?

`AS_PRESENT_AT` connects an ASN to the facilities it occupies; `IX_MEMBER` to the exchanges it joins. Both are one hop off the network.

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
// Cloudflare's physical footprint: facilities
MATCH (a:ASN {name: "AS13335"})-[:AS_PRESENT_AT]->(f:FACILITY)
RETURN f.name AS facility
ORDER BY facility
LIMIT 25
```
**Returns:** `facility`

```cypher expect=rows>0 seed=AS13335 verified=2026-09-02
// Which internet exchanges is the network a member of?
MATCH (a:ASN {name: "AS13335"})-[:IX_MEMBER]->(ix:INTERNET_EXCHANGE)
RETURN ix.name AS internet_exchange
ORDER BY internet_exchange
LIMIT 25
```

**Costs:** milliseconds; one anchored hop each.

**From here, →** [IXP membership density](#ixp-membership-density).

### IXP membership density

How many networks a given exchange aggregates is a measure of regional interconnection. Anchor on the exchange and count members.

```cypher expect=rows>0 seed="LINX LON1" verified=2026-09-02
// Member-network count at a major exchange
MATCH (a:ASN)-[:IX_MEMBER]->(ix:INTERNET_EXCHANGE {name: "LINX LON1"})
RETURN ix.name AS exchange, count(DISTINCT a) AS member_networks
LIMIT 1
```
**Returns:** `exchange, member_networks`

**Sample output:**
```json
[{"exchange": "LINX LON1", "member_networks": 835}]
```

**Costs:** milliseconds; one inbound hop from an indexed exchange name, aggregated.

**From here, →** [Submarine-cable landing topology](#submarine-cable-landing-topology).

### Submarine-cable landing topology

Subsea cables (`SUBMARINE_CABLE`) land at `CABLE_LANDING` points, which sit near `FACILITY` buildings. Trace a cable's landings to study coastal interconnection.

```cypher expect=rows>0 seed=2Africa verified=2026-09-02
// Where does the 2Africa cable land?
MATCH (cable:SUBMARINE_CABLE {name: "2Africa"})-[:CABLE_LANDS_AT]->(lp:CABLE_LANDING)
OPTIONAL MATCH (lp)-[:LANDING_NEAR]->(f:FACILITY)
RETURN lp.name AS landing_point, collect(DISTINCT f.name) AS nearby_facilities
ORDER BY landing_point
LIMIT 25
```
**Returns:** `landing_point, nearby_facilities`

**Sample output:**
```json
[{"landing_point": "Dakar, Senegal", "nearby_facilities": ["ONIX Senegal", "PAIX Dakar"]}]
```

**Costs:** milliseconds; one anchored hop plus an optional facility hop; keep `LANDING_NEAR` optional.

**From here, →** [Facility co-location — who else is in this building?](#facility-co-location-who-else-is-in-this-building).

### Facility co-location — who else is in this building?

The carrier-hotel adjacency that explains a lot of real-world peering. Anchor on a facility and count its resident networks.

```cypher expect=rows>0 verified=2026-09-02
// Networks present in a major carrier hotel
MATCH (a:ASN)-[:AS_PRESENT_AT]->(f:FACILITY {name: "Equinix DA1 - Dallas"})
RETURN count(DISTINCT a) AS resident_networks
LIMIT 1
```
**Returns:** `resident_networks`

**Sample output:**
```json
[{"resident_networks": 510}]
```

**Costs:** milliseconds; one inbound hop from an indexed facility name; `FIBER_SEGMENT` (facility → facility) and `CDN_POP_AT` (CDN point of presence → facility) extend the same building into its fiber links and CDN tenants.

**From here, →** [Prefixes mapped to cloud regions](#prefixes-mapped-to-cloud-regions).

### Prefixes mapped to cloud regions

`PREFIX_IN_REGION` ties address space to cloud-provider regions (`aws:eu-west-1` style names), the foundation for studying cloud address-space distribution. Anchor on the region or the prefix; the source label is `PREFIX`, not `ANNOUNCED_PREFIX`.

```cypher expect=rows>0 seed=aws:eu-west-1 verified=2026-09-02
// Prefixes the graph places in a specific cloud region
MATCH (p:PREFIX)-[:PREFIX_IN_REGION]->(r:CLOUD_REGION {name: "aws:eu-west-1"})
WITH p LIMIT 1000
RETURN count(p) AS sampled_prefixes
```
**Returns:** `sampled_prefixes`

**Sample output:**
```json
[{"sampled_prefixes": 131}]
```

**Costs:** milliseconds; one inbound hop from an indexed region name, bounded; `MATCH (r:CLOUD_REGION) RETURN r.name` lists the region names, it is a small reference label.

> **Empty result:** cloud-region mapping is partial. **A zero-row result here means the address space is not mapped to a region, not that it is not in a cloud.** Zero rows is never a verdict.

**From here, →** [Where are the DNS root-server instances?](#where-are-the-dns-root-server-instances).

### Where are the DNS root-server instances?

Studying the physical distribution of the root zone (how many anycast instances each root letter operates and where they sit) is a question about internet structure rather than about any one network. `DNS_ROOT_INSTANCE` is a small reference label, so group it directly.

```cypher expect=rows>0 verified=2026-09-02
// Root-server instances by country
MATCH (d:DNS_ROOT_INSTANCE)
RETURN d.countryCode AS country, count(*) AS instances
ORDER BY instances DESC
LIMIT 10
```
**Returns:** `country, instances`

**Sample output:**
```json
[
  {"country": "US", "instances": 271},
  {"country": "BR", "instances": 64},
  {"country": "CA", "instances": 46},
  {"country": "DE", "instances": 44}
]
```

Narrow to one country to see the individual instances, which letter each serves, and whether it answers globally or only locally:

```cypher expect=rows>0 seed=NL verified=2026-09-02
// The root instances in one country
MATCH (d:DNS_ROOT_INSTANCE)
WHERE d.countryCode = "NL"
RETURN d.name AS instance, d.rootLetter AS root_letter,
       d.town AS town, d.type AS instance_type
LIMIT 10
```

**Sample output:**
```json
[
  {"instance": "a3.nl-ams.root", "root_letter": "J", "town": "Amsterdam", "instance_type": "Global"},
  {"instance": "amnl1.droot.maxgigapop.net", "root_letter": "D", "town": "Amsterdam", "instance_type": "Global"}
]
```

**Costs:** milliseconds; a grouped scan of a small reference label; `rootLetter` is upper-case (`"K"`, not `"k"`), the usual reason a filter on it comes back empty.

> `type` separates `Global` instances, announced to the whole internet, from `Local` ones, whose announcement is deliberately constrained to a region; a country served only by `Local` instances has a different resilience story from one hosting a `Global` node. Instance naming is inconsistent across operators by nature, so treat `name` as an operator label rather than a resolvable hostname.

**From here, →** [Which TLDs does a registry operator run?](#which-tlds-does-a-registry-operator-run).

---

## Naming and registry ecosystem

### Which TLDs does a registry operator run?

Studying the registry ecosystem, you want which TLDs a given operator runs. `TLD_OPERATOR-[:OPERATES]->TLD` is the link, and the operator is the fast direction to anchor on.

```cypher expect=rows>0,no-null-columns seed="NISSAN MOTOR CO., LTD." verified=2026-09-02
// TLDs operated by a registry, anchored on the operator
MATCH (op:TLD_OPERATOR {name: "NISSAN MOTOR CO., LTD."})-[:OPERATES]->(t:TLD)
RETURN op.name AS operator, collect(DISTINCT t.name) AS tlds
LIMIT 5
```
**Returns:** `operator, tlds`

**Sample output:**
```json
[{"operator": "NISSAN MOTOR CO., LTD.", "tlds": ["datsun", "infiniti", "nissan"]}]
```

**Costs:** milliseconds; one anchored hop; operator names are exact strings, punctuation included.

> A brand running several vanity TLDs (as here) is a neat illustration of the post-2012 gTLD landscape. `MATCH (t:TLD) RETURN count(t)` sizes the TLD population; it is a small label and safe to scan.

**From here, →** [Which apexes are shared hosting in disguise?](#which-apexes-are-shared-hosting-in-disguise).

### Which apexes are shared hosting in disguise?

Co-hosting results keep surfacing apexes with thousands of unrelated subdomains: CDN and multi-tenant platforms that make every tenant look like a neighbour. Identify them so you can weight them down before you report co-tenancy as a relationship.

```cypher expect=rows>0 verified=2026-09-02
// Apexes whose certificate and subdomain fan-out looks like shared hosting
CALL whisper.threatIntel.candidateCdnApex(5)
YIELD apex, subCount, certCount, wildcardCount, recommendation
RETURN apex, subCount, certCount, wildcardCount, recommendation
LIMIT 5
```
**Returns:** `apex, subCount, certCount, wildcardCount, recommendation`

**Sample output:**
```json
[
  {"apex": "microsoft.com", "subCount": 13021, "certCount": 54832, "wildcardCount": 9987, "recommendation": "add-to-deny-list"},
  {"apex": "narkive.com", "subCount": 7327, "certCount": 10614, "wildcardCount": 0, "recommendation": "add-to-deny-list-no-wildcard"}
]
```

**Costs:** milliseconds; a precomputed snapshot read; the argument is a number of candidates, not a domain name.

> A high `certCount` relative to `subCount` means many independent certificates under one apex, the signature of a platform serving unrelated tenants. `recommendation` is the graph's own read on whether the apex is worth denying outright or just watching. Use the output as a suppression list in any co-tenancy study.

**From here, →** [Tor-exit egress distribution by network](#tor-exit-egress-distribution-by-network).

---

## Cross-layer studies

The payoff of a pre-joined graph is the join across layers that normally live in separate tools. A few research-shaped combinations.

### Tor-exit egress distribution by network

`OPERATES_EXIT_NODE` links an IP to its `TOR_RELAY` identity (which survives IP rotation). Join Tor exits to the networks that route them to see where exit capacity concentrates. Sign in to run it.

```cypher expect=rows>0 verified=2026-09-02
// Which networks host a sample of Tor exit IPs?
MATCH (ip:IPV4)-[:OPERATES_EXIT_NODE]->(relay:TOR_RELAY)
WITH ip LIMIT 1000
MATCH (ip)-[:ANNOUNCED_BY]->(ap:ANNOUNCED_PREFIX)-[:ROUTES]->(a:ASN)
RETURN a.name AS asn, count(DISTINCT ip) AS exit_ips
ORDER BY exit_ips DESC
LIMIT 15
```
**Returns:** `asn, exit_ips`

**Sample output:**
```json
[
  {"asn": "AS62744", "exit_ips": 100},
  {"asn": "AS60729", "exit_ips": 60},
  {"asn": "AS53667", "exit_ips": 21}
]
```

**Costs:** under a second; a bounded sample of the exit population, then two hops into routing; the `WITH ip LIMIT 1000` is what keeps it a sample.

**From here, →** [Threat density across a network's prefixes](#threat-density-across-a-network-s-prefixes).

> **Read `coverage` before `band`.** Only `known-clean` licenses the word "clean"; `no-data` means
> *unknown*, which is a different thing again; `malicious-evidenced` and `ambiguous` mean there is
> evidence, whatever the band says.
> Full contract: [Coverage — what we looked at](/docs/whisper-graph/procedures/coverage).

### Threat density across a network's prefixes

Aggregate threat is rolled up onto `ANNOUNCED_PREFIX` (`threatScore`, `threatLevel`) and onto the `ASN` itself (`maxThreatScore`, `avgThreatScore`, `hasThreateningPrefixes`), so you can study reputation distribution without walking every IP. For per-network triage, read the ASN aggregate directly; for the contributing prefixes, anchor and bound.

```cypher expect=static seed=AS13335 verified=2026-09-02 reason="multi-origin conflict state changes as prefixes are withdrawn and re-announced, so this block shows a captured result rather than a live run"
// Threat-listed prefixes within a network, ordered by aggregate score
MATCH (a:ASN {name: "AS13335"})-[:ROUTES]->(ap:ANNOUNCED_PREFIX)
WHERE ap.threatScore > 0
RETURN ap.name AS prefix, ap.threatScore AS score, ap.threatLevel AS level
ORDER BY score DESC
LIMIT 25
```
**Returns:** `prefix, score, level`

**Sample output:**
```json
[{"prefix": "172.70.207.0/24", "score": 40, "level": "HIGH"}, {"prefix": "172.70.206.0/24", "score": 40, "level": "HIGH"}]
```

**Costs:** milliseconds; one anchored hop with a property filter.

> For a comparable per-address measure across networks, `CALL whisper.asnThreatDensity("AS13335")` returns `listedIps`, `announcedIpv4` and `densityRatio`. For a clean per-indicator verdict with its reasoning, prefer `CALL explain("AS13335")` over hand-walking `ASN → PREFIX → IP → LISTED_IN`; the manual walk does not finish on a large network. See [explain()](/docs/whisper-graph/procedures/explain).

**From here, →** [TLS-fingerprint reuse across IPs](#tls-fingerprint-reuse-across-ips).

### TLS-fingerprint reuse across IPs

`EMITS_TLS_FINGERPRINT` ties an IP to a JA3/JARM fingerprint, useful for studying how a server signature spreads across address space. Seed it with an IP that actually carries a fingerprint: this is a thin plane, so most addresses have none.

```cypher expect=rows>0,no-null-columns seed=18.189.12.168 verified=2026-09-02
// IPs sharing a given JARM/JA3 fingerprint
MATCH (ip:IPV4 {name: "18.189.12.168"})-[:EMITS_TLS_FINGERPRINT]->(fp:TLS_FINGERPRINT)
WITH fp LIMIT 1
MATCH (other:IPV4)-[:EMITS_TLS_FINGERPRINT]->(fp)
RETURN fp.name AS fingerprint, count(DISTINCT other) AS ips_with_fingerprint
LIMIT 1
```
**Returns:** `fingerprint, ips_with_fingerprint`

**Sample output:**
```json
[{"fingerprint": "jarm:07d14d16d21d21d07c42d41d00041d24a458a375eef0c576d23a7bab9a9fb1",
  "ips_with_fingerprint": 141}]
```

**Costs:** milliseconds; one anchored hop out and one back; pick the seed from the graph rather than from an incident.

> **Empty result:** the layer holds 271 `EMITS_TLS_FINGERPRINT` edges, so expect no match on almost any indicator. **A zero-row result here means Whisper holds no observation, not that the host shares no infrastructure.** `MATCH (ip:IPV4)-[:EMITS_TLS_FINGERPRINT]->(fp) RETURN ip.name, fp.name LIMIT 5` gives you a seed that works. Zero rows is never a verdict.

> Anchor `TLS_FINGERPRINT` on its actual `.name` (the `jarm:`- or `ja3:`-prefixed hash). `CALL whisper.lookupTlsFingerprint("jarm:…")` classifies a hash you already hold.

**From here, →** [Actor → ATT&CK technique map](#actor-att-ck-technique-map).

### Actor → ATT&CK technique map

Named threat actors (`ACTOR`, case-sensitive) link to the MITRE techniques they use via `USES_TECHNIQUE`. Study an actor's technique footprint without leaving the graph.

```cypher expect=rows>0 seed=APT28 verified=2026-09-02
// MITRE ATT&CK techniques mapped to APT28 in public reporting
MATCH (actor:ACTOR {name: "APT28"})-[:USES_TECHNIQUE]->(t:ATTACK_PATTERN)
RETURN DISTINCT t.name AS technique
ORDER BY technique
LIMIT 50
```
**Returns:** `technique`

**Sample output:**
```json
[{"technique": "Additional Email Delegate Permissions"}, {"technique": "Application Access Token"}, {"technique": "Archive Collected Data"}]
```

**Costs:** milliseconds; one anchored hop; techniques are `ATTACK_PATTERN`, never `TECHNIQUE`, and `ACTOR.aliases` holds the vendor names.

> WhisperGraph carries the MITRE ATT&CK knowledge base as graph structure: 9,256 `USES_TECHNIQUE` edges from `ACTOR` to `ATTACK_PATTERN` and 872 `USES_TACTIC` edges, across 1,944 actors and 712 attack patterns. **This is a curated reference layer, not Whisper's own attribution.** It reflects what public reporting has mapped, not what Whisper observed. `ATTRIBUTED_TO`, the edge from an indicator to an actor, holds 302 edges; these queries return technique and tactic rollups. **They do not attribute anything.**
>
> Convergence on a shared technique is a lead about the reporting, not about the infrastructure. For the infrastructure-side pivots (co-tenancy, shared registrant, nameserver siblings) see [Campaign Pivoting](/docs/recipes/threat-intel).

**From here, →** [What's actually in the graph](#what-s-actually-in-the-graph) to check the next label before you build on it.

---

## Programmatic bulk runs

For aggregate studies you'll script the endpoint rather than click. Cypher over REST at `https://graph.whisper.security/api/query`:

```bash
curl -s https://graph.whisper.security/api/query \
  -H "Content-Type: application/json" \
  -H "X-API-Key: $WHISPER_API_KEY" \
  -d '{"query":"UNWIND [\"AS13335\",\"AS3356\",\"AS15169\",\"AS2914\"] AS asn MATCH (a:ASN {name:asn})-[:BGP_NEIGHBOR]->(p:ASN) RETURN asn, count(p) AS degree ORDER BY degree DESC"}'
```

Each call returns `columns`, `rows`, and execution `statistics` as JSON, so you can fan a sample of anchors out across calls and reassemble the distribution locally. AI agents get the same surface via MCP at `https://mcp.whisper.security`; point any MCP client at it and it runs these queries mid-analysis (see [AI & Agents](/docs/ai)).

For more patterns by workflow, see [Workflows](/docs/workflows); for the complete label/edge/property model, the [Graph Schema](/docs/whisper-graph/schema).
