Skip to content
Recipes
Skip navigation
Recipes
View as Markdown

Internet Measurement

Bulk, aggregate Cypher for internet measurement: schema/feed catalogs, peering topology, RPKI coverage, root-server and cross-layer studies.

Published Last updated

On this page (33)

Internet Measurement Documentation

You study the internet itself: topology, deployment trends, ecosystem structure, not one incident at a time but in aggregate, across billions of edges. The hard part isn't the analysis; it's getting clean, joined, planet-scale data to analyze. These recipes take you to the joined data: DNS, BGP peering and observed paths, RPKI, the root-server and registry ecosystems, and the physical internet, all in one Cypher surface. The recipes below are bulk and aggregate queries, written with the bounding that keeps them fast on a graph of 7.5B nodes and 39.8B edges.

A few rules that keep research queries honest at this scale:

  • Anchor or aggregate, never bare-scan a large label. A query that touches all of HOSTNAME, IPV4 or NAMESERVER_FOR without an anchored start will not finish. Anchor on a {name:"..."} node, or aggregate behind a CALL db.* histogram or a whisper.* ranking procedure.
  • Bound high-fan-out hops with WITH ... LIMIT before you expand again.
  • Small reference labels are safe to scan. FEED_SOURCE, CATEGORY, VENDOR, CDN_POP, DNS_ROOT_INSTANCE and THREAT_SIGNAL_TYPE are catalogues, and listing them is instant. Reach the large labels through an edge.
  • Treat LINKS_TO as a sampled crawl layer, not a web-scale link graph. Anchor on a host, read its degree as a floor, and do not build a link-structure study on it alone.

See Getting Started for keys, the Graph Schema for the full label/edge model, and the Procedures reference for the CALL surface.

Run it live. Several of these measurement jobs have a guided, browser-runnable version that opens with a result on your own indicator:

  • Digital Infrastructure Mapping — one indicator mapped to its owner and full footprint across every layer.
  • Supply-Chain Dependency Mapping — every external provider a domain depends on, grouped by function, with dependency chains and single-vendor (SPOF) signals across facilities, cable landings and subsea cables.
  • Investigate an Indicator — verdict, hosting, routing, and the shared infrastructure around any domain, IP, ASN, or prefix.

More live flows are on the Research & OSINT use cases page.

Key concepts: BGP routing · RPKI ROA · MOAS conflict · Internet exchange point · Submarine cable · Tor exit node · MITRE ATT&CK.


Schema exploration

What's actually in the graph

Before you write a traversal, confirm the label and edge exist. The most common cause of an empty result set is anchoring on a label that doesn't exist (there is no DOMAIN or FQDN label; every name is a HOSTNAME). The db.* procedures return precomputed histograms, instant even at this scale.

cypher · runnablegraph.whisper.securitySign in to run
// Every node label with its live count
CALL db.labels()

Returns: label, count

Sample output (the five largest of 42 labels):

json
[
  {"label": "HOSTNAME", "count": 2752403048},
  {"label": "IPV4", "count": 621441120},
  {"label": "EMAIL", "count": 237065663},
  {"label": "ORGANIZATION", "count": 119189847},
  {"label": "PHONE", "count": 60194142}
]
cypher · runnablegraph.whisper.securitySign in to run
// Every edge type with its live count, source and target labels
CALL db.relationshipTypes()

Sample output (the five largest of 55 edge types, sourceLabels/targetLabels elided):

json
[
  {"type": "NAMESERVER_FOR", "count": 9173662411},
  {"type": "ANNOUNCED_BY", "count": 4331089630},
  {"type": "RESOLVES_TO", "count": 3125689316},
  {"type": "CHILD_OF", "count": 2451196569},
  {"type": "REGISTERED_BY", "count": 916255242}
]
cypher · runnablegraph.whisper.securitySign in to run
// Every property name in the graph
CALL db.propertyKeys() YIELD propertyKey RETURN propertyKey ORDER BY propertyKey LIMIT 200

Costs: milliseconds; histogram reads, no traversal; the column for edges is type, not relationshipType.

Why this matters: the edge histogram is a research dataset, the shape of the global internet one CALL away. Use these counts to plan which traversals are cheap (anchored) and which need aggregation, and read sourceLabels/targetLabels off db.relationshipTypes() to learn an edge's direction before you write it.

From here, → Confirm a property before you filter on it.

Confirm a property before you filter on it

Filtering on a property that doesn't exist returns empty, not an error. A quick keys() read (or db.propertyKeys()) saves you from WHERE h.fqdn = ... when the property is name.

cypher · runnablegraph.whisper.securitySign in to run
// What properties does a threat-listed IP carry?
MATCH (ip:IPV4 {name: "185.220.101.1"})
RETURN keys(ip) AS properties
LIMIT 1

Returns: properties

Sample output (trimmed):

json
[{"properties": ["id", "label", "name", "threatScore", "threatSources", "isThreat", "isTor", "threatLevel", "verdictLevel", "verdictScore", "verdictCoverage", "verdictBlocking"]}]

Costs: milliseconds; one indexed anchor and a property read; the key set differs by label, so check the label you are about to filter.

From here, → The threat feed catalog.

The threat feed catalog

Feed coverage is a study in itself. The catalogue is small enough to list directly, and any indicator's LISTED_IN edges tell you which feeds and categories cover it.

cypher · runnablegraph.whisper.securitySign in to run
// All threat-feed sources, by display name
MATCH (f:FEED_SOURCE)
RETURN f.displayName AS feed
ORDER BY f.displayName
LIMIT 15

Returns: feed

Sample output:

json
[{"feed": "1Hosts Xtra"}, {"feed": "AlienVault Reputation"}, {"feed": "Bad Hosting ASN"}]
cypher · runnablegraph.whisper.securitySign in to run
// Which feeds and categories cover a known Tor exit?
MATCH (ip:IPV4 {name: "185.220.101.1"})-[:LISTED_IN]->(f:FEED_SOURCE)
MATCH (f)-[:BELONGS_TO]->(c:CATEGORY)
RETURN c.displayName AS category, collect(f.displayName) AS feeds
ORDER BY category
LIMIT 25

Sample output:

json
[
  {"category": "General Blacklists", "feeds": ["GreenSnow Blacklist", "IPsum", "FireHOL Level 2", "duggytuxy-datashield-critical"]},
  {"category": "Spam", "feeds": ["StopForumSpam Listed IPs (7 day)"]},
  {"category": "TOR Network", "feeds": ["Tor Exit Nodes"]}
]

Costs: milliseconds; a scan of a small reference label, and two anchored hops for the per-indicator view; swap FEED_SOURCE for CATEGORY to list the categories.

The graph indexes 134 feeds across 32 categories with 14.5M LISTED_IN edges. The catalog spans block lists and trust lists, so the same query surface answers "known-bad?" and "known-good?". .name is the slug (firehol-level2, tor) and .displayName the readable label; the full list is in Threat Feeds & Categories.

From here, → BGP peering-degree, network by network.


Internet topology

BGP peering-degree, network by network

Peering data lives in PeeringDB and route-collector dumps you have to download, parse and join yourself. In the graph, BGP_NEIGHBOR is the canonical ASN↔ASN adjacency edge, already materialized. Anchor on a set of ASNs and count peers in one round-trip.

cypher · runnablegraph.whisper.securitySign in to run
// Peering degree for a sample of well-known networks
UNWIND ["AS13335", "AS3356", "AS15169", "AS2914"] AS asn_name
MATCH (a:ASN {name: asn_name})-[:BGP_NEIGHBOR]->(peer:ASN)
RETURN asn_name, count(peer) AS peer_count
ORDER BY peer_count DESC

Returns: asn_name, peer_count

Sample output:

json
[
  {"asn_name": "AS3356", "peer_count": 6196},
  {"asn_name": "AS2914", "peer_count": 1461},
  {"asn_name": "AS13335", "peer_count": 1284},
  {"asn_name": "AS15169", "peer_count": 139}
]

Costs: milliseconds; one anchored hop per element, aggregated; PEERS_WITH still resolves as an alias, but write BGP_NEIGHBOR.

Reading it: transit-heavy carriers (AS3356 Lumen, AS2914 NTT) sit at the top of the degree distribution; content networks (AS15169 Google) peer selectively. The degree gap is the structural difference between transit and content ASNs, visible in one query. For the whole distribution rather than a sample, CALL whisper.bgpDegreeDistribution() returns one row per (inDegree, outDegree) pair; see BGP & RPKI.

From here, → Rank the densest networks by prefix count.

Rank the densest networks by prefix count

For a top-of-distribution view without enumerating every ASN, use the ranking procedure. It reads a precomputed ranking, so it's instant where the equivalent per-ASN sweep would be expensive.

cypher · runnablegraph.whisper.securitySign in to run
// The ASNs announcing the most prefixes
CALL whisper.topAsnsByPrefixCount(15)
YIELD asn, prefixCount
RETURN asn, prefixCount
LIMIT 15

Returns: asn, prefixCount

Sample output:

json
[
  {"asn": "AS16509", "prefixCount": 22511},
  {"asn": "AS9808", "prefixCount": 21519},
  {"asn": "AS577", "prefixCount": 16258}
]

To study a single network's footprint, anchor and bound the fan-out:

cypher · runnablegraph.whisper.securitySign in to run
// How many prefixes does Cloudflare announce?
MATCH (a:ASN {name: "AS13335"})-[:ROUTES]->(ap:ANNOUNCED_PREFIX)
RETURN count(ap) AS announced_prefixes
LIMIT 1

Costs: milliseconds; a precomputed ranking, then one anchored aggregated hop; the argument is a number of networks, not an ASN.

The ordering shifts as the routing table does, so cite it with a date.

From here, → Second-degree peering reach.

Second-degree peering reach

The peering graph's structure shows up in the two-hop neighborhood: how many distinct networks are within two BGP hops. Bound the first hop hard before expanding; a large carrier's neighbor set is in the thousands.

cypher · runnablegraph.whisper.securitySign in to run
// Distinct networks reachable within two BGP hops of Cloudflare
MATCH (a:ASN {name: "AS13335"})-[:BGP_NEIGHBOR]->(n1:ASN)
WITH DISTINCT n1 LIMIT 500
MATCH (n1)-[:BGP_NEIGHBOR]->(n2:ASN)
RETURN count(DISTINCT n2) AS two_hop_reach
LIMIT 1

Returns: two_hop_reach

Sample output:

json
[{"two_hop_reach": 33555}]

Costs: milliseconds; two explicit hops with a WITH ... LIMIT 500 between them, which is load-bearing.

Tip. BGP_NEIGHBOR also works inside a bounded variable-length pattern: MATCH p = (a:ASN {name: "AS13335"})-[:BGP_NEIGHBOR*2..2]->(n:ASN) WHERE n <> a RETURN n.name, length(p) LIMIT 10 samples the second ring directly. Keep WHERE n <> a: a peering mesh is undirected in practice, so a two-hop walk routinely lands back on the origin.

From here, → ASN home jurisdiction distribution.

ASN home jurisdiction distribution

HAS_COUNTRY runs from ASN straight to COUNTRY, no city hop needed for the AS's registered jurisdiction. Aggregate across a sample to study where networks are domiciled.

cypher · runnablegraph.whisper.securitySign in to run
// Home country for a sample of networks
UNWIND ["AS13335", "AS15169", "AS3356", "AS2914", "AS4837"] AS asn_name
MATCH (a:ASN {name: asn_name})-[:HAS_COUNTRY]->(c:COUNTRY)
RETURN c.name AS country, count(*) AS networks
ORDER BY networks DESC

Returns: country, networks

Sample output:

json
[{"country": "US", "networks": 4}, {"country": "CN", "networks": 1}]

Costs: milliseconds; one anchored hop per element, aggregated.

From here, → Which countries hold the most autonomous systems?.

Which countries hold the most autonomous systems?

Characterising the shape of the routed internet, you want the distribution of autonomous systems by registered country across the whole population, not a sample. The ranking procedure reads it precomputed.

cypher · runnablegraph.whisper.securitySign in to run
// Which countries hold the most autonomous systems
CALL whisper.asnCountries(10) YIELD country, asns
RETURN country, asns
ORDER BY asns DESC
LIMIT 10

Returns: country, asns

Sample output:

json
[
  {"country": "US", "asns": 31320},
  {"country": "BR", "asns": 8933},
  {"country": "IN", "asns": 6113},
  {"country": "RU", "asns": 5620}
]

Costs: milliseconds; a precomputed ranking; the argument is the number of countries to return, not a country code.

These are registered countries from the routing registries, a legal-entity fact rather than a physical one: a network registered in one country routinely announces prefixes that terminate somewhere else. For where the traffic actually lands, join through the facility and exchange layer below.

From here, → Outbound link degree from a domain.


LINKS_TO is a crawl-derived hyperlink edge between hostnames, in the same query surface as DNS, WHOIS and BGP, which means you can join link structure against routing without exporting a CSV. It is a sample of the crawled web, not a census of it. The one rule: always anchor, and read every degree as a floor.

How many distinct hosts a site links out to is the simplest link-structure measurement, and anchored on the host it is an instant read.

cypher · runnablegraph.whisper.securitySign in to run
// How many distinct hosts does github.com link out to?
MATCH (h:HOSTNAME {name: "github.com"})-[:LINKS_TO]->(target:HOSTNAME)
RETURN count(DISTINCT target) AS outbound_hosts
LIMIT 1

Returns: outbound_hosts

Sample output:

json
[{"outbound_hosts": 189369}]

Costs: milliseconds; one anchored hop, aggregated; degree reflects what the crawl sample captured for that host.

From here, → Join the link graph to routing — what networks does a site link out to?.

The join this makes: follow each outbound link to where it actually resolves and who routes it. Cap the link fan-out first, then traverse DNS→BGP for each.

cypher · runnablegraph.whisper.securitySign in to run
// Outbound links → resolve each target → which networks host them
MATCH (h:HOSTNAME {name: "github.com"})-[:LINKS_TO]->(target:HOSTNAME)
WITH DISTINCT target LIMIT 200
MATCH (target)-[:RESOLVES_TO]->(ip:IPV4)-[:ANNOUNCED_BY]->(ap:ANNOUNCED_PREFIX)-[:ROUTES]->(a:ASN)
RETURN a.name AS asn, count(DISTINCT target) AS linked_hosts
ORDER BY linked_hosts DESC
LIMIT 15

Returns: asn, linked_hosts

Sample output:

json
[
  {"asn": "AS13335", "linked_hosts": 43},
  {"asn": "AS16509", "linked_hosts": 38},
  {"asn": "AS396982", "linked_hosts": 8}
]

Costs: milliseconds; a four-layer join (link → DNS → BGP announcement → ASN) kept bounded by WITH DISTINCT target LIMIT 200; sign in to run the three-hop leg.

Why it's hard otherwise: this crosses three datasets that normally live in three different tools. Here it's one statement. Widen the sample deliberately, and re-anchor on a linked host to walk further rather than writing a variable-length pattern over LINKS_TO.

From here, → Is a network's announced space ROA-covered?.


RPKI coverage

RPKI Route Origin Authorizations (ROA) are first-class nodes (3M of them), linked to the prefixes and origin ASNs they authorize. You can study deployment and validity without pulling and parsing the RIR trust-anchor dumps yourself.

Is a network's announced space ROA-covered?

Cross-referencing announced prefixes against the RPKI repository means reconciling two separate feeds. In the graph both are nodes; ROA_AUTHORIZES_ORIGIN joins them, and the prefix each ROA covers is a property on the ROA itself.

cypher
// RPKI ROAs that authorize Cloudflare as an origin AS, with the max-length they permit
MATCH (roa:ROA)-[:ROA_AUTHORIZES_ORIGIN]->(a:ASN {name: "AS13335"})
WHERE roa.authSource = "rpki-roa"
RETURN roa.prefix AS authorized_prefix, roa.maxLength AS max_length,
       roa.trustAnchor AS trust_anchor
LIMIT 25

Returns: authorized_prefix, max_length, trust_anchor

Sample output (captured 2026-10-06):

json
[
  {"authorized_prefix": "102.219.82.0/24", "max_length": 24, "trust_anchor": "afrinic"},
  {"authorized_prefix": "154.193.133.0/24", "max_length": 24, "trust_anchor": "afrinic"},
  {"authorized_prefix": "154.193.184.0/24", "max_length": 24, "trust_anchor": "afrinic"}
]

Costs: milliseconds; one inbound hop from an indexed ASN plus property reads; count(roa) sizes the set before you list it.

Empty result: no rows means no ROA names this ASN as an origin, which is the unsigned state, not missing data. Zero rows is never a verdict.

A ROA node has no name; its identity is the (prefix, asn) pair it authorizes. Two kinds of record share the label, told apart by authSource. An RPKI ROA (rpki-roa) carries id, label, authSource, asn, prefix, maxLength, trustAnchor, validUntil. A route object from an internet routing registry (irr-route) carries id, label, authSource, asn, prefix, mnt_by, source_rir and no maxLength, trustAnchor or validUntil, and it is the larger share of the label. keys(roa) shows which kind a sample is, and WHERE roa.authSource = "rpki-roa" keeps a study to RPKI.

From here, → Cross-check an announcement against its authorization.

Cross-check an announcement against its authorization

Walk from an IP to its announced prefix, then read the precomputed validation state and count the ROAs covering that exact prefix: the building block of route-origin validation, written as explicit single hops.

cypher · runnablegraph.whisper.securitySign in to run
// Does the prefix covering 1.1.1.1 validate, and which RPKI ROAs cover it?
MATCH (ip:IPV4 {name: "1.1.1.1"})-[:ANNOUNCED_BY]->(ap:ANNOUNCED_PREFIX)
OPTIONAL MATCH (roa:ROA)-[:ROA_AUTHORIZES_PREFIX]->(ap)
WHERE roa.authSource = "rpki-roa"
RETURN ap.name AS announced_prefix, ap.rpkiStatus AS rpki_status,
       ap.roaAsn AS roa_asn, ap.roaMaxLength AS roa_max_length,
       count(roa) AS covering_roas, collect(DISTINCT roa.asn) AS roa_origins
LIMIT 5

Returns: announced_prefix, rpki_status, roa_asn, roa_max_length, covering_roas, roa_origins

Sample output:

json
[{"announced_prefix": "1.1.1.0/24", "rpki_status": "valid", "roa_asn": 13335, "roa_max_length": 24, "covering_roas": 1, "roa_origins": [13335]}]

Costs: milliseconds; one anchored hop plus one optional ROA hop; rpkiStatus takes valid, invalid or not-found, and not-found is the unsigned population.

Empty result: covering_roas: 0 with rpki_status: "not-found" is unsigned space, not a missing edge. Zero rows is never a verdict.

From here, → MOAS conflicts — the early hijack signal.

MOAS conflicts — the early hijack signal

A prefix announced by more than one origin AS is the leading early signal of a BGP hijack or route leak. For a study you want the population, not one network, so anchor on the conflict edge itself, bound it, and rank by how many origins are competing.

cypher
// The most heavily contested prefixes, and how many origins announce them
MATCH (p:ANNOUNCED_PREFIX)-[:CONFLICTS_WITH]->(other:ASN)
WITH DISTINCT p LIMIT 3000
MATCH (p)-[:CONFLICTS_WITH]->(o:ASN)
WITH p, collect(DISTINCT o.name) AS conflicting_origins
WHERE size(conflicting_origins) > 1
RETURN p.name AS prefix, size(conflicting_origins) AS origins,
       conflicting_origins[0..6] AS sample_origins
ORDER BY origins DESC
LIMIT 25

Returns: prefix, origins, sample_origins

Sample output:

json
[
  {"prefix": "192.58.128.0/24", "origins": 21, "sample_origins": ["AS396549", "AS396738", "AS396739", "AS396707", "AS396576", "AS396686"]},
  {"prefix": "192.30.45.0/24", "origins": 12, "sample_origins": ["AS396549", "AS396578", "AS20362", "AS211369", "AS396555", "AS396566"]}
]

Costs: milliseconds with the WITH DISTINCT p LIMIT 3000 bound; the same aggregation without the bound runs for many seconds.

Read the tail before you read the head. The graph holds 9,721 CONFLICTS_WITH edges, and the ranking is dominated by prefixes that are supposed to have many origins: 192.58.128.0/24 (J-root) is anycast working correctly, not a stack of hijacks. A MOAS study's real work is separating anycast and legitimate multi-homing from the two- or three-origin cases that are anomalies; p.moasIsLegitimate is the graph's own read on that. Pair it with the per-announcement RPKI state (rpkiStatus / roaAsn / roaMaxLength): a MOAS conflict where one origin is RPKI-invalid is a far stronger signal than the conflict alone. See BGP & RPKI.

Starting from a named network instead (MATCH (a:ASN {name: "…"})-[:ROUTES]->(p) WHERE p.isMoas) is a valid question with a usually-empty answer, because most networks are not in conflict. That empty result means no conflict, not no data.

From here, → Where is a network physically present?.


Physical infrastructure distributions

The layer DNS-only datasets don't have: data centers, internet exchanges, submarine cables and root-server instances as queryable nodes, joined to the networks that sit in them. This is where you study the physical topology of the internet.

Where is a network physically present?

AS_PRESENT_AT connects an ASN to the facilities it occupies; IX_MEMBER to the exchanges it joins. Both are one hop off the network.

cypher · runnablegraph.whisper.securitySign in to run
// Cloudflare's physical footprint: facilities
MATCH (a:ASN {name: "AS13335"})-[:AS_PRESENT_AT]->(f:FACILITY)
RETURN f.name AS facility
ORDER BY facility
LIMIT 25

Returns: facility

cypher · runnablegraph.whisper.securitySign in to run
// Which internet exchanges is the network a member of?
MATCH (a:ASN {name: "AS13335"})-[:IX_MEMBER]->(ix:INTERNET_EXCHANGE)
RETURN ix.name AS internet_exchange
ORDER BY internet_exchange
LIMIT 25

Costs: milliseconds; one anchored hop each.

From here, → IXP membership density.

IXP membership density

How many networks a given exchange aggregates is a measure of regional interconnection. Anchor on the exchange and count members.

cypher · runnablegraph.whisper.securitySign in to run
// Member-network count at a major exchange
MATCH (a:ASN)-[:IX_MEMBER]->(ix:INTERNET_EXCHANGE {name: "LINX LON1"})
RETURN ix.name AS exchange, count(DISTINCT a) AS member_networks
LIMIT 1

Returns: exchange, member_networks

Sample output:

json
[{"exchange": "LINX LON1", "member_networks": 835}]

Costs: milliseconds; one inbound hop from an indexed exchange name, aggregated.

From here, → Submarine-cable landing topology.

Submarine-cable landing topology

Subsea cables (SUBMARINE_CABLE) land at CABLE_LANDING points, which sit near FACILITY buildings. Trace a cable's landings to study coastal interconnection.

cypher · runnablegraph.whisper.securitySign in to run
// Where does the 2Africa cable land?
MATCH (cable:SUBMARINE_CABLE {name: "2Africa"})-[:CABLE_LANDS_AT]->(lp:CABLE_LANDING)
OPTIONAL MATCH (lp)-[:LANDING_NEAR]->(f:FACILITY)
RETURN lp.name AS landing_point, collect(DISTINCT f.name) AS nearby_facilities
ORDER BY landing_point
LIMIT 25

Returns: landing_point, nearby_facilities

Sample output:

json
[{"landing_point": "Dakar, Senegal", "nearby_facilities": ["ONIX Senegal", "PAIX Dakar"]}]

Costs: milliseconds; one anchored hop plus an optional facility hop; keep LANDING_NEAR optional.

From here, → Facility co-location — who else is in this building?.

Facility co-location — who else is in this building?

The carrier-hotel adjacency that explains a lot of real-world peering. Anchor on a facility and count its resident networks.

cypher · runnablegraph.whisper.securitySign in to run
// Networks present in a major carrier hotel
MATCH (a:ASN)-[:AS_PRESENT_AT]->(f:FACILITY {name: "Equinix DA1 - Dallas"})
RETURN count(DISTINCT a) AS resident_networks
LIMIT 1

Returns: resident_networks

Sample output:

json
[{"resident_networks": 510}]

Costs: milliseconds; one inbound hop from an indexed facility name; FIBER_SEGMENT (facility → facility) and CDN_POP_AT (CDN point of presence → facility) extend the same building into its fiber links and CDN tenants.

From here, → Prefixes mapped to cloud regions.

Prefixes mapped to cloud regions

PREFIX_IN_REGION ties address space to cloud-provider regions (aws:eu-west-1 style names), the foundation for studying cloud address-space distribution. Anchor on the region or the prefix; the source label is PREFIX, not ANNOUNCED_PREFIX.

cypher · runnablegraph.whisper.securitySign in to run
// Prefixes the graph places in a specific cloud region
MATCH (p:PREFIX)-[:PREFIX_IN_REGION]->(r:CLOUD_REGION {name: "aws:eu-west-1"})
WITH p LIMIT 1000
RETURN count(p) AS sampled_prefixes

Returns: sampled_prefixes

Sample output:

json
[{"sampled_prefixes": 131}]

Costs: milliseconds; one inbound hop from an indexed region name, bounded; MATCH (r:CLOUD_REGION) RETURN r.name lists the region names, it is a small reference label.

Empty result: cloud-region mapping is partial. A zero-row result here means the address space is not mapped to a region, not that it is not in a cloud. Zero rows is never a verdict.

From here, → Where are the DNS root-server instances?.

Where are the DNS root-server instances?

Studying the physical distribution of the root zone (how many anycast instances each root letter operates and where they sit) is a question about internet structure rather than about any one network. DNS_ROOT_INSTANCE is a small reference label, so group it directly.

cypher · runnablegraph.whisper.securitySign in to run
// Root-server instances by country
MATCH (d:DNS_ROOT_INSTANCE)
RETURN d.countryCode AS country, count(*) AS instances
ORDER BY instances DESC
LIMIT 10

Returns: country, instances

Sample output:

json
[
  {"country": "US", "instances": 271},
  {"country": "BR", "instances": 64},
  {"country": "CA", "instances": 46},
  {"country": "DE", "instances": 44}
]

Narrow to one country to see the individual instances, which letter each serves, and whether it answers globally or only locally:

cypher · runnablegraph.whisper.securitySign in to run
// The root instances in one country
MATCH (d:DNS_ROOT_INSTANCE)
WHERE d.countryCode = "NL"
RETURN d.name AS instance, d.rootLetter AS root_letter,
       d.town AS town, d.type AS instance_type
LIMIT 10

Sample output:

json
[
  {"instance": "a3.nl-ams.root", "root_letter": "J", "town": "Amsterdam", "instance_type": "Global"},
  {"instance": "amnl1.droot.maxgigapop.net", "root_letter": "D", "town": "Amsterdam", "instance_type": "Global"}
]

Costs: milliseconds; a grouped scan of a small reference label; rootLetter is upper-case ("K", not "k"), the usual reason a filter on it comes back empty.

type separates Global instances, announced to the whole internet, from Local ones, whose announcement is deliberately constrained to a region; a country served only by Local instances has a different resilience story from one hosting a Global node. Instance naming is inconsistent across operators by nature, so treat name as an operator label rather than a resolvable hostname.

From here, → Which TLDs does a registry operator run?.


Naming and registry ecosystem

Which TLDs does a registry operator run?

Studying the registry ecosystem, you want which TLDs a given operator runs. TLD_OPERATOR-[:OPERATES]->TLD is the link, and the operator is the fast direction to anchor on.

cypher · runnablegraph.whisper.securitySign in to run
// TLDs operated by a registry, anchored on the operator
MATCH (op:TLD_OPERATOR {name: "NISSAN MOTOR CO., LTD."})-[:OPERATES]->(t:TLD)
RETURN op.name AS operator, collect(DISTINCT t.name) AS tlds
LIMIT 5

Returns: operator, tlds

Sample output:

json
[{"operator": "NISSAN MOTOR CO., LTD.", "tlds": ["datsun", "infiniti", "nissan"]}]

Costs: milliseconds; one anchored hop; operator names are exact strings, punctuation included.

A brand running several vanity TLDs (as here) is a neat illustration of the post-2012 gTLD landscape. MATCH (t:TLD) RETURN count(t) sizes the TLD population; it is a small label and safe to scan.

From here, → Which apexes are shared hosting in disguise?.

Which apexes are shared hosting in disguise?

Co-hosting results keep surfacing apexes with thousands of unrelated subdomains: CDN and multi-tenant platforms that make every tenant look like a neighbour. Identify them so you can weight them down before you report co-tenancy as a relationship.

cypher · runnablegraph.whisper.securitySign in to run
// Apexes whose certificate and subdomain fan-out looks like shared hosting
CALL whisper.threatIntel.candidateCdnApex(5)
YIELD apex, subCount, certCount, wildcardCount, recommendation
RETURN apex, subCount, certCount, wildcardCount, recommendation
LIMIT 5

Returns: apex, subCount, certCount, wildcardCount, recommendation

Sample output:

json
[
  {"apex": "microsoft.com", "subCount": 13021, "certCount": 54832, "wildcardCount": 9987, "recommendation": "add-to-deny-list"},
  {"apex": "narkive.com", "subCount": 7327, "certCount": 10614, "wildcardCount": 0, "recommendation": "add-to-deny-list-no-wildcard"}
]

Costs: milliseconds; a precomputed snapshot read; the argument is a number of candidates, not a domain name.

A high certCount relative to subCount means many independent certificates under one apex, the signature of a platform serving unrelated tenants. recommendation is the graph's own read on whether the apex is worth denying outright or just watching. Use the output as a suppression list in any co-tenancy study.

From here, → Tor-exit egress distribution by network.


Cross-layer studies

The payoff of a pre-joined graph is the join across layers that normally live in separate tools. A few research-shaped combinations.

Tor-exit egress distribution by network

OPERATES_EXIT_NODE links an IP to its TOR_RELAY identity (which survives IP rotation). Join Tor exits to the networks that route them to see where exit capacity concentrates. Sign in to run it.

cypher · runnablegraph.whisper.securitySign in to run
// Which networks host a sample of Tor exit IPs?
MATCH (ip:IPV4)-[:OPERATES_EXIT_NODE]->(relay:TOR_RELAY)
WITH ip LIMIT 1000
MATCH (ip)-[:ANNOUNCED_BY]->(ap:ANNOUNCED_PREFIX)-[:ROUTES]->(a:ASN)
RETURN a.name AS asn, count(DISTINCT ip) AS exit_ips
ORDER BY exit_ips DESC
LIMIT 15

Returns: asn, exit_ips

Sample output:

json
[
  {"asn": "AS62744", "exit_ips": 100},
  {"asn": "AS60729", "exit_ips": 60},
  {"asn": "AS53667", "exit_ips": 21}
]

Costs: under a second; a bounded sample of the exit population, then two hops into routing; the WITH ip LIMIT 1000 is what keeps it a sample.

From here, → Threat density across a network's prefixes.

Read coverage before band. Only known-clean — coverage: known-clean. In coverage, no malicious evidence. licenses the word "clean"; no-data — coverage: no-data. Not in coverage. This is not a verdict — nothing was looked at. means unknown, which is a different thing again; malicious-evidenced — coverage: malicious-evidenced. In coverage, with positive evidence of malice. and ambiguous — coverage: ambiguous. In coverage, and the evidence points both ways. mean there is evidence, whatever the band says. Full contract: Coverage — what we looked at.

Threat density across a network's prefixes

Aggregate threat is rolled up onto ANNOUNCED_PREFIX (threatScore, threatLevel) and onto the ASN itself (maxThreatScore, avgThreatScore, hasThreateningPrefixes), so you can study reputation distribution without walking every IP. For per-network triage, read the ASN aggregate directly; for the contributing prefixes, anchor and bound.

cypher
// Threat-listed prefixes within a network, ordered by aggregate score
MATCH (a:ASN {name: "AS13335"})-[:ROUTES]->(ap:ANNOUNCED_PREFIX)
WHERE ap.threatScore > 0
RETURN ap.name AS prefix, ap.threatScore AS score, ap.threatLevel AS level
ORDER BY score DESC
LIMIT 25

Returns: prefix, score, level

Sample output:

json
[{"prefix": "172.70.207.0/24", "score": 40, "level": "HIGH"}, {"prefix": "172.70.206.0/24", "score": 40, "level": "HIGH"}]

Costs: milliseconds; one anchored hop with a property filter.

For a comparable per-address measure across networks, CALL whisper.asnThreatDensity("AS13335") returns listedIps, announcedIpv4 and densityRatio. For a clean per-indicator verdict with its reasoning, prefer CALL explain("AS13335") over hand-walking ASN → PREFIX → IP → LISTED_IN; the manual walk does not finish on a large network. See explain().

From here, → TLS-fingerprint reuse across IPs.

TLS-fingerprint reuse across IPs

EMITS_TLS_FINGERPRINT ties an IP to a JA3/JARM fingerprint, useful for studying how a server signature spreads across address space. Seed it with an IP that actually carries a fingerprint: this is a thin plane, so most addresses have none.

cypher · runnablegraph.whisper.securitySign in to run
// IPs sharing a given JARM/JA3 fingerprint
MATCH (ip:IPV4 {name: "18.189.12.168"})-[:EMITS_TLS_FINGERPRINT]->(fp:TLS_FINGERPRINT)
WITH fp LIMIT 1
MATCH (other:IPV4)-[:EMITS_TLS_FINGERPRINT]->(fp)
RETURN fp.name AS fingerprint, count(DISTINCT other) AS ips_with_fingerprint
LIMIT 1

Returns: fingerprint, ips_with_fingerprint

Sample output:

json
[{"fingerprint": "jarm:07d14d16d21d21d07c42d41d00041d24a458a375eef0c576d23a7bab9a9fb1",
  "ips_with_fingerprint": 141}]

Costs: milliseconds; one anchored hop out and one back; pick the seed from the graph rather than from an incident.

Empty result: the layer holds 271 EMITS_TLS_FINGERPRINT edges, so expect no match on almost any indicator. A zero-row result here means Whisper holds no observation, not that the host shares no infrastructure. MATCH (ip:IPV4)-[:EMITS_TLS_FINGERPRINT]->(fp) RETURN ip.name, fp.name LIMIT 5 gives you a seed that works. Zero rows is never a verdict.

Anchor TLS_FINGERPRINT on its actual .name (the jarm:- or ja3:-prefixed hash). CALL whisper.lookupTlsFingerprint("jarm:…") classifies a hash you already hold.

From here, → Actor → ATT&CK technique map.

Actor → ATT&CK technique map

Named threat actors (ACTOR, case-sensitive) link to the MITRE techniques they use via USES_TECHNIQUE. Study an actor's technique footprint without leaving the graph.

cypher · runnablegraph.whisper.securitySign in to run
// MITRE ATT&CK techniques mapped to APT28 in public reporting
MATCH (actor:ACTOR {name: "APT28"})-[:USES_TECHNIQUE]->(t:ATTACK_PATTERN)
RETURN DISTINCT t.name AS technique
ORDER BY technique
LIMIT 50

Returns: technique

Sample output:

json
[{"technique": "Additional Email Delegate Permissions"}, {"technique": "Application Access Token"}, {"technique": "Archive Collected Data"}]

Costs: milliseconds; one anchored hop; techniques are ATTACK_PATTERN, never TECHNIQUE, and ACTOR.aliases holds the vendor names.

WhisperGraph carries the MITRE ATT&CK knowledge base as graph structure: 9,256 USES_TECHNIQUE edges from ACTOR to ATTACK_PATTERN and 872 USES_TACTIC edges, across 1,944 actors and 712 attack patterns. This is a curated reference layer, not Whisper's own attribution. It reflects what public reporting has mapped, not what Whisper observed. ATTRIBUTED_TO, the edge from an indicator to an actor, holds 302 edges; these queries return technique and tactic rollups. They do not attribute anything.

Convergence on a shared technique is a lead about the reporting, not about the infrastructure. For the infrastructure-side pivots (co-tenancy, shared registrant, nameserver siblings) see Campaign Pivoting.

From here, → What's actually in the graph to check the next label before you build on it.


Programmatic bulk runs

For aggregate studies you'll script the endpoint rather than click. Cypher over REST at https://graph.whisper.security/api/query:

bash
curl -s https://graph.whisper.security/api/query \
  -H "Content-Type: application/json" \
  -H "X-API-Key: $WHISPER_API_KEY" \
  -d '{"query":"UNWIND [\"AS13335\",\"AS3356\",\"AS15169\",\"AS2914\"] AS asn MATCH (a:ASN {name:asn})-[:BGP_NEIGHBOR]->(p:ASN) RETURN asn, count(p) AS degree ORDER BY degree DESC"}'

Each call returns columns, rows, and execution statistics as JSON, so you can fan a sample of anchors out across calls and reassemble the distribution locally. AI agents get the same surface via MCP at https://mcp.whisper.security; point any MCP client at it and it runs these queries mid-analysis (see AI & Agents).

For more patterns by workflow, see Workflows; for the complete label/edge/property model, the Graph Schema.