Syntax & Clauses
The Cypher clauses, operators, subqueries, parameters, and path syntax WhisperGraph supports, with a working example for each.
On this page (20)
Syntax & Clauses Documentation
WhisperGraph implements a read-only Cypher dialect. You send Cypher over HTTP and get back columns and rows. Write clauses (CREATE, MERGE, SET, DELETE, REMOVE, FOREACH) are not supported — the parser recognizes them and rejects them with a readonly_engine suggestion before anything runs.
Two rules make every query fast: anchor the starting node by its name (an indexed lookup), and add a LIMIT. This page walks each clause with a runnable example, then covers parameters, batching, and the plan you get from EXPLAIN. For the labels and edge directions you traverse, see the Graph Schema; for the procedures you call with CALL, the Procedures reference. For the golden rules and pitfalls, see Best Practices.
MATCH
MATCH finds patterns in the graph. The fastest form anchors a node by its name, which is an indexed lookup.
MATCH (h:HOSTNAME {name: "google.com"}) RETURN h.name
Chain a relationship to reach the node on the other end:
MATCH (a:ASN {name: "AS13335"})-[:ROUTES]->(p:ANNOUNCED_PREFIX)
RETURN p.name LIMIT 5
A label-only match with no {name: ...} scans every node of that label. That is fine on small labels like CATEGORY, but it never finishes on billion-node labels like HOSTNAME or IPV4 — always anchor those. State the label as well as the name: a bare MATCH (h {name: "..."}) is not planned the same way and can miss a sparsely connected name.
Names are stored lowercase, without a trailing dot, and they are matched exactly. Normalize in your own code before you anchor: Google.com does not reach the google.com node, and gmail.com. is not the same node as gmail.com. Names with special characters anchor as plain strings, so *spf.google.com and punycode names such as xn--80ak6aa92e.com need no escaping.
An unknown label or edge name in an anchored pattern is not an error: it matches nothing. When a correct-looking query returns zero rows, check CALL db.labels() and CALL db.relationshipTypes() first. Labels from other graph products are the one exception: Domain, IpAddress, and Certificate are rejected with a schema-drift error that names the label to use instead (HOSTNAME, IPV4, and CT_OBSERVATION).
OPTIONAL MATCH
OPTIONAL MATCH keeps the driving row even when the optional pattern has no match, filling the missing columns with null. Use it for sparse fields like WHOIS contacts or geolocation, where a plain MATCH would drop the whole row.
MATCH (h:HOSTNAME {name: "google.com"})
OPTIONAL MATCH (h)-[:HAS_EMAIL]->(e:EMAIL)
OPTIONAL MATCH (h)-[:HAS_REGISTRAR]->(r:REGISTRAR)
RETURN h.name, collect(DISTINCT e.name) AS emails, collect(DISTINCT r.name) AS registrars
WHERE
WHERE filters bound rows. The supported operators:
- Comparison:
=,<>,<,>,<=,>= - Logical:
AND,OR,NOT,XOR - Null checks:
IS NULL,IS NOT NULL - List membership:
IN - String predicates:
STARTS WITH,ENDS WITH,CONTAINS,=~(regex)
MATCH (h:HOSTNAME)
WHERE h.name STARTS WITH "cloudflare."
RETURN h.name LIMIT 5
MATCH (h:HOSTNAME)
WHERE h.name ENDS WITH ".cloudflare.com"
RETURN h.name LIMIT 5
MATCH (a:ASN)
WHERE a.name IN ["AS13335", "AS15169"]
RETURN a.name
STARTS WITH and ENDS WITH on .name are index-backed. Keep the suffix narrow and leading-dot (ENDS WITH ".cloudflare.com", never ENDS WITH "com"), and remember that only hostnames carry the suffix index: on PREFIX or ASN, anchor instead. CONTAINS is fine once the query is anchored or paired with STARTS WITH; never run it across an unanchored label. For a token you cannot classify, use CALL whisper.search("token"), which routes to an indexed lookup and never scans. On ASN, .name is the AS number (AS13335), so match it exactly or with STARTS WITH "AS", not with CONTAINS.
Regex is a full match against the whole value and only gets an index when it is a plain prefix or a .*literal.* shape, so prefer the string predicates and use =~ on rows you have already anchored:
MATCH (a:ASN {name: "AS13335"}) WHERE a.name =~ "AS[0-9]+" RETURN a.name
"AS133" would not match AS13335; the pattern has to cover the whole name. An IN list is rewritten into indexed lookups; for a long list, switch to UNWIND (below).
RETURN
RETURN selects what comes back. Use AS for aliases and DISTINCT to deduplicate. RETURN * returns every bound variable, and literals of any type (numbers, strings, booleans, lists, maps) can be returned directly.
MATCH (h:HOSTNAME {name: "google.com"})-[:RESOLVES_TO]->(i)
RETURN DISTINCT labels(i)[0] AS family
One case to plan around: across a chain that runs through an announced prefix (ANNOUNCED_BY, then ROUTES), aggregate with count(DISTINCT ...) or de-duplicate in your client rather than writing RETURN DISTINCT over the projection. Best Practices has the pattern.
WITH
WITH pipes results from one part of a query to the next. It is how you aggregate or narrow a set before traversing further, and you can filter after it with WHERE.
MATCH (h:HOSTNAME {name: "google.com"})<-[:NAMESERVER_FOR]-(ns:HOSTNAME)
WITH ns LIMIT 3
MATCH (ns)-[:NAMESERVER_FOR]->(sibling:HOSTNAME)
RETURN ns.name AS nameserver, collect(DISTINCT sibling.name)[0..8] AS domains
LIMIT 3
This anchor-then-narrow-then-expand shape is the single most useful pattern in the language. Bounding the intermediate set with WITH ... LIMIT keeps a two-stage query from exploding, and it puts the bound where the fan-out happens: a LIMIT at the end of the query does not bound the traversal that feeds it.
ORDER BY, LIMIT, SKIP
ORDER BY sorts, LIMIT caps the row count, and SKIP offsets for pagination. Always include a LIMIT, and use literal numbers in SKIP / LIMIT.
MATCH (sub:HOSTNAME)-[:CHILD_OF]->(:HOSTNAME {name: "google.com"})
RETURN sub.name AS subdomain
ORDER BY sub.name SKIP 0 LIMIT 15
To page, keep a stable ORDER BY and walk SKIP forward: SKIP 0 LIMIT 15, then SKIP 15 LIMIT 15, and so on. LIMIT 0 returns no rows, and a SKIP past the end returns an empty page. If you bind SKIP or LIMIT to a parameter and the value resolves to null, the bound is dropped and the response carries a null-pagination-param advisory; pass a number.
Read
coveragebeforeband. Only known-clean — coverage: known-clean. In coverage, no malicious evidence. licenses the word "clean"; no-data — coverage: no-data. Not in coverage. This is not a verdict — nothing was looked at. means unknown, which is a different thing again; malicious-evidenced — coverage: malicious-evidenced. In coverage, with positive evidence of malice. and ambiguous — coverage: ambiguous. In coverage, and the evidence points both ways. mean there is evidence, whatever the band says. Full contract: Coverage — what we looked at.
UNWIND
UNWIND turns a list into rows, one per element. It is the right pattern for batch lookups: each element becomes its own anchored query.
UNWIND ["185.220.101.1", "104.16.123.96", "8.8.8.8"] AS addr
MATCH (ip:IPV4 {name: addr})
RETURN ip.name AS ip, ip.threatLevel AS level, ip.isThreat AS isThreat
The MATCH after UNWIND is still anchored — each row binds ip on its indexed name. UNWIND followed by CALL runs a procedure once per element (see below), and UNWIND $names AS n MATCH (h:HOSTNAME {name: n}) is the form to use when an IN list grows long.
UNION
UNION combines results from multiple queries and deduplicates; UNION ALL keeps duplicates. Every branch must return the same column names, and each branch can carry its own LIMIT.
MATCH (a:ASN {name: "AS13335"})-[:ROUTES]->(p) RETURN p.name AS n LIMIT 5
UNION
MATCH (a:ASN {name: "AS13335"})-[:BGP_NEIGHBOR]-(peer:ASN) RETURN peer.name AS n LIMIT 5
CALL procedures
CALL runs a procedure. Standalone, or with YIELD to name the columns you want and feed the rest of the query. See Procedures for the full set.
CALL explain("1.1.1.1") YIELD indicator, score, level
RETURN indicator, score, level LIMIT 1
A bare CALL explain("1.1.1.1") returns every column the procedure defines, including transport fields such as available and cached and an advisory slot that stays empty when there is nothing to advise. Name the columns you read with YIELD, as above, and the result stays stable as the procedure grows.
CALL whisper.variants("paypal.com")
YIELD variant, method, exists
WHERE exists
RETURN variant, method LIMIT 10
A CALL placed after UNWIND, WITH, or MATCH runs once per incoming row, so you can score a whole list in one query:
UNWIND ["1.1.1.1", "8.8.8.8"] AS ip
CALL explain(ip) YIELD indicator, score, level
RETURN indicator, score, level LIMIT 2
Three rules keep procedure calls out of trouble:
- Quote every argument.
CALL whisper.identify(ubuntu.com)is a bad-argument error, and an unquoted IPv6 literal is parsed as something else entirely. AlwaysCALL whisper.identify("ubuntu.com"). YIELDcolumns are exact contracts. A column the procedure does not emit is rejected, not ignored, and the error lists the columns it does emit.db.relationshipTypes()emitstype, notrelationshipType.YIELD *is rejected on a procedure whose columns depend on what you passed in.explainandwhisper.historyare both multi-shape. Name columns from one shape, or call the single-shape variant:whisper.history.whois(domain)for WHOIS columns,whisper.history.bgp(ip|asn|prefix)for routing columns.
Schema-introspection procedures are cheap and answer immediately, so they are the fastest way to confirm a label or edge exists before you anchor on it:
CALL db.labels() YIELD label RETURN label ORDER BY label LIMIT 20
CALL db.relationshipTypes() YIELD type, count RETURN type, count ORDER BY type LIMIT 5
CALL subqueries
CALL { ... } scopes a subquery, importing outer variables with WITH. A standalone CALL { ... } with no preceding clause is not allowed — give it an importing clause. Bounding each branch inside its own CALL {} is the reliable way to run several per-branch aggregations from one anchor, and the subquery's own LIMIT bounds each branch on its own. It is also where a multi-hop routing leg belongs after a WITH: keep ANNOUNCED_BY and ROUTES together inside the subquery, or give each WITH stage one computed hop.
MATCH (a:ASN {name: "AS13335"})
CALL {
WITH a
MATCH (a)-[:ROUTES]->(p:ANNOUNCED_PREFIX)
RETURN count(p) AS pc
}
RETURN a.name, pc
MATCH (h:HOSTNAME {name: "google.com"})
CALL {
WITH h
MATCH (h)-[:RESOLVES_TO]->(ip:IPV4)
RETURN ip LIMIT 2
}
RETURN h.name, ip.name
EXISTS and COUNT subqueries
EXISTS { ... } tests whether a pattern has at least one match without binding it, and NOT EXISTS { ... } negates it. COUNT { ... } returns how many matches the pattern has, as an expression you can project or filter on.
MATCH (h:HOSTNAME {name: "google.com"})
WHERE EXISTS { MATCH (h)-[:RESOLVES_TO]->(:IPV4) }
RETURN h.name
MATCH (h:HOSTNAME {name: "google.com"})
WHERE NOT EXISTS { MATCH (h)-[:SPF_EXISTS]->(:HOSTNAME) }
RETURN h.name
MATCH (h:HOSTNAME {name: "google.com"})
RETURN h.name, COUNT { MATCH (h)-[:RESOLVES_TO]->(:IPV4) } AS ipv4_count
List and pattern comprehensions
A list comprehension filters and maps a list inline: [x IN list WHERE predicate | expression]. A pattern comprehension does the same over a graph pattern, collecting one value per match, and it can be sliced like any list.
RETURN [x IN [1, 2, 3] WHERE x > 1 | x * 10] AS scaled
MATCH (a:ASN {name: "AS13335"})
RETURN a.name, [(a)-[:ROUTES]->(p) | p.name][0..3] AS prefixes
The slice runs after the list is built, so a pattern comprehension over a wide fan-out collects everything first. Bound the anchor set with WITH ... LIMIT before you comprehend over it.
CASE
CASE is a conditional expression, in both the simple form (match a value) and the searched form (evaluate conditions).
MATCH (a:ASN {name: "AS13335"})
RETURN CASE a.name WHEN "AS13335" THEN "cloudflare" ELSE "other" END AS label
UNWIND [1, 5, 10] AS x
RETURN CASE WHEN x < 3 THEN "low" WHEN x < 8 THEN "mid" ELSE "high" END AS bucket
Parameters
Write $name placeholders in the query and send the values in the request body's parameters object (the field is parameters, not params). Parameters keep the query plan cacheable and save you from escaping quotes inside JSON. They are available on POST /api/query only.
MATCH (h:HOSTNAME {name: $n}) RETURN h.name AS n
{"query": "MATCH (h:HOSTNAME {name: $n}) RETURN h.name AS n", "parameters": {"n": "google.com"}}
A query that references $n without a matching entry in parameters is rejected with 400 query-error and Missing parameter: $n; the placeholder is never treated as a literal. See Parameter binding for the request in five languages.
Batching statements
Several statements separated by a top-level ; in one query string run as a batch, and the response changes shape: instead of one columns/rows envelope you get a results array with one entry per statement.
RETURN 1 AS a; RETURN 2 AS b
Each entry carries the statement's own result (columns, rows, rowCount, executionTimeMs), an outcome (OK, PARSE_ERROR, EXECUTION_ERROR, or DEADLINE_EXCEEDED), and a success flag. A failed statement adds errorMessage and errorType and does not stop the others, so read outcome per entry rather than the HTTP status. Splitting happens on top-level semicolons only; one inside a string literal is left alone. Details in Batch statements.
Patterns & paths
A pattern is a sequence of node and relationship descriptions. The forms:
- Directed —
(:HOSTNAME {name: "google.com"})-[:RESOLVES_TO]->(ip) - Reverse —
(:HOSTNAME {name: "google.com"})<-[:MAIL_FOR]-(mx)walks the edge backward (mail and nameserver edges point server → domain) - Undirected —
(:ASN {name: "AS13335"})-[:BGP_NEIGHBOR]-(peer)matches either arrow; use it for peering, which is symmetric in practice - Multi-type —
(:HOSTNAME {name: "google.com"})-[:RESOLVES_TO|HAS_EMAIL]->(x)matches either edge - Variable-length —
(:HOSTNAME {name: "www.mail.google.com"})-[:CHILD_OF*1..3]->(parent)walks one to three hops - Named paths — bind the path with
p = (...)to usenodes(),relationships(), andlength() - Anonymous nodes —
(:ASN {name: "AS13335"})-[:ROUTES]->()when you don't need the far node bound
MATCH p = (h:HOSTNAME {name: "github.com"})-[:RESOLVES_TO]->(:IPV4)-[:ANNOUNCED_BY]->(:ANNOUNCED_PREFIX)
RETURN length(p) AS hops, [n IN nodes(p) | n.name] AS chain LIMIT 5
MATCH (h:HOSTNAME {name: "google.com"})-[r:RESOLVES_TO|HAS_EMAIL]->(x)
RETURN type(r) AS edge, x.name AS target LIMIT 5
When you expand outward from an announced prefix (ANNOUNCED_PREFIX, REGISTERED_PREFIX), give the relationship a type (<-[:ROUTES]-, -[:CONFLICTS_WITH]->) or start from the PREFIX-labelled node with the same name; an untyped, undirected -[r]- from a computed prefix is not a supported shape. On stored labels such as HOSTNAME the untyped form is fine.
Variable-length relationships
Repeat an edge between a lower and an upper bound. Always write the upper bound yourself: a bare [*] is read as a very deep walk, which is almost never what you meant.
MATCH (h:HOSTNAME {name: "www.mail.google.com"})-[:CHILD_OF*1..3]->(parent)
RETURN parent.name LIMIT 5
The chain ends at the TLD: www.mail.google.com → mail.google.com → google.com → com. A range wider than the label chain costs nothing extra, but it also finds nothing extra.
Edges computed at query time (BGP_NEIGHBOR, ROUTES, ANNOUNCED_BY, LISTED_IN, BELONGS_TO, CONFLICTS_WITH) work inside [*1..N] as long as one endpoint is anchored, and length(p) reports correctly. Keep the range tight, and on a peering walk filter WHERE n <> a: a mesh routinely walks back to the origin AS.
MATCH p = (a:ASN {name: "AS13335"})-[:BGP_NEIGHBOR*1..2]-(n:ASN)
WHERE n <> a
RETURN n.name AS asn, length(p) AS distance LIMIT 5
Over a high-fan-out edge, prefer explicit single hops joined with WITH ... LIMIT: the variable-length form expands every intermediate row before it returns anything. See Best Practices.
shortestPath
shortestPath finds the minimum-hop path between two anchored nodes. Write the variable-length range explicitly and keep it tight; both endpoints must be anchored, and when no path exists within the bound the result is empty rather than an error.
MATCH p = shortestPath(
(a:HOSTNAME {name: "www.google.com"})-[*1..4]-(b:HOSTNAME {name: "google.com"})
)
RETURN length(p) AS hops
EXPLAIN
EXPLAIN returns the query plan without executing it, as a single plan column holding the operator tree. Use it to confirm an anchored lookup hits the index rather than scanning.
EXPLAIN MATCH (h:HOSTNAME {name: "google.com"})-[:RESOLVES_TO]->(ip) RETURN ip.name
A NodeLookup at the leaf of the plan means the anchor is indexed. A bare label scan there is a warning that the query will be slow — anchor it before you run it.
PROFILE
PROFILE runs the query and returns the same rendered plan together with the rows it produced and its executionTimeMs, so you can see what a query costs before you put it in a loop.
PROFILE MATCH (h:HOSTNAME {name: "google.com"})-[:RESOLVES_TO]->(ip) RETURN ip.name LIMIT 1